Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,5 +25,5 @@
- `in2lambda convert FILE PartsOneSol` now exports the worked solution a document writes in a `solution` environment. Pandoc writes that environment as a Div whose classes hold `solution`, and the filter recognised only a Div whose first block reads `Solution`, so a document using the environment exported every question with an empty worked solution. A Div whose first block reads `Solution` is still recognised.
- Importing `in2lambda.katex_convert` no longer writes a file called `log` into the working directory. That module reports what it changed in an expression to the `in2lambda.katex_convert` logger, which is silent unless the application configures logging.
- `in2lambda convert` now reads a .docx that holds an image. in2lambda looks in the document for the directories a `\graphicspath` names, and read the document as UTF-8 text to find them. A .docx is a zip file, so converting a Word document holding a figure raised `UnicodeDecodeError`. in2lambda now reads a document that is not UTF-8 text as naming no directory, which is what a .docx names.
- `in2lambda compare BUILT_ZIP EXPORT_DIR` compares two Lambda Feedback sets, each given as a folder or a zip, so that a set in2lambda wrote can be checked against the export it should reproduce. Each question's main text is compared, and each part's text and worked solution, and every difference is printed naming the question, the part and the field. Three differences in wording are taken off both sides first: a run of whitespace is compared as one space, an image is compared by the file's name, and a part holding neither text nor a worked solution is dropped where it is the question's only part. `--known FILE` names the differences the two sets are known to have, one line per difference with the ticket that would close it written after ` # `, and `in2lambda compare` exits 1 where the differences found are not the differences that file names. `in2lambda.compare.differences` and `in2lambda.compare.known` are the two functions behind the command.
- `in2lambda compare BUILT_ZIP EXPORT_DIR` compares two Lambda Feedback sets, each given as a folder or a zip, so that a set in2lambda wrote can be checked against the export it should reproduce. Each question's main text is compared, and each part's text and worked solution, and every difference is printed naming the question, the part and the field. The differences in wording that are not differences in what a question says are taken off both sides first: a run of whitespace is compared as one space, a line holding nothing but hyphens is dropped, each curly quote is compared as the straight quote, an image is compared by the file's name, and a part holding neither text nor a worked solution is dropped where it is the question's only part. Inside every `$ ... $` and `$$ ... $$`, `\left` and `\right` are removed, `~`, `\,` and `\space` are compared as a space, and every run of whitespace is dropped, except that a run between a control word and a following letter is compared as one space, so that `$z=2+3 i$` and `$z=2+3i$` are the same expression and `$\alpha x$` and `$\alphax$` are two. `--known FILE` names the differences the two sets are known to have, one line per difference with the ticket that would close it written after ` # `, and `in2lambda compare` exits 1 where the differences found are not the differences that file names. `in2lambda.compare.differences` and `in2lambda.compare.known` are the two functions behind the command.
- The rest of the Python API is unchanged: `in2lambda.main.runner` and everything under `in2lambda.api` take the same arguments and return the same objects.
76 changes: 70 additions & 6 deletions in2lambda/compare.py
Original file line number Diff line number Diff line change
@@ -1,17 +1,30 @@
"""Compares two sets question by question, naming every place they say something else.
r"""Compares two sets question by question, naming every place they say something else.

`in2lambda convert` writes a set, `in2lambda build` writes a set from a draft, and Lambda
Feedback exports a set. :func:`differences` compares any two of them in question and part
order - each question's main text, and each part's text and worked solution - and returns
one line per difference, naming the question, the part and the field as
`in2lambda.validation` names them.

Three differences in wording are not differences in what a question says, and are taken
These differences in wording are not differences in what a question says, and are taken
off both sides before comparing:

- **Whitespace.** Every run of whitespace is compared as one space, because a draft
quotes the lines pandoc wrapped where `in2lambda convert` writes a paragraph on one
line.
- **Separator lines.** A line holding nothing but three or more hyphens is dropped,
because Lambda Feedback writes one around a display maths block where a document
writes nothing.
- **Quotes.** ``‘`` and ``’`` are compared as ``'``, and ``“`` and ``”`` as ``"``,
because pandoc's LaTeX reader writes the curly quote where its commonmark_x writer
writes the straight one.
- **Maths notation.** Inside every ``$ ... $`` and ``$$ ... $$``, ``\left`` and
``\right`` are removed, ``~``, ``\,`` and ``\space`` are compared as a space, and
every run of whitespace is dropped, except that a run between a control word and a
following letter is compared as one space. LaTeX renders ``$z=2+3 i$`` and
``$z=2+3i$`` the same, and ``\mathrm{~m}`` and ``\mathrm{m}`` the same, where
``\alpha x`` and ``\alphax`` are two different expressions. An export writes
``$z = 2+3 i$`` where `in2lambda convert` writes ``$z=2+3i$``.
- **Image references.** An image is compared by the file's name, because
`in2lambda convert` writes every image as ``![pictureTag](path)`` where a draft keeps
the alt text the document wrote, and an export names each file as ``media/`` holds it
Expand All @@ -25,26 +38,77 @@
after `` # ``.
"""

import re
from itertools import zip_longest
from pathlib import Path
from typing import Any, Optional

from in2lambda.api.question import Question
from in2lambda.api.set import Set
from in2lambda.json_convert.json_convert import _IMAGE
from in2lambda.validation import _location
from in2lambda.validation import _COMMAND, _MATHS, _location

_TICKET = " # "
"""What a line of a differs.txt names the ticket closing it after."""

_RULE = re.compile(r"(?m)^[ \t]*(?:-{3,}|\*{3,}|_{3,})[ \t]*$")
"""A line holding nothing but three or more hyphens, asterisks or underscores: a markdown
separator, which Lambda Feedback writes around a display maths as `---` or `***`."""

_QUOTES = str.maketrans({"‘": "'", "’": "'", "“": '"', "”": '"'})
"""Each curly quote and the straight quote it is compared as."""

_SIZE = re.compile(r"\\(?:left|right)(?![a-zA-Z])")
r"""``\left`` and ``\right``, which size a delimiter without changing which it is."""

_LATEX_SPACE = re.compile(r"~|\\,|\\space(?![a-zA-Z])")
"""The three ways of writing a space inside maths."""

_SPACING = re.compile(rf"({_COMMAND.pattern})\s+(?=[a-zA-Z])|\s+")
"""A run of whitespace inside maths, with the control word it ends where one precedes it
and a letter follows it."""


def _spacing(whitespace: re.Match[str]) -> str:
r"""One space where a run of whitespace ends a control word, and nothing elsewhere.

``\alpha x`` is two symbols and ``\alphax`` is a control word nothing defines, so
the space between a control word and a letter is the only whitespace LaTeX renders.
"""
return f"{whitespace[1]} " if whitespace[1] else ""


def _maths(expression: re.Match[str]) -> str:
"""One ``$ ... $`` or ``$$ ... $$`` with the notation that is not the maths folded.

Args:
expression: A match of `in2lambda.validation._MATHS`, holding the display maths
it found in its first group and the inline maths in its second.
"""
display = expression[1] is not None
tex = expression[1] if display else expression[2]
delimiter = "$$" if display else "$"
folded = _LATEX_SPACE.sub(" ", _SIZE.sub("", tex))
return f"{delimiter}{_SPACING.sub(_spacing, folded)}{delimiter}"


def _text(markdown: str) -> str:
"""A field with the differences in wording that are not differences taken off.

Every run of whitespace becomes one space, and every image reference is written as
the file's name alone. The module docstring says why.
A separator line is dropped, each curly quote becomes a straight quote, the maths
notation inside every ``$ ... $`` and ``$$ ... $$`` is folded, every image reference
is written as the file's name alone, and every run of whitespace becomes one space.
The module docstring says why.
"""
named = _IMAGE.sub(lambda reference: f"![]({Path(reference[1]).name})", markdown)
# Before the whitespace collapse below, which writes the field on one line and
# leaves no line for _RULE to match.
# The platform writes a space as the entity ` ` (and a hard one as ` `).
without_rules = _RULE.sub("", markdown.replace(" ", " ").replace(" ", " "))
folded = _MATHS.sub(_maths, without_rules.translate(_QUOTES))
# A space touching a maths delimiter from outside renders the same either way:
# `of $y$` and `of$y$` are one expression, so the whitespace beside `$` is dropped.
folded = re.sub(r"\s*(\${1,2})\s*", r"\1", folded)
named = _IMAGE.sub(lambda reference: f"![]({Path(reference[1]).name})", folded)
return " ".join(named.split())


Expand Down
13 changes: 8 additions & 5 deletions in2lambda/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -638,12 +638,15 @@ def compare(built_zip: str, export_dir: str, known_path: Optional[str]) -> None:

Each argument is a Lambda Feedback set, as a folder or as a zip. Each question's main
text is compared, and each part's text and worked solution, and every difference is
printed naming the question, the part and the field. Three differences in wording are
taken off both sides first: a run of whitespace is compared as one space, an image is
printed naming the question, the part and the field. The differences in wording that
are not differences in what a question says are taken off both sides first: a run of
whitespace is compared as one space, a line of hyphens is dropped, a curly quote is
compared as a straight quote, the notation inside maths is folded, an image is
compared by the file's name, and a lone empty part is dropped. in2lambda.compare says
why. --known names a file of the differences the two sets are known to have, one per
line as this command prints it, with a ticket written after " # ". in2lambda compare
exits 1 where the differences found are not the differences --known names.
which notation and why. --known names a file of the differences the two sets are known to
have, one per line as this command prints it, with a ticket written after " # ".
in2lambda compare exits 1 where the differences found are not the differences --known
names.
"""
with _message_not_traceback():
found = in2lambda.compare.differences(
Expand Down
2 changes: 0 additions & 2 deletions tests/fixtures/against_convert/PartsOneSol/differs.txt

This file was deleted.

18 changes: 13 additions & 5 deletions tests/fixtures/against_convert/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,12 +26,24 @@ document itself, as the ranges in `fixtures/sources` are. They were produced wit
The comparison is `in2lambda.compare.differences`, whose docstring states these rules, so
that this README and the code do not drift apart.

Each question's main text is compared, and each part's text and worked solution. Three
Each question's main text is compared, and each part's text and worked solution. These
differences between the routes are not differences in what a question says, and are taken
off both sides before comparing:

- **Line breaks.** The draft quotes the lines pandoc wrapped; convert writes a paragraph
on one line. Every run of whitespace is compared as one space.
- **Separator lines.** A line holding nothing but three or more hyphens is dropped,
because Lambda Feedback writes one around a display maths block where a document
writes nothing.
- **Quotes.** Pandoc's LaTeX reader writes `’` where its `commonmark_x` writer writes
`'`, so `aren’t` and `aren't` are the same wording. Each curly quote is compared as
the straight quote.
- **Maths notation.** Inside every `$ ... $` and `$$ ... $$`, `\left` and `\right` are
removed, `~`, `\,` and `\space` are compared as a space, and every run of whitespace
is dropped, except that a run between a control word and a following letter is
compared as one space. LaTeX renders `$z=2+3 i$` and `$z=2+3i$` the same, and
`\mathrm{~m}` and `\mathrm{m}` the same, where `\alpha x` and `\alphax` are two
different expressions.
- **Image references.** Convert writes every image as `![pictureTag](path)`; the draft keeps
the alt text the document wrote, which is empty for `\includegraphics`. The set read back
from the zip names each file as it sits in the export's `media/`, where the set convert
Expand All @@ -50,7 +62,3 @@ A folder with no such file is a document the two routes say the same thing about

Finding a difference the file does not list fails the test, and so does agreeing where it
lists one: closing a ticket below means deleting the lines it names.

- **Smart quotes** - pandoc's LaTeX reader writes `’` where its `commonmark_x` writer
writes `'`, so convert uploads `aren’t` for `PartsOneSol/example.tex` and the draft
uploads `aren't`.
71 changes: 67 additions & 4 deletions tests/test_compare.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,51 @@ def test_a_run_of_whitespace_is_one_space() -> None:
)


def test_a_separator_line_is_dropped() -> None:
"""Lambda Feedback writes a `---` line where a document writes nothing."""
assert (
differences(
_set("The mass is\n\n---\n\n$$m = 1$$"), _set("The mass is\n$$m=1$$")
)
== []
)


def test_a_curly_quote_is_a_straight_quote() -> None:
"""Pandoc's LaTeX reader writes `’` where its commonmark_x writer writes `'`."""
assert differences(_set("It isn’t large."), _set("It isn't large.")) == []
assert differences(_set("The “load”."), _set('The "load".')) == []


def test_whitespace_inside_maths_is_dropped() -> None:
"""LaTeX renders `$z=2+3 i$` and `$z=2+3i$` the same."""
assert differences(_set("$z = 2+3 i$"), _set("$z=2+3i$")) == []
assert differences(_set("$$\nF = pA\n$$"), _set("$$F=pA$$")) == []


def test_a_space_between_a_control_word_and_a_letter_is_kept() -> None:
r"""`\alpha x` is two symbols and `\alphax` is a control word nothing defines."""
assert differences(_set(r"$\alpha x$"), _set(r"$\alpha x$")) == []

assert differences(_set(r"$\alpha x$"), _set(r"$\alphax$")) == [
'Question 1 "", main text: the draft says '
r"'$\\alpha x$' and convert says '$\\alphax$'"
]


def test_a_sized_delimiter_is_the_delimiter() -> None:
r"""`\left(` and `(` render the same bracket."""
assert differences(_set(r"$\left( x+1 \right)$"), _set("$(x+1)$")) == []


def test_each_latex_space_is_a_space() -> None:
r"""`~`, `\,` and `\space` are the three ways of writing a space inside maths."""
assert differences(_set("$a~b$"), _set("$ab$")) == []
assert differences(_set(r"$a\,b$"), _set("$ab$")) == []
assert differences(_set(r"$a\space b$"), _set("$ab$")) == []
assert differences(_set(r"$5\mathrm{~m}$"), _set(r"$5\mathrm{m}$")) == []


def test_an_image_is_compared_by_the_file_name() -> None:
"""The alt text and the directory differ between the routes; the file name does not."""
assert (
Expand Down Expand Up @@ -67,12 +112,12 @@ def test_a_difference_names_the_question_the_part_and_the_field() -> None:
expected = _set("Find the load.", ("State the pressure.", "$F = 2pA$"))

assert differences(built, expected) == [
"Question 1 \"\", part (a), worked solution: the draft says '$F = pA$' and "
"convert says '$F = 2pA$'"
"Question 1 \"\", part (a), worked solution: the draft says '$F=pA$' and "
"convert says '$F=2pA$'"
]
assert differences(built, expected, "set.zip", "the export") == [
"Question 1 \"\", part (a), worked solution: set.zip says '$F = pA$' and "
"the export says '$F = 2pA$'"
"Question 1 \"\", part (a), worked solution: set.zip says '$F=pA$' and "
"the export says '$F=2pA$'"
]


Expand Down Expand Up @@ -164,3 +209,21 @@ def test_the_command_refuses_a_path_that_is_not_a_set(
result.exception, (ValueError, zipfile.BadZipFile)
), result.output
assert f"{path} is not a Lambda Feedback set" in result.output


def test_an_html_space_entity_is_a_space() -> None:
a = _set("$16y''-\\pi^2y=0$ (Use $A$ and $B$ for your constants.)")
b = _set("$16y''-\\pi^2y=0$     (Use $A$ and $B$ for your constants.)")
assert differences(a, b) == []


def test_whitespace_touching_a_maths_delimiter_from_outside_is_ignored() -> None:
a = _set("equation of $y''+y'-6y=0$: then")
b = _set("equation of$y''+y'-6y=0$: then")
assert differences(a, b) == []


def test_a_separator_of_asterisks_is_dropped_like_one_of_hyphens() -> None:
a = _set("Then:\n\n$$\nx=1\n$$\n\n***\n\nRecall the rule.")
b = _set("Then:\n\n$$\nx=1\n$$\n\nRecall the rule.")
assert differences(a, b) == []
Loading