diff --git a/CHANGELOG.md b/CHANGELOG.md index c379426..e47c09a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,7 @@ - beartype is now `^0.22`. At 0.20.0 and below, the beartype import hook leaves `cli` a plain function instead of a group, so the command line either fails to import or runs `convert` whatever the arguments are. 0.20.1 is the first version that works. - `in2lambda source add FILE` freezes a document: `in2lambda source add` converts .docx and .tex to markdown beside the file, and writes `FILE.draft.json` beside the file, holding the markdown's hash and every block in the markdown with the lines it spans, so that another tool can quote the source by line range. The draft is named after the source, so a folder holding a term's sheets holds one draft per sheet. `in2lambda source show` prints that markdown numbered with the block ids. Freezing a file that has changed since is refused unless `--start-over` discards the draft. Showing a file that has changed is refused as well, because its block ids would name lines they were not taken from. `in2lambda source add` needs pandoc and the `convert` extra, as `in2lambda convert` does. `in2lambda source show` reads the markdown the draft was frozen from, and needs neither. - `in2lambda source add` writes the markdown of a converted .docx or .tex unwrapped, so that a paragraph is one line however long the paragraph is, and an inline `$ ... $` is never broken across two lines. `in2lambda source add` also moves every `$$ ... $$` that pandoc wrote on one line onto lines of its own, in a paragraph and in a list item, indented to the item's width where the maths is written in a list item. A `$$` that opens or closes in a table cell, in a block quote or in a code block is left unchanged, as is the maths that an unpaired `$$` elsewhere in the document, such as one in inline code, pairs with. `in2lambda validate` reports the maths left unchanged. pandoc's writer produces both forms, and Lambda Feedback renders neither, so `in2lambda validate` reported them against every converted sheet. The markdown beside a document frozen before this release differs, and so do its line ranges. `in2lambda source add --start-over` freezes the document again. +- `in2lambda source add` now drops the underline, highlight and small capitals a converted .docx holds, keeping the text each marked, so a field reads `**Question 2:**` in place of `**[Question 2:]{.underline}**`. Pandoc's docx reader writes those three formats as the bracketed spans `[text]{.underline}`, `[text]{.mark}` and `[text]{.smallcaps}`, which the PDF generator's pandoc turns into LaTeX commands its template does not define, so a set built from such a field failed to compile. Lambda Feedback's markdown renders none of the three. A bracketed span in a code block or in inline code is left unchanged. The markdown beside a .docx frozen before this release differs, and so do its line ranges where a span stood; `in2lambda source add --start-over` freezes the document again. - A draft now holds a `log` of every command that changed it, and a `fields` map of the fields those commands wrote. Each field records the layer that wrote it (1 a spec, 2 a predicate, 3 a line range, 4 a literal), the source ranges it was copied from, whether it has been edited, and by whom. `in2lambda draft mark ignore BLOCK` is the first such command. `in2lambda draft replay` rebuilds the draft from the frozen markdown and the log, and refuses unless the draft it builds matches the draft on disk byte for byte. A draft written before this release holds no `log` and is refused as a draft in2lambda did not write; `in2lambda source add --start-over` freezes the document again. - `in2lambda draft question add`, `in2lambda draft part add QUESTION` and `in2lambda draft question solution QUESTION` fill a draft in. Each takes `--text` to copy the wording out of the frozen source, as a block id such as `b3` or as lines such as `s10:14`. Each takes `--literal TEXT` for wording the source does not hold in a form the field can take, which records the field as edited and written by layer 4 in place of layer 3. in2lambda works the question and part numbers out from the fields already written, so a replay arrives at the same ids. `in2lambda draft split block BLOCK AT` cuts a block the parser made one of two things into `b3a` and `b3b`, so that each half can be quoted on its own. A command writing a field that is already written is refused, naming the field. A command quoting lines another field was taken from is refused, naming both fields. - `in2lambda draft part solution PART --text RANGE` gives one part of a question its worked solution, for a sheet that writes a solution under each part rather than one solution answering the whole question. PART is written `q1.p2`. The command writes `q1.p2.solution` from the lines named, or from `--literal TEXT`, as `in2lambda draft question solution` writes a question's solution. A PART naming a question, or naming a part the draft has not written, is refused, naming the part. A sheet laid out this way can now be answered part by part by commands, which only a spec's `PartPartSolSol` and `PartSolPartSol` layouts could do before, so the warning `in2lambda validate` prints about a part nothing answers can be acted on. diff --git a/in2lambda/source/__init__.py b/in2lambda/source/__init__.py index f493c17..53dd491 100644 --- a/in2lambda/source/__init__.py +++ b/in2lambda/source/__init__.py @@ -97,6 +97,12 @@ def _field_fault(field: Any) -> str: The writer and the reader must agree. Pandoc's ``markdown`` writer emits fenced divs and bracketed spans that a commonmark reader reads as ordinary text. ``commonmark_x`` also covers the ``$...$`` maths and the ``{width=...}`` attributes a converted document holds. + +``commonmark_x`` does write the bracketed spans a .docx holds, which +:func:`_spans_unwrapped` unwraps afterwards. The ``-bracketed_spans`` the writer takes +does not help: with ``raw_html`` on, which ``commonmark_x`` keeps, the writer falls back +to ``Question 2:`` and ``oil``, and turning ``raw_html`` +off as well changes how a figure is written. """ _POSITION = re.compile(r"(?:[^@;]*@)?(\d+):\d+-(\d+):(\d+)") @@ -422,6 +428,81 @@ def blocked(position: int) -> bool: return "".join(written) +_ATTRIBUTE = r"""[.#][^\s{}]+|[\w-]+=(?:"[^"\n]*"|[^\s{}]+)""" +"""One attribute of a pandoc attribute list: a class, an id, or a key and its value.""" + +_SPAN = re.compile(rf"\[([^\[\]\n]*)\]\{{(?:{_ATTRIBUTE})(?: +(?:{_ATTRIBUTE}))*\}}") +"""A bracketed span as ``commonmark_x`` writes one, opened and closed on the one line. + +The ``]{`` is what tells one from a link's ``](`` and from an image's ``){width=...}``. +The braces must hold an attribute list, because LaTeX writes brackets before braces as +well: ``$\\sqrt[3]{x + 1}$`` is a cube root, and dropping its braces would leave +``$\\sqrt3$``, which KaTeX renders and no check reports. +""" + + +def _spans_unwrapped(markdown: str) -> str: + r"""Markdown pandoc wrote, with the attributes of its bracketed spans dropped. + + Pandoc's docx reader turns Word's underline into ``[Question 2:]{.underline}``, its + highlight into ``[oil]{.mark}`` and its small capitals into ``[Note:]{.smallcaps}``. + A field quoted out of markdown holding one of those is read back by the PDF + generator's pandoc as underline, highlight or small capitals, and written to LaTeX as + a command the generator's template does not define, so the set fails to compile. + Lambda Feedback's markdown renders none of the three, so the attribute is dropped and + the text it marked is kept. The reader emits a ``[text]{custom-style=...}`` span only + with its ``+styles`` extension, which the freeze does not enable; a document holding + one is unwrapped the same way. + + Every rewrite stays within the line it began on, so no line range moves. A span + pandoc broke over two lines is left as written, and ``in2lambda validate`` reports the + field quoting it. A span in a code block or in inline code is left as written as well: + those are characters the document shows. + + Examples: + >>> from in2lambda.source import _spans_unwrapped + >>> _spans_unwrapped("# [Hydraulic scale]{.underline}\n") + '# Hydraulic scale\n' + >>> _spans_unwrapped("**[Question 2:]{.underline}** joined by [oil]{.mark}.\n") + '**Question 2:** joined by oil.\n' + >>> _spans_unwrapped("[**[a]{.mark}**]{.underline}\n") + '**a**\n' + >>> _spans_unwrapped("[Note:]{.smallcaps} the oil is incompressible.\r\n") + 'Note: the oil is incompressible.\r\n' + >>> _spans_unwrapped("Type `[a]{.mark}` first.\n") + 'Type `[a]{.mark}` first.\n' + >>> _spans_unwrapped("Type this:\n\n [a]{.mark}\n") + 'Type this:\n\n [a]{.mark}\n' + >>> _spans_unwrapped("::: {.solution}\nThe load is $F = pA$.\n:::\n") + '::: {.solution}\nThe load is $F = pA$.\n:::\n' + >>> _spans_unwrapped('![](figure.png){width="1in"}\n') + '![](figure.png){width="1in"}\n' + >>> _spans_unwrapped("The root is $\\sqrt[3]{x + 1}$.\n") + 'The root is $\\sqrt[3]{x + 1}$.\n' + """ + if "\r\n" in markdown: + # Pandoc writes the line endings of whoever is running it, and the file on disk + # is hashed as it is written, so a Windows freeze stays a Windows file. + return _spans_unwrapped(markdown.replace("\r\n", "\n")).replace("\n", "\r\n") + + verbatim = _verbatim_lines(markdown) + + def unwrapped(match: re.Match[str]) -> str: + before = markdown[markdown.rfind("\n", 0, match.start()) + 1 : match.start()] + if markdown.count("\n", 0, match.start()) + 1 in verbatim or ( + before.count("`") % 2 + ): + return match.group() + return match.group(1) + + rewritten = _SPAN.sub(unwrapped, markdown) + if rewritten == markdown: + return markdown + # A span holding a span - `[**[a]{.mark}**]{.underline}` - unwraps from the inside, + # because the brackets of the outer one hold the brackets of the inner one. + return _spans_unwrapped(rewritten) + + def _digest(data: bytes) -> str: """How a frozen markdown is named in its draft, so that a change to it is reported. @@ -870,7 +951,9 @@ def add( are written to ``FILE.draft.json``, so that code quoting a source by line range can check that those lines still hold the text they held. A converted file is written unwrapped: a paragraph is one line, however long, and each ``$$ ... $$`` is written - on lines of its own. Lambda Feedback renders display maths written that way. + on lines of its own. Lambda Feedback renders display maths written that way. The + underline, highlight and small capitals a .docx holds are dropped and the text they + marked is kept, because Lambda Feedback renders none of the three. The files are numbered in the order they are given. A file already frozen into the draft beside them is checked against the hash it was frozen at, and is not frozen @@ -929,11 +1012,14 @@ def add( raw, markdown = _source(path) frozen_path = path else: - # Unwrapped, and with the display maths blocked out, before anything is - # hashed: both are habits of pandoc's writer rather than anything the author - # did, and both are what a field quoting these lines would have to render. + # Unwrapped, with the display maths blocked out and the bracketed spans + # dropped, before anything is hashed: all three are habits of pandoc's writer + # rather than anything the author did, and all three are what a field quoting + # these lines would have to render. markdown = _display_maths_blocked( - _pandoc(str(path), _MARKDOWN, "--wrap=none").decode("utf-8") + _spans_unwrapped( + _pandoc(str(path), _MARKDOWN, "--wrap=none").decode("utf-8") + ) ) raw = markdown.encode("utf-8") frozen_path = path.with_suffix(".md") diff --git a/tests/fixtures/sources/README.md b/tests/fixtures/sources/README.md index 677bcad..5605d78 100644 --- a/tests/fixtures/sources/README.md +++ b/tests/fixtures/sources/README.md @@ -62,3 +62,12 @@ image usually has no alt text, so this is also the ordinary case. The image the docx embeds is referenced as `media/rId9.png`, which is not extracted - nothing here reads the image, only the lines around it. + +`bracketed_spans/source.docx` holds an underlined heading, a bold and underlined `Question 2:` +beside a highlighted word, and a run in small capitals. `make_source.py` beside it wrote that +document with python-docx, which is not a dependency of in2lambda: the document is committed and +the script records what it holds. Pandoc's docx reader turns the three formats into the bracketed +spans `[Hydraulic scale]{.underline}`, `[oil]{.mark}` and `[Note:]{.smallcaps}`, which the freeze +unwraps to the text alone. The test asserting that no fixture's frozen markdown holds a `]{` +looks for that rather than for `{.`, because `fenced_div` freezes to `::: {.solution}`, which is a +div the freeze keeps. diff --git a/tests/fixtures/sources/bracketed_spans/expected.json b/tests/fixtures/sources/bracketed_spans/expected.json new file mode 100644 index 0000000..1cc3329 --- /dev/null +++ b/tests/fixtures/sources/bracketed_spans/expected.json @@ -0,0 +1,20 @@ +[ + { + "end": 1, + "id": "b1", + "start": 1, + "type": "heading" + }, + { + "end": 3, + "id": "b2", + "start": 3, + "type": "paragraph" + }, + { + "end": 5, + "id": "b3", + "start": 5, + "type": "paragraph" + } +] diff --git a/tests/fixtures/sources/bracketed_spans/make_source.py b/tests/fixtures/sources/bracketed_spans/make_source.py new file mode 100644 index 0000000..e5978fe --- /dev/null +++ b/tests/fixtures/sources/bracketed_spans/make_source.py @@ -0,0 +1,36 @@ +"""Writes the ``source.docx`` beside this file, which was run by hand. + +python-docx is not a dependency of in2lambda and is not added as one: the document is +committed, and this script records what it holds so that anyone can write it again. The +import stands inside the main block because the test suite imports every module under +``tests`` to collect its doctests. + +Word's underline, highlight and small capitals are the three formats pandoc's docx +reader turns into a bracketed span, which `in2lambda source add` unwraps. +""" + +if __name__ == "__main__": + from pathlib import Path + + from docx import Document + from docx.enum.text import WD_COLOR_INDEX + + document = Document() + + heading = document.add_heading("", level=1) + heading.add_run("Hydraulic scale").underline = True + + paragraph = document.add_paragraph() + label = paragraph.add_run("Question 2:") + label.bold = True + label.underline = True + paragraph.add_run(" A hydraulic scale has two pistons joined by ") + paragraph.add_run("oil").font.highlight_color = WD_COLOR_INDEX.YELLOW + paragraph.add_run(".") + + note = document.add_paragraph() + note.add_run("Note: ").font.small_caps = True + note.add_run("the oil is incompressible.") + + # Beside this file, so that the fixture is written wherever the script is run from. + document.save(Path(__file__).with_name("source.docx")) diff --git a/tests/fixtures/sources/bracketed_spans/source.docx b/tests/fixtures/sources/bracketed_spans/source.docx new file mode 100644 index 0000000..e44c9ac Binary files /dev/null and b/tests/fixtures/sources/bracketed_spans/source.docx differ diff --git a/tests/test_source.py b/tests/test_source.py index ab12e32..d58cd97 100644 --- a/tests/test_source.py +++ b/tests/test_source.py @@ -14,15 +14,18 @@ import pytest from click.testing import CliRunner -from conftest import SOURCES, SOURCES_DIR +from conftest import SOURCES, SOURCES_DIR, needs_compiler from in2lambda.draft import _quoted from in2lambda.main import cli -from in2lambda.validation import MathDelimiterError, math_delimiter_checker +from in2lambda.validation import MathDelimiterError, math_delimiter_checker, pdf MARKDOWN = SOURCES_DIR / "markdown" """The case the tests below happen to use; what they check holds for any of them.""" +SPAN_ATTRIBUTE = re.compile(r"\]\{[^}]*\}") +"""The attributes of a bracketed span, which no frozen markdown may hold.""" + def _frozen(directory: Path) -> Path: """The document a folder holds, whatever format it is in.""" @@ -53,6 +56,35 @@ def test_source_add_finds_the_expected_blocks(folder: Path, tmp_path: Path) -> N quoted = _quoted(draft, markdown, 1, block["start"], block["end"]) assert math_delimiter_checker(quoted) is MathDelimiterError.PASSED, quoted + # A bracketed span is underline, highlight or small capitals, which the PDF + # generator's pandoc writes as a LaTeX command its template does not define, so no + # freeze may leave one. A fenced div's `::: {.solution}` is not one of these. + assert not SPAN_ATTRIBUTE.search(markdown) + + +@needs_compiler +def test_a_frozen_docx_compiles_as_the_pdf_generator_does(tmp_path: Path) -> None: + """A .docx holding underline, highlight and small capitals is what faulted at build. + + The three sheets that reported it are private, so the fixture stands for them: every + block of it is quoted as a field would be and compiled as Lambda Feedback compiles a + set. + """ + shutil.copytree(SOURCES_DIR / "bracketed_spans", tmp_path, dirs_exist_ok=True) + + result = CliRunner().invoke(cli, ["source", "add", str(tmp_path / "source.docx")]) + + assert result.exit_code == 0, result.output + draft = json.loads((tmp_path / "source.draft.json").read_text()) + (source,) = draft["sources"] + markdown = (tmp_path / source["source"]).read_text() + fields = [ + (block["id"], _quoted(draft, markdown, 1, block["start"], block["end"])) + for block in source["blocks"] + ] + + assert pdf.problems(fields, []) == [] + def test_freezing_again_is_refused_once_the_source_has_changed( tmp_path: Path, monkeypatch