Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
- A first argument that is neither a command nor a file is refused, and the message names the `in2lambda convert` command to run. Earlier versions printed the usage message and exited 0.
- beartype is now `^0.22`. At 0.20.0 and below, the beartype import hook leaves `cli` a plain function instead of a group, so the command line either fails to import or runs `convert` whatever the arguments are. 0.20.1 is the first version that works.
- `in2lambda source add FILE` freezes a document: `in2lambda source add` converts .docx and .tex to markdown beside the file, and writes `FILE.draft.json` beside the file, holding the markdown's hash and every block in the markdown with the lines it spans, so that another tool can quote the source by line range. The draft is named after the source, so a folder holding a term's sheets holds one draft per sheet. `in2lambda source show` prints that markdown numbered with the block ids. Freezing a file that has changed since is refused unless `--start-over` discards the draft. Showing a file that has changed is refused as well, because its block ids would name lines they were not taken from. `in2lambda source add` needs pandoc and the `convert` extra, as `in2lambda convert` does. `in2lambda source show` reads the markdown the draft was frozen from, and needs neither.
- `in2lambda source add` writes the markdown of a converted .docx or .tex unwrapped, so that a paragraph is one line however long the paragraph is, and an inline `$ ... $` is never broken across two lines. `in2lambda source add` also moves every `$$ ... $$` that pandoc wrote on one line onto lines of its own, in a paragraph and in a list item, indented to the item's width where the maths is written in a list item. A `$$` that opens or closes in a table cell, in a block quote or in a code block is left unchanged, as is the maths that an unpaired `$$` elsewhere in the document, such as one in inline code, pairs with. `in2lambda validate` reports the maths left unchanged. pandoc's writer produces both forms, and Lambda Feedback renders neither, so `in2lambda validate` reported them against every converted sheet. The markdown beside a document frozen before this release differs, and so do its line ranges. `in2lambda source add --start-over` freezes the document again.
- `in2lambda source add` writes the markdown of a converted .docx or .tex unwrapped, so that a paragraph is one line however long the paragraph is. Where the author broke a line inside an inline `$ ... $`, pandoc writes that newline back, and `in2lambda source add` joins the lines of the maths with a space. `in2lambda source add` also moves every `$$ ... $$` onto lines of its own, in a paragraph and in a list item, indented to the item's width where the maths is written in a list item, whether pandoc wrote the maths on one line or opened it on one line and closed it on the next. A `$$` that opens or closes in a table cell, in a block quote or in a code block is left unchanged, as is the maths that an unpaired `$$` elsewhere in the document, such as one in inline code, pairs with. An inline `$ ... $` that opens or closes in a table cell, in a block quote or in a code block, and the text between two unpaired `$`, are left unchanged for the same reason. `in2lambda validate` reports the maths left unchanged. pandoc's writer produces both forms, and Lambda Feedback renders neither, so `in2lambda validate` reported them against every converted sheet. The markdown beside a document frozen before this release differs, and so do its line ranges. `in2lambda source add --start-over` freezes the document again.
- A draft now holds a `log` of every command that changed it, and a `fields` map of the fields those commands wrote. Each field records the layer that wrote it (1 a spec, 2 a predicate, 3 a line range, 4 a literal), the source ranges it was copied from, whether it has been edited, and by whom. `in2lambda draft mark ignore BLOCK` is the first such command. `in2lambda draft replay` rebuilds the draft from the frozen markdown and the log, and refuses unless the draft it builds matches the draft on disk byte for byte. A draft written before this release holds no `log` and is refused as a draft in2lambda did not write; `in2lambda source add --start-over` freezes the document again.
- `in2lambda draft question add`, `in2lambda draft part add QUESTION` and `in2lambda draft question solution QUESTION` fill a draft in. Each takes `--text` to copy the wording out of the frozen source, as a block id such as `b3` or as lines such as `s10:14`. Each takes `--literal TEXT` for wording the source does not hold in a form the field can take, which records the field as edited and written by layer 4 in place of layer 3. in2lambda works the question and part numbers out from the fields already written, so a replay arrives at the same ids. `in2lambda draft split block BLOCK AT` cuts a block the parser made one of two things into `b3a` and `b3b`, so that each half can be quoted on its own. A command writing a field that is already written is refused, naming the field. A command quoting lines another field was taken from is refused, naming both fields.
- `in2lambda draft part solution PART --text RANGE` gives one part of a question its worked solution, for a sheet that writes a solution under each part rather than one solution answering the whole question. PART is written `q1.p2`. The command writes `q1.p2.solution` from the lines named, or from `--literal TEXT`, as `in2lambda draft question solution` writes a question's solution. A PART naming a question, or naming a part the draft has not written, is refused, naming the part. A sheet laid out this way can now be answered part by part by commands, which only a spec's `PartPartSolSol` and `PartSolPartSol` layouts could do before, so the warning `in2lambda validate` prints about a part nothing answers can be acted on.
Expand Down
173 changes: 149 additions & 24 deletions in2lambda/source/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -270,6 +270,9 @@ def _pandoc(file: str, to: str, *options: str) -> bytes:
_MARKER = re.compile(r" *(?:[-+*]|\(?(?:\d+|[ivxlcdm]+|[IVXLCDM]+|[A-Za-z])[.)]) {1,4}")
"""A list item's marker on its first line, as `commonmark_x` reads one."""

_DOUBLE_DOLLAR = re.compile(r"(?<!\\)\$\$")
"""One delimiter of a display maths, wherever it stands on a line."""


def _verbatim_lines(markdown: str) -> set[int]:
r"""The lines of some markdown whose ``$$`` is code rather than maths.
Expand All @@ -283,6 +286,13 @@ def _verbatim_lines(markdown: str) -> set[int]:
:func:`dedented` takes off again, so a line this leaves alone is a line the field
quoting it reads as code too.

A line inside a ``$$ ... $$`` opened on an earlier line is maths however far it is
indented: pandoc writes the author's own indent inside the maths on top of the
item's, which puts the closing delimiter of maths in a list item past the column
code starts at. Such a line neither counts as code nor closes the item it stands
in. Maths does not cross a blank line, so a blank line ends the span, and an
unpaired ``$$`` - one in inline code, say - leaves the lines after it as they were.

Examples:
>>> from in2lambda.source import _verbatim_lines
>>> sorted(_verbatim_lines("Text\n\n $$x = y$$\n"))
Expand All @@ -293,9 +303,14 @@ def _verbatim_lines(markdown: str) -> set[int]:
[3]
>>> sorted(_verbatim_lines("``` python\n$$x = y$$\n```\n"))
[1, 2, 3]
>>> sorted(_verbatim_lines("1. Item $$x =\n y$$ more\n"))
[]
>>> sorted(_verbatim_lines("1. Item\n\n $$x =\n y$$\n"))
[3, 4]
"""
verbatim = set()
fence = ""
maths = False # Whether a `$$` opened on an earlier line is still open.
items: list[int] = [] # The content column of each list item open at this line.
for number, line in enumerate(markdown.split("\n"), start=1):
stripped = line.lstrip(" ")
Expand All @@ -307,7 +322,11 @@ def _verbatim_lines(markdown: str) -> set[int]:
elif not stripped:
# Commonmark closes an item at the next non-blank line indented less than
# its content column, not at the blank line before that one.
maths = False
continue
elif maths:
if len(_DOUBLE_DOLLAR.findall(line)) % 2 == 1:
maths = False
else:
while items and indent < items[-1]:
items.pop()
Expand All @@ -317,11 +336,103 @@ def _verbatim_lines(markdown: str) -> set[int]:
verbatim.add(number)
elif indent >= base + 4:
verbatim.add(number)
elif marker := _MARKER.match(line):
items.append(marker.end())
else:
if marker := _MARKER.match(line):
items.append(marker.end())
# Only a line that is not code opens a span: a `$$` in a code block is
# characters the document shows rather than a delimiter.
maths = len(_DOUBLE_DOLLAR.findall(line)) % 2 == 1
return verbatim


def _blocked(markdown: str, verbatim: set[int], position: int) -> bool:
"""Whether the delimiter at this offset stands in a table row, a quote or code.

The rewrites below move a maths onto lines of their own or join the lines it stands
on, and neither can carry a pipe table's cell, a block quote's ``> `` or the meaning
of a code block's characters across the lines it changes.
"""
opening = markdown[markdown.rfind("\n", 0, position) + 1 : position]
return opening.lstrip()[:1] in ("|", ">") or (
markdown.count("\n", 0, position) + 1 in verbatim
)


_INLINE_MATHS = re.compile(
r"(?<![\\$])\$(?!\$)(?!\s)((?:[^$\\]|\\.)+?)(?<!\s)\$(?!\$)", re.DOTALL
)
"""Inline maths: a single ``$``, content holding no unescaped ``$``, a single ``$``.

The content neither starts nor ends with whitespace, which is the rule pandoc's own
``tex_math_dollars`` reader applies. Without it, the ``$`` pandoc writes unescaped
inside a URL pairs with the opening ``$`` of the next maths in the document.
"""


def _inline_maths_joined(markdown: str) -> str:
r"""Markdown pandoc wrote, with each inline ``$ ... $`` on one line.

An author who broke a line inside a ``$ ... $`` in the document has that newline
written back by ``commonmark_x``, and the delimiter checks refuse a newline inside
inline maths. The newline carries nothing the maths renders, so the lines of the
span are joined with a space. Pandoc escapes a dollar the document shows as ``\$``,
which is what makes an unescaped single ``$`` in its output a delimiter.

A match holding a backtick, or running across a blank line, is left as written, as
is one whose opening or closing ``$`` stands on a code block's line: such a match is
an unpaired ``$`` - one in inline code or in a shell prompt - closed by the ``$`` of
a later maths, and joining the two would run the lines between them together. A
match that opens or closes on a pipe table's row or on a block quote's line is left
as written too: the join keeps the first line as it stands and strips the rest, so
it would carry the quote's ``> `` or the next row's ``|`` into the maths.
``in2lambda validate`` reports the maths left in any of these. A ``$$`` opens no
match here, so display maths is left to :func:`_display_maths_blocked`.

An unescaped ``$`` pandoc writes inside a URL joins nothing either: the text from
that ``$`` to the opening ``$`` of the next maths ends with the space before that
delimiter, and :data:`_INLINE_MATHS` matches no span whose content ends in
whitespace.

Examples:
>>> from in2lambda.source import _inline_maths_joined
>>> _inline_maths_joined("A speed of $v =\n576$ here.\n")
'A speed of $v = 576$ here.\n'
>>> _inline_maths_joined("1. The energy is $U =\n 5a$ now.\n")
'1. The energy is $U = 5a$ now.\n'
>>> _inline_maths_joined("The load is $$F =\npA$$ here.\n")
'The load is $$F =\npA$$ here.\n'
>>> _inline_maths_joined("Type `$` and\nmore `$` here.\n")
'Type `$` and\nmore `$` here.\n'
>>> _inline_maths_joined("Costs $5 today.\n\nAnd more($6) here.\n")
'Costs $5 today.\n\nAnd more($6) here.\n'
>>> _inline_maths_joined("``` sh\ncost=$a and\nsum=$b\n```\n")
'``` sh\ncost=$a and\nsum=$b\n```\n'
>>> _inline_maths_joined("> The energy is $U =\n> 5a$ here.\n")
'> The energy is $U =\n> 5a$ here.\n'
>>> _inline_maths_joined("A fee at <http://x/$1>\\\nand $E = mc^2$ here.\n")
'A fee at <http://x/$1>\\\nand $E = mc^2$ here.\n'
"""
verbatim = _verbatim_lines(markdown)
written: list[str] = []
end = 0
for match in _INLINE_MATHS.finditer(markdown):
if "\n" not in match.group():
continue
if "`" in match.group(1) or any(
not line.strip() for line in match.group().split("\n")
):
continue
if _blocked(markdown, verbatim, match.start()) or _blocked(
markdown, verbatim, match.end()
):
continue
body = " ".join(line.strip() for line in match.group(1).split("\n"))
written.append(f"{markdown[end : match.start()]}${body}$")
end = match.end()
written.append(markdown[end:])
return "".join(written)


def _display_maths_blocked(markdown: str) -> str:
r"""Markdown pandoc wrote, with its display maths moved onto lines of its own.

Expand All @@ -332,7 +443,10 @@ def _display_maths_blocked(markdown: str) -> str:

The inserted lines take the indent of the line the maths began on - a list item's
marker width included, so maths in an item stays in the item - and whatever stood
either side of it on that line becomes a paragraph of its own.
either side of it on that line becomes a paragraph of its own. A maths the author
broke a line inside, which pandoc writes as an opening ``$$`` on one line and a
closing ``$$`` on a later one, is rewritten the same way, each of its lines becoming
a line of the block.

A ``$$`` that opens or closes on a pipe table's row, on a block quote's line or on a
code block's line is left as pandoc wrote it: a table cell cannot hold a block, an
Expand All @@ -349,8 +463,8 @@ def _display_maths_blocked(markdown: str) -> str:
'The load is\n\n$$\nF = pA\n$$\n\nhere.\n'
>>> _display_maths_blocked("1. Find $$F = pA$$\n")
'1. Find\n\n $$\n F = pA\n $$\n'
>>> _display_maths_blocked("A load $$F = pA$$\r\n")
'A load\r\n\r\n$$\r\nF = pA\r\n$$\r\n'
>>> _display_maths_blocked("1. Given $$U = a\n b$$ where r is.\n")
'1. Given\n\n $$\n U = a\n b\n $$\n\n where r is.\n'
>>> _display_maths_blocked("> The load is $$F = pA$$ here.\n")
'> The load is $$F = pA$$ here.\n'
>>> _display_maths_blocked("Type this:\n\n $$x = y$$\n")
Expand All @@ -364,21 +478,7 @@ def _display_maths_blocked(markdown: str) -> str:
>>> _display_maths_blocked("The load is $$F = pA\n> and $$ here.\n")
'The load is $$F = pA\n> and $$ here.\n'
"""
if "\r\n" in markdown:
# Pandoc writes the line endings of whoever is running it, and the file on disk
# is hashed as it is written, so a Windows freeze stays a Windows file.
blocked = _display_maths_blocked(markdown.replace("\r\n", "\n"))
return blocked.replace("\n", "\r\n")

verbatim = _verbatim_lines(markdown)

def blocked(position: int) -> bool:
"""Whether the `$$` at this offset stands in a table row, a quote or code."""
opening = markdown[markdown.rfind("\n", 0, position) + 1 : position]
return opening.lstrip()[:1] in ("|", ">") or (
markdown.count("\n", 0, position) + 1 in verbatim
)

written: list[str] = []
end = 0
for match in _DISPLAY_MATHS.finditer(markdown):
Expand All @@ -390,7 +490,9 @@ def blocked(position: int) -> bool:
# the opening `$$` of a later maths. Rewriting it would make a maths block
# of the words standing between the two.
continue
if blocked(match.start()) or blocked(match.end()):
if _blocked(markdown, verbatim, match.start()) or _blocked(
markdown, verbatim, match.end()
):
# A pipe table's cell cannot hold a block; an inserted line carries the
# indent of the line the maths began on but not a block quote's `> `, so the
# rewrite would put the maths and the words after it outside the quote; and
Expand Down Expand Up @@ -422,6 +524,27 @@ def blocked(position: int) -> bool:
return "".join(written)


def _maths_rewritten(markdown: str) -> str:
r"""Markdown pandoc wrote, with its maths written as the delimiter checks want it.

:func:`_inline_maths_joined` runs first: joining the lines of an inline maths moves
every line below it, and :func:`_display_maths_blocked` reads line numbers.

Examples:
>>> from in2lambda.source import _maths_rewritten
>>> _maths_rewritten("A load $$F = pA$$\r\n")
'A load\r\n\r\n$$\r\nF = pA\r\n$$\r\n'
>>> _maths_rewritten("A speed of $v =\n576$ here.\n")
'A speed of $v = 576$ here.\n'
"""
if "\r\n" in markdown:
# Pandoc writes the line endings of whoever is running it, and the file on disk
# is hashed as it is written, so a Windows freeze stays a Windows file.
rewritten = _maths_rewritten(markdown.replace("\r\n", "\n"))
return rewritten.replace("\n", "\r\n")
return _display_maths_blocked(_inline_maths_joined(markdown))


def _digest(data: bytes) -> str:
"""How a frozen markdown is named in its draft, so that a change to it is reported.

Expand Down Expand Up @@ -929,10 +1052,12 @@ def add(
raw, markdown = _source(path)
frozen_path = path
else:
# Unwrapped, and with the display maths blocked out, before anything is
# hashed: both are habits of pandoc's writer rather than anything the author
# did, and both are what a field quoting these lines would have to render.
markdown = _display_maths_blocked(
# Unwrapped, with the display maths blocked out and the inline maths joined
# onto one line, before anything is hashed: the wrapping and the one-line
# `$$ ... $$` are habits of pandoc's writer rather than anything the author
# did, the newline inside an inline `$ ... $` carries nothing the maths
# renders, and all three are what a field quoting these lines would render.
markdown = _maths_rewritten(
_pandoc(str(path), _MARKDOWN, "--wrap=none").decode("utf-8")
)
raw = markdown.encode("utf-8")
Expand Down
Loading
Loading