Skip to content

to_entropy() cannot read the Japanese mnemonics to_mnemonic() writes - #145

Open
fametrano wants to merge 1 commit into
trezor:masterfrom
fametrano:japanese_delimiter_read_back
Open

fametrano wants to merge 1 commit into
trezor:masterfrom
fametrano:japanese_delimiter_read_back

Conversation

@fametrano

@fametrano fametrano commented Aug 5, 2026 •

Copy link
Copy Markdown

Mnemonic("japanese") joins words with U+3000, the ideographic space
(#L73,
#L210),
as the BIP39 Japanese vectors do. to_entropy and expand split on " ",
so they see a single word. to_entropy raises; expand returns the
sentence unchanged. check accepts the sentence, because it normalizes
first, and NFKD maps U+3000 to U+0020.

The README's own steps reproduce it. On master, Python 3.12:

>>> from mnemonic import Mnemonic
>>> m = Mnemonic("japanese")
>>> words = m.generate(strength=256)   # README: "Generate word list"
>>> m.check(words)
True
>>> m.to_entropy(words)                # README: "calculate original entropy"
Traceback (most recent call last):
  ...
ValueError: Number of words must be one of the following: [12, 15, 18, 21, 24], but it is not (1).

This block is a doctest. It passes on master. With this patch it fails,
because to_entropy returns the entropy. On vectors.json, check
accepts all 288 mnemonics. to_entropy fails on all 24 Japanese ones, and
on none of the other 264.

#110 fixed this by splitting on self.delimiter. It was merged as
71cf5203 and reverted as df3e1500 for failing CI. That approach cannot
work in check, which normalizes before it splits: after NFKD there is
no U+3000 left to split on, so every Japanese mnemonic fails validation.
So this PR splits on any run of whitespace. detect_language already
does this.

The changes to to_entropy, check and expand fix the Japanese bug.
The change to to_seed is separate. Today to_seed only applies NFKD.
So input that check rejects, such as a leading space, a trailing newline
or a tab, silently gives a different seed from the same words separated
by single spaces. If you do not want that change, drop the to_seed line
and the last assertion of test_whitespace_runs. The fix for the Japanese
bug still holds.

Compatibility: all 288 vectors are unchanged. test_vectors asserts the
mnemonic, seed and xprv of each, and passes. No input that check accepts
today gives a different seed. Only input that it rejects changes: it now
gives the canonical seed. The passphrase is not touched, because its
whitespace is part of the secret.

CI, since #110 was reverted for failing it: python tests/test_mnemonic.py
is OK on Python 3.8 through 3.14 and on PyPy 3.11. black --check,
isort --check-only, flake8 and pyright report nothing on
src tests tools.

Written with machine assistance and reviewed against master b57a5ad
before filing. Found while checking btclib's BIP39 reading against this
implementation: btclib-org/btclib#258.

@fametrano

Copy link
Copy Markdown
Author

@prusnak could you approve the workflow run here? This is my first pull request from this fork, so CI has never started, and the local run in the description is the only evidence in the thread.

The branch is a single commit now. Since filing I added one test for expand, the fourth line the patch changes and the one that had no coverage.

@fametrano

Copy link
Copy Markdown
Author

@prusnak the workflow runs expired before approval, so CI still hasn't run. Could you re-run them, or should I push a new commit?

to_mnemonic joins a Japanese mnemonic with U+3000, while to_entropy,
check and expand split on " ". So the library cannot read the sentences
it writes: to_entropy raises on all 24 Japanese vectors of vectors.json,
and expand returns the sentence unchanged.

check splits after NFKD, which maps U+3000 to U+0020. So it cannot
split on self.delimiter: that was trezor#110, reverted in df3e150 for failing
CI. Split on any run of whitespace instead, as detect_language already
does.

to_seed follows the same rule. Before, input that check rejects gave a
different seed from the same words separated by single spaces.

The tests fail without the fix: to_entropy raises on every Japanese
vector, check rejects a sentence separated by anything but one space,
and expand returns a tab- or U+3000-separated sentence unexpanded.
The existing round-trip test hides the first failure, because it splits
the sentence itself before passing it in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@fametrano fametrano changed the title to_entropy() cannot read the japanese mnemonics to_mnemonic() writes to_entropy() cannot read the Japanese mnemonics to_mnemonic() writes Sep 29, 2026
@fametrano
fametrano force-pushed the japanese_delimiter_read_back branch from bd50ae2 to b5e1375 Compare September 29, 2026 21:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant