Skip to content

Record which rule produced each braille cell - #207

Open
owjs3901 wants to merge 150 commits into
mainfrom
owjs3901/rule-trace
Open

owjs3901 wants to merge 150 commits into
mainfrom
owjs3901/rule-trace

Conversation

@owjs3901

@owjs3901 owjs3901 commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

점자 한 칸마다 그 칸을 적은 규정 조항을 기록하는 추적(encode_with_trace)으로 시작해서, 과학 점자 규정 전체와 문맥 인자까지 담은 PR입니다.

  • 규정 fixture 5740/5914. 실패 174건은 모두 이번에 넣거나 고친 fixture로, 엔진이 아직 PDF대로 적지 못하는 곳입니다(아래 '누락 예제와 정정'). 나머지는 모두 통과합니다. 인쇄의 글자체를 유니코드로 적을 수 없는 5건은 limitation으로 건너뜁니다.
  • 말뭉치 456,626/467,121 (97.75%). 문장 단위로 비교해 잃은 문장은 3개이며, 모두 말뭉치가 같은 모양을 두 가지로 적은 곳입니다(아래 '문장 부호 띄어쓰기' 항목).
  • Linux 라인 커버리지 100% (d0f75dc 기준. 위 174건 때문에 지금 CI의 테스트 단계는 실패합니다)

검수 부탁드릴 fixture — 누락 예제와 정정

두 규정 PDF(docs/2024 개정 한국 점자 규정.pdf, docs/Rules-of-Unified-English-Braille-2024.pdf)에 인쇄된 점자를 모두 뽑아 fixture와 맞췄습니다. 쪽수는 PDF 뷰어 기준입니다. 커밋: 31554d3, 659bffd, 008215c, 0070e44, 8bead1d, 8bcf9ee, dc06cf1, c9c7b41, c1f6dbd, eb33c50, 7bed3a4, ee5985e, d68c221.

누락 예제 추가

  • 영어 표·예문(31554d3, dc06cf1)
    • 각 절 첫머리 기호·약자 표에서 단독 fixture가 없던 칸: 단어 약자, 글자 약자, 단축어, 그리스 문자, 옛 영어 글자, 화살표, 화폐, 문장 부호 등
    • 부록 1(단축어 목록)과 부록 2(낱말 목록) 중 다른 fixture에 없던 낱말
    • §2.4.7, 2.6.2, 2.6.3, 4.4.1, 9.1.3, 15.2의 빠진 예문
    • 새 그룹: english/appendix_1, english/appendix_2, english/rule_2_4_7, english/rule_4_4_1
    • 부록 1(d68c221): 점자로 인쇄된 항목을 모두 PDF 순서대로 담았습니다. 10.9.5 예외 3, 단축어 표제어 75, 목록 구성 규칙 2~6의 예 26건입니다. 규칙 3의 [not] 형태는 note에만 적었습니다. 표제어 아래의 긴 낱말 328개는 PDF에 점자가 없어 넣지 않았습니다.
  • 영어 기호(eb33c50, ee5985e)
    • 부록 3(기호 목록, english/appendix_3) 61건: 유니코드 코드가 있고 단독 fixture가 없던 기호. 대부분 기술 자료 기호(∞ ≤ ≥ ∈ ∀ ∂ √ ± ∴ ∵ 등)입니다. 같은 글자에 점자가 여럿 인쇄된 경우 모두 받습니다(“ ⠦/⠘⠦, ” ⠴/⠘⠴, ’ ⠄/⠠⠴).
    • §6.1 숫자 10건(english/rule_6_1_1), §15.3 성조 글자 ˦ ˧ ˨(english/rule_15_3_1), §3.27 [open tn]·[close tn] 단독
    • 부록 2의 Al(인쇄 그대로 ⠠⠁⠇). §10.9.7은 홀로 선 Al을 ⠰⠠⠁⠇로 적어 서로 어긋납니다.
    • §3.22 마지막 아이콘 예(물음표를 동그라미로 감싼 아이콘): 입력은 ? + 둘러싸는 동그라미 U+20DD
  • 한글(31554d3, c1f6dbd)
    • 제52항 //, 제59항 ;, 제60항 *·※ 단독, 제72항 예제 첫 줄
    • 제53항 줄임표 네 형태(…… … ⠠⠠⠠, ...... ... ⠲⠲⠲)
    • 제61항 단독 ’(⠄). 제49항의 단독 닫는 작은따옴표(⠴⠄)와 입력이 같아 둘 중 하나는 반드시 실패합니다.
  • 과학(31554d3, c9c7b41)
    • 제1항 첫 줄, 제4항 2 NH₃ 구조식·전자 점식, 제18항 5·6, 제21항 전극 ∣과 염다리 ∥(과학 문맥), 제2항 이온 기호 ⁺ ⁻ 단독
    • 제13항 4·5·7 환핵 그림 6건(science_spatial): 케쿨레 벤젠, 나프탈렌, H가 붙은 벤젠, 메타자일렌, 사이클로펜테인, 사이클로프로페인
    • [부록](science/appendix, 과학 문맥) 53건: 구조식 6(기호 표기 + 점역자 주), 일기 현상 5, 운량 11, 전선 4, 그 밖의 일기 기호 4, 전기·전자 회로 소자 23

입력 형식을 새로 정한 것:

  • 환핵 그림: 묵자 그림을 줄마다 옮겨 사선은 / \, 세로선은 |로 적었습니다. 점자가 칸을 두 번 적는 이중 결합은 // \\ ||로, 가로선은 ─(점자의 ⠒ 수만큼)로 적었습니다.
    • PDF는 빈칸을 반 칸 폭으로 조판해 사선 줄이 반 칸씩 어긋나 보입니다. 사선이 줄마다 한 칸씩 옮겨 가도록 칸을 정했고, 한 줄 안의 빈칸 수는 PDF 그대로입니다.
    • [부록]의 구조식은 같은 그림을 과학 문맥(기호 표기)으로 읽은 것이라, 제13항과 같은 입력을 씁니다.
  • 그림뿐인 기호: [그림: 명칭]으로 적었습니다(영어 fixture의 [open tn]과 같은 자리표시). 해당하는 것은 눈, 운량 0~10, 전선, 태풍, 열대성 저기압, 회로 소자입니다. PDF가 글자로 인쇄한 기호(≡ ∇˙ ☈ ● H L)는 그 글자를 씁니다.

판단한 것:

  • 홀로 선 → ← ↑ ↵ ? " 6건은 두 형태를 alternatives로 받습니다.
    • 기호 표 형태: ⠰ 없음(49·103쪽)
    • 글 속에 홀로 선 형태: ⠰를 앞세움(51·85·106·109·114쪽 예)
  • ※는 제60항 본문의 ⠐⠔과 [다만 1]의 ⠸⠔을 둘 다 받습니다.
  • 넣지 않은 것:
    • §1.3.2~1.3.5: UEB가 아닌 영어 점자의 기호입니다.
    • 점역자가 정하는 기호: §2.6.5·3.26 점역자 정의 기호, 한글 제72항 점역자 정의 글머리 기호. 묵자가 정해져 있지 않습니다.
    • 점자 전용 표지: 대문자·글자체·수표·줄 끝 연속 표지 등. 묵자 글자가 없습니다.
    • 묵자가 없는 칸: §4.2 대문자 앞 변형 부호(글자 없이는 입력이 없음), 한글 제73항 표의 빈칸 ⠿⠿
    • §15.3의 나머지 성조: 같은 글자(➚ ➘)에 점자가 둘이거나, ↑ ↓ ↗ ↘처럼 화살표와 입력이 같습니다.
    • 음악 점자(한국·서양): 오선 악보를 글자로 적을 입력 형식이 없고, 엔진도 지원하지 않습니다.

잘못된 fixture 정정

  • 영어 정답이 엔진 출력이던 11건(659bffd): 29f1e19에서 들어왔고, 그중 9건에는 note "aligned to current encoder output"가 있었습니다. 정답을 PDF 점자로 바꿨습니다.
    • 약자: afternoon ⠁⠋⠝·from ⠋·herself ⠓⠻⠋(2.6.3), ar ⠜(4.2.2), ea ⠂(8.5.6), and ⠯(9.1.3), ed ⠫·en ⠢(9.3.2)
    • 배치: 줄 표시 ⠸로 이은 시(2.6.3), UEB 대문자 단어표와 밑줄 끝 표지 뒤의 마침표(9.1.3), 마침표 앞 밑줄 기호 표지 없음(9.8.1)
    • 인쇄 모양을 잃었던 입력: 2.6.3 이탤릭, 4.2.2 굵은 Ž(굵은 Z + U+030C), 8.4.2 이탤릭 X, 9.5.1 점선 밑줄 U+0323
  • 영어 29건(7bed3a4): 모든 영어 fixture의 점자가 이제 PDF에 있습니다. 27건은 점자가 PDF 어디에도 없었고, 4건은 칸은 같지만 배치가 달랐습니다.
    • 선 모드(16.2.2, 16.3.1, 16.4.2, 16.4.3, 16.5.1): 정답을 PDF 점자 격자 그대로 칸마다 맞췄습니다.
      • 표 두 개가 한 줄로 이어져 있었습니다.
      • 상자가 한 칸 좁았고, 이음 칸에 세로선 칸이 하나 더 있었습니다.
      • 점 위치 표지로 싼 도식의 끝이 ⠐⠐⠿⠰⠄가 아니라 ⠰⠄였습니다.
      • 입력도 3건 바꿨습니다: 교차 도식은 인쇄의 7줄로, 조직도는 빠진 세로선을 넣었고, 삼목(noughts and crosses)은 인쇄대로 소문자 x o로 적었습니다.
    • 글자체: 입력에는 점자가 적는 글자체만 넣었습니다.
      • 인쇄는 굵지만 점자는 글자체를 적지 않은 제목·외국어 문장(8.5.6, 13.6.4)과, 인쇄 근거 없이 이탤릭이던 번역문·제목(8.5.7, 13.6.4, 16.5.1)은 보통 글씨로 바꿨습니다.
      • 점자가 이탤릭을 적는 letter d와 FAHRENHEIT(5.7.1, 8.5.3)는 이탤릭을 넣었습니다.
    • limitation 5건: 인쇄의 글자체를 유니코드로 적을 수 없어 입력만으로는 PDF 점자를 낼 수 없는 행입니다. 정답은 PDF 점자로 두었습니다.
      • 굵은 ! ? $ €(9.2.1), 이탤릭 숫자(8.5.3), 필기체 숫자(9.3.2)
      • 하네스는 이 행을 건너뛰되, 엔진이 이 행을 통과하면 오류로 알립니다.
    • 11.5.1: 근호 앞 1급 기호를 인쇄대로 ⠰ 하나로 적었습니다.
    • 7.2: 제가 넣었던 단독 긴 줄표의 입력을 U+2014에서 U+2015로 고쳤습니다. U+2014는 §7.2.1 예와 부록 3대로 줄표 ⠠⠤입니다.
  • 줄바꿈(8bead1d, 8bcf9ee): 줄이 바뀌는 fixture는 배치를 하나로 고르지 않고 alternatives로 아래 형태를 모두 받습니다. 점자 칸은 모든 형태가 같습니다.
    • PDF에 인쇄된 점자 줄 그대로
    • 인쇄의 줄바꿈만 따른 형태(서명·제목이 따로 선 줄) — PDF 줄과 다를 때만
    • 한 줄로 이은 형태
    • 영어: 11.6.1, 15.1.2 2건, 15.2.1 3건, 2.6.3 시, 8.5.6, 8.5.7, 8.5.3, 9.1.3 CHAPTER 6, 9.4.4(PDF 형태에만 첫 줄 끝 ⠐⠐)
    • 한국 점자: 한글 제42항 2건·제43항, 수학 제66항 2건(LaTeX 포함 3행), 과학 제19항·제24항 2건. 줄 끝 연결표 ⠠와 수학 제66항의 연산 기호 뒤 줄바꿈이 이제 fixture에 있습니다.
    • 낱말 나누기 예(10.13, 8.4.3, 8.4.4)는 한 형태로 둡니다. 나누기는 줄 끝에서만 성립하고, 한 줄로 이으면 두 낱말로 읽힙니다.
  • 수학 4건(008215c): math_4 -5<x<-2의 둘째 빼기표, math_66 f(x+a)f(x-a)의 둘째 f, 그리고 각각의 LaTeX 행.
  • 판단한 것:
    • 8.5.6 DIGEST: August의 한 줄 형태 빈칸은 internal대로 한 칸입니다. PDF가 이 자리에서 줄을 바꿔, 인쇄의 세 칸 공백이 점자로 몇 칸인지 드러나지 않습니다.
    • rule_9_7_3의 5~7행은 rule_9_8_1과 겹친 §9.8.1 예문이라 지웠습니다.
  • bun 무결성 테스트가 이제 english 폴더도 검사합니다(0070e44).

엔진이 아직 PDF대로 적지 못하는 174건

CI의 테스트 단계는 이 174건에서 실패합니다.

  • 과학 63건: [부록] 52, 제13항 환핵 그림 6, 염다리 ∥, 제1항 두 문단, 제18항 5·6의 3건(화살표 위아래 글자)
  • 영어 107건
    • 부록 3 56, 선 모드 13, 부록 2 3(Al, hoity-toity, sh), §15.3 성조 3
    • 옛 글자 대문자 Ȝ·Þ·Ð·Ƿ 4, ŋ·Ŋ·Ə·meningococcus 4, ¡·¿, 단독 ′·″·∶·∷·^
    • 2.4.7 2, 2.6.3 2, 9.1.3 2, 9.3.2 2, 3.22 아이콘, 4.2.2, 8.5.6, 9.5.1, 9.8.1, 11.5.1, 13.6.4, 15.2 운율 표시, 7.2 단독 줄표 —(U+2014)
  • 한글 2건: 제60항 단독 *(끝에 빈칸이 붙음), 제61항 단독 ’
  • 수학 2건: math_4 둘째 빼기표

검수 부탁드릴 fixture — test_cases/science/ 33개 파일, 212건

모두 docs/2024 개정 한국 점자 규정.pdf 과학 점자에 인쇄된 예제입니다. world·jeomsarang은 비워 두었습니다.

  • 문맥을 밝힌 188건

    • "context": "korean" 13건: 제30항 단독 단위(μm, in, yard, min, hPa, kgf/m² 등)
    • "context": "science" 159건: 한글이 없는 과학 fixture 83건 전부(한글 없는 글은 문맥이 없으면 영어로 읽기 때문, 아래 '문맥 없는 글' 참조)와, 한글이 있어도 묵자 모양만으로는 과학 기호인지 알 수 없는 글 15건(pOH, 유전자, 치식, 고리 화합물 등), 그리고 이번에 넣은 누락 예제·[부록] 61건
    • "context": "science_spatial" 16건: 공간 표기 형식(아래, 제13항 환핵 그림 6건 포함)
    • 과학 fixture는 모두 과학 문맥에서도 같은 답을 내도록 테스트로 묶었습니다.
  • 여러 답을 허용한 alternatives 5건

  • 여러 줄 입력 17건 (2D 도식): 묵자 배치를 줄마다 옮긴 입력입니다.

    • 구조식: 가로 결합 - = ≡, 세로 결합 | ‖ ⦀
    • 전자: : ‥는 2개, ⋮는 3개, ∷는 4개. 제16항 4의 위쪽 전자 넷은 PDF에 점 네 개로만 인쇄되어 있어 ∷로 적었습니다.
    • 가계도(제25항)는 세대마다 한 줄씩 \n으로 적습니다.
  • 공간 표기 형식 9건 (제9항 2, 제11항 4건, 제15항 2, 제17항 2건, 제18항 6): 같은 묵자를 기호 표기로도 적을 수 있어(제9·15항) science_spatial 문맥으로 고릅니다.

    • PDF는 빈칸을 반 칸 폭으로 조판했습니다. 기호가 보이는 위치를 칸으로 옮겼고, 도식 전체의 왼쪽 여백은 뺐습니다. 모든 예에서 세로 결합선·전자와 위아래 원소는 중심 원소의 첫 칸(대문자 기호표 칸)에 맞춰져 있습니다.
    • 각 fixture의 note에 이 읽기를 적었습니다. HCOOH는 PDF가 도식 전체를 한 칸 들여 썼는데, 이 여백을 뺐습니다.
  • science_31 만유인력 상수: alternatives로 두 형태를 받습니다.

    • PDF 네 줄 그대로
    • 줄을 바꾸지 않은 한 줄 형태(줄을 바꿀 때만 적는 수학 제66항 연결표 ⠠를 뺌)
    • sec 앞의 로마자표는 규정 문구가 없어 이 예제를 따랐습니다(국립국어원 질의 초안 15번).
  • 고리 화합물(제12~14항) 12건: 환핵 도형 문자를 묵자 배치대로 놓아 입력합니다.

    • 세로 방향 육각 환핵 ⬡(U+2B21), 가로 방향 ⎔(U+2394), 오각 환핵 ⬠(U+2B20), 사각 환핵 □(U+25A1), 삼각 환핵 △(U+25B3).
    • □·△는 기본 문맥에서 제72항 글머리 기호이므로, science 문맥에서 글 전체가 그 글자일 때만 환핵으로 읽습니다. 제12항 3·4는 그림 없이 점형만 정의했습니다(오각 환핵과 같음).
    • 세로 결합선을 나눈 두 핵은 같은 줄에 두 칸, 사선 결합선을 나눈 두 핵은 한 칸·한 줄 떨어뜨립니다(나프탈렌 ⬡ ⬡).
    • ⎔는 기본 문맥에서 수학 제40항 육각형이므로 science 문맥에서만 환핵으로 읽습니다.
    • 제14항 3의 페난트렌은 PDF가 같은 그림을 중심-위-아래, 중심-아래-위 두 순서로 적었습니다. 순서가 그림에 드러나지 않아 alternatives로 둘 다 받습니다.
  • 제13항 6 사각 환핵 공간 표기 1건(science_spatial): 원소가 달린 사각형을 사슬 화합물 결합선으로 적습니다. 세로 결합선이 H₂C의 c 칸에 닿도록, 원소 묶음에서 수소가 아닌 첫 원소의 칸에 맞췄습니다. PDF가 도식 본문을 한 칸 들여 쓴 것은 HCOOH처럼 뺐습니다.

제13항 4·5·7(육각·오각·삼각 환핵 그림)은 PDF가 사선을 반 칸씩 옮겨 조판했습니다(삼각 환핵의 왼쪽 사선은 줄마다 0.55칸). 사선이 줄마다 한 칸씩 옮겨 가도록 칸을 정해 넣었습니다(위 '누락 예제와 정정'). 국립국어원 질의 초안 14번에 이 읽기를 묻는 내용이 있습니다.

바뀐 동작

  • 추적: 모든 칸에 규정 조항을 기록합니다. 한 칸이 여러 조항에서 나오면 모두 적습니다(국립국어원 회신 6).
  • 과학 점자: 화학식·반응식·구조식·전자 점식·가계도·전지·치식·화식·단위를 과학 규정대로 적습니다.
    • PCO₂처럼 한 글자 대문자 뒤가 그대로 분자(CO₂)이면 그 글자는 로마자로 읽습니다(제7항 2).
    • science_spatial 문맥에서는 구조식과 전자 점식을 공간 표기 형식(제11·17항)으로 적습니다.
    • 연산 기호가 있는 식에서 수 뒤의 단위(10⁻⁸cm³g⁻¹sec⁻²)를 단위로 적습니다(한글 제69항).
      • 한 글자(3m)는 변수로 남기고, LaTeX는 기존 수식 경로를 따릅니다.
  • 국어 문장 속 LaTeX 화학식: 수식 두 칸 띄우기 대신 과학 규정 배치를 따릅니다.
  • 문맥 인자: 모든 바인딩에 추가했습니다(Rust encode_in_context, JS·Python·Ruby·C·.NET·JVM·Go). 이름은 fixture의 context와 같습니다.
  • 문맥 없는 글: 문맥이 없고 한글도 없는 글은 main처럼 영어 점자로 적습니다. 화학식·구조식·전지·유전자·전자 배치·도식·물리량 같은 과학 기호로는 science 문맥에서만 읽습니다(UEB §8.8.3의 화학식은 영어 점자로 적힌 예). 수식 모양($…$, 첨자가 붙은 짧은 낱말 H₂O)이 수학 경로로 가는 것은 main 그대로입니다. 한글 문장 속 화학식은 지금처럼 과학 규정으로 적습니다.
    • 그래서 english/rule_8_8_3의 HOCH₂는 main과 같이 문맥 없이 통과하고, 한글이 없는 과학 fixture 83건에 science를 달았습니다(그중 48건은 달지 않으면 틀림).
    • 말뭉치 변화 0문장. 문제지는 33행이 바뀝니다. 한글 없는 31행(대부분 화학 보기 ① $CH_{3}CH_{2}CH_{2}Br$ …)은 수학·영어 경로로 적히고, 한글 문장 속 ${I}_{3}^{-}$ 2행은 제68항대로 로마자표 ⠴를 앞세웁니다.
    • 그 31행에서 드러난 기존 결함: 수식 첫머리의 괄호 붙은 대문자(A(g), Y(s))는 수학 경로에서 대문자표가 빠집니다. 한글 문장 속 $F(x)$도 같아서 이번 변경과 별개로 남아 있습니다.
  • 국립국어원 회신 반영:
    • 동그라미 대문자는 로마자표·대문자표를 적습니다.
    • 총곱 ∏는 점역하지 않습니다.
    • 수식 안의 말줄임표는 수학 규정을 따릅니다.
    • 식 안의 가운뎃점은 곱셈으로 읽습니다.
  • 쌍점(제51항): 표제 뒤에 한글 내용이 오면 쌍점 뒤를 한 칸 띄웁니다(+68문장).
    • 로마자·숫자 표제(A:우리나라는)와 괄호로 끝난 표제(A(정 셰프):도저히)가 해당합니다.
    • 비율 200만:1과 맞세운 대비 찬성(5):반대(5)는 [다만 2]대로 붙입니다.
    • 한글 대비 쌍 꼴(청군:백군)은 그 쌍만 적었거나 쌍점 항목이 더 있을 때만 붙입니다. 문장 속에 홀로 선 쌍(뉴:홈 공급, 에오스:블루 온라인)은 제목과 부제를 가르는 쌍점이라 띄웁니다(회신 2026-09-11, +11).
    • 비율 뒤 괄호에서 항을 밝힌 쌍(69:31(남성:여성), 4대 6(지방비:국비))은 '대'로 읽어 붙입니다(+9).
  • 방점 판정(회신 5): 낱말 사이의 가운뎃점(5·18)을 방점으로 보지 않습니다.
    • 그래서 가운뎃점과 한자가 함께 든 문장(5·18 … 고(故))이 더는 오류로 끝나지 않습니다.
    • 인쇄되지 않는 이체자 선택 문자(U+FE0F)는 버립니다(회신 11).
  • 칠한 세모: 어절 끝에 붙은 ▲(씨▲ 도의상)도 어절 첫머리와 같은 글머리 기호로 적습니다(+1).
  • 위 첨자: 이어진 위 첨자(g⁻¹)를 위첨자 기호 하나 뒤에 적습니다(⠘⠔⠼⠁, 수학 제18항).
  • LaTeX \right.: 그려지지 않는 구분자이므로 적지 않습니다(회신 11).
    • 이전에는 \left.x\right.가 ⠭⠄로 나왔습니다.
    • \left\{ … \right.만 수학 제6항 1 연립식 괄호의 둘째 칸 ⠄으로 남깁니다.
  • 활자 변이 기호: 따옴표로 친 억음 부호 `는 작은따옴표(제49항), 두 낱말 사이의 ˙(U+02D9)는 가운뎃점(제50항), 굵은 세로선 ┃(U+2503)는 세로선 |(제71항)로 적습니다(+3).
    • ▶는 제72항이 점역자 정의 글머리 기호로만 받아 점역자 주가 필요하고, 회신 11이 다른 기호의 점형을 빌리지 말라고 했으므로 오류로 둡니다.
  • 수 뒤 +(제68항 [붙임 2]): 위 첨자로 적는 것은 등급(1++등급)뿐입니다. 50+캠퍼스, 4.1+실업률의 +는 덧셈표로 적습니다.
  • 한글 낱말에 붙은 괄호 주석(제34항)은 수식이 아닙니다: 다음은 모두 수학 제11항 두 칸 없이 한글 괄호로 적습니다(+62).
    • 수의 변화 영업이익(100→95): 일반 문장의 화살표는 한글 제70항(회신 8)
    • 나열 동구(1·2), 단위 종목(1000m, …), 부호 전기장비(+12p)
    • 대괄호 한미약품[128940], 다음 어절에서 닫히는 괄호 시간대(09:00 18:00)
    • 한글 낱말 뒤의 빗금·가운뎃점 무선충전/NFC, 한·EU: 앞 항이 한글이므로 분수선·곱셈 점이 아닙니다(제49항).
  • 규칙 적용 순서: 쌍점·쌍반점 나누기와 괄호 안쪽 틈 닫기를 정규화 단계 끝으로 옮겼습니다. 나뉜 낱말이 뒤따르는 대문자 규칙을 거치게 되어 평화음악제:MONOLOG, 회의( WCPFC,의 대문자표가 살아납니다(+87). 문제지의 ○장기총비용함수:$TC(q)=…$ 같은 LaTeX도 표제에서 떨어져 수식으로 점역됩니다.
  • 별표·참고표(제60항): 별표가 없는 국어 글의 ※는 별표와 같은 ⠐⠔으로 적습니다([다만 1]의 ⠸⠔은 둘을 가려야 할 때만). 글을 여는 별표(*조용제 지사 이야기는) 뒤는 한 칸 띄웁니다(+23).
    • 말뭉치가 양쪽으로 적어 되돌린 것: 수 뒤 마침표 다음의 혼동 한글 띄기((26.토트넘), 로마자 뒤 쌍반점(R&D;에), 가운뎃점 규칙의 단계 이동.
  • 토큰 규칙 엔진: 할 일이 없는(Noop) 규칙이 낱말을 같은 단계의 다음 규칙에 넘기도록 모든 단계를 맞췄습니다. 지금까지는 정규화·마지막 단계만 넘기고 가운데 네 단계는 첫 규칙에서 멈춰, 뒤 규칙이 한 번도 실행되지 않았습니다. 그렇게 죽어 있던 규칙 4개를 지웠습니다(출력 변화 0, 말뭉치·fixture·문제지 모두 동일).
  • 수식 경로에서 뺀 산문 표기: 로마자 낱말 뒤 느낌표(ON!, 제49항), 로마자 약어 사이 가운뎃점(IT·SW, 제50항), 수 없는 단위(kWh), 괄호가 맞물리지 않는 수식 앞부분(secretary(비서)라), 앞 어절 괄호를 닫는 부호(김남준·29),), 수 사이 × 주석(풀HD(1920×1080)), 수+로마자 한 글자(아이폰5S,), 붙임표 범위(100-130mm)와 곱한 단위(kgf·m, 제69항), 로마자 사이 !·?(제29항), 열린 로마자 주석((RCEP, 29.0%), 제34항). 말뭉치 +127, 잃은 문장 0.
  • 문장 부호 띄어쓰기와 로마자 구간 경계: 제49항이 따르는 『문장 부호 해설』대로 글 끝 마침표와 다음 어절을 여는 쉼표는 앞말에 붙이고, 한쪽만 띄운 붙임표((8강)- 2010)·빗금(업그레이드/ 새일센터)은 붙입니다. 제60항 홀로 선 별표 앞 두 칸을 한 칸으로, 제18항 약어는 부호 뒤에서도(유일한(그리고), 가운데 세 점은 줄임표로 적습니다. [3] 은 수 표기, 디지털 주소 뒤 쉼표·괄호 앞 종료표 생략(제33·34항), OPS(.837) 의 소수점은 한글 쪽, FM 98.1 MHz·GS 450h 의 단위는 같은 로마자 구간(제35항, UEB 6.5.2), KIA,27개·Bq: 1초에·PM 22:35에 는 한글 쉼표·쌍점(제33·51항), AA-(안정적) 은 등급 뺄셈. 말뭉치 +192, 잃은 문장 3 — 말뭉치가 같은 모양을 두 가지로 적은 곳(‘코드북(CODE BOOK)- 우리들의 ↔ ‘세일(Sayl)- 게이밍, ··· 를 가운뎃점으로 적은 2줄 ↔ 줄임표로 적은 7줄).
    • 로마 숫자 규칙: 제36항은 로마 숫자를 해당 로마자로 적으므로 일반 로마자 경로가 이미 맡고 있습니다. 살려 보면 LG V10, i3에서 로마자 구간을 닫아 제35항에 어긋나고 말뭉치 972문장과 제36항 fixture를 잃었습니다.
    • LaTeX 분수 규칙: 수식 규칙이 같은 일을 먼저 합니다.
    • 본문 N/N 분수 규칙: /는 빗금, 분수는 \frac이라는 fixture 규약과 반대입니다.
    • 보조 용언 띄어쓰기 규칙: 아무것도 하지 않던 등록용 규칙입니다.
  • 빈 fixture 그룹 정리: 항목이 하나도 없던 math/math_5를 지웠습니다. english/rule_4_4_1은 이번에 §4.4.1 예제 5건을 넣어 다시 두었습니다.

배포

changepack을 갱신했습니다. 새 공개 API가 생겨 9개 패키지 모두 Minor로 올리고, 이번에 문맥 인자가 추가된 JVM도 목록에 넣었습니다.

Asking why a word came out the way it did meant reading the rule engine and
guessing. The encoder now offers to say so itself: encode_with_trace returns
the cells alongside a Trace, and the Trace names, for every cell, the rule
that wrote it.

Tracing is opt-in. encode never builds a Trace, so the untraced path keeps
its shape, and the numbers say the same: fixtures stay 5141 of 5141, the
corpus stays at 455,975 of 467,121, and the marker bench still reads
837 / 145 / 305 / 398.

Attribution is partitioned by the engine that owns the input, because the
encoder is really several engines and a caller must be able to tell an
uninstrumented one from a rule that declined to fire. Korean syllables are
split to the article rather than reported as one composite entry, so 안 names
제6항 for its vowel and 제3항 for its 받침 instead of naming the syllable rule
twice. A rule that matched and then skipped produced nothing and is not
recorded: that it ran and that it explains the output are different claims.

Cells no rule object writes are still accounted for. The blank between two
words, a pre-encoded run whose token rule declared no article, and 제29항's
roman indicator, continuation and terminator each name themselves through
the emitter, so an unexplained cell means a genuine gap rather than a
structural one.

Over all 467,112 corpus sentences the trace now explains every one of the
88,927,183 cells. The one shape that had been slipping through was 제35항's
numeric bridge resuming into a lowercase a-j, where UEB 6.5.2 makes the
emitter write a continuation cell — 298 of them, plus two roman indicators
on the same path.

What remains unexplained is the capitals and grade-1 indicators on the
pure-UEB path, where attribution places whole attempts of the contraction
search and an indicator belongs to no attempt. A Korean document never
reaches it; exactly one corpus sentence takes that path at all. The counts
are pinned in a test rather than absorbed into a catch-all slot, because a
slot that swallows anything unclaimed would make the gap unmeasurable.
@github-actions

github-actions Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Braillify testcase report

Suite Passed Total Failed Success rate
Standard testcases 5284 5284 0 100.00%
NIKL 2021 corpus 91556 93115 1559 98.33%
NIKL 2022 corpus 107917 108795 878 99.19%
NIKL 2023 corpus 122828 126693 3865 96.95%
NIKL 2024 corpus 53589 54990 1401 97.45%
NIKL 2025 corpus 80736 83528 2792 96.66%
NIKL corpus (all years) 456626 467121 10495 97.75%

Command: cargo test test_by_testcase -- --nocapture

devfive added 28 commits September 21, 2026 14:56
The linux gate wants every line reached, and the tracing work left two
stretches untested: the registry lookups for the jamo articles, the emitter
slots and an index past its engine's partition, and the whole trace shim in
the node package.

The shim's per-event conversion and its two label tables are now named
functions, because a match arm inlined in a closure can only be reached by
finding an input that produces it — and the math path, for one, no Korean
sentence reaches.
The remaining gap was the traced side of guards no test entered: a rule that
matches and then skips, the token engine's rewrite shapes, the forced UEB
encode, and the math route both when it lands and when it is thrown away.
Each now has a test that goes through the public entry point where one exists.

Two of them no input could reach. The encoder only dropped its origin table
when a transform had already changed the token count, which the traced path
never does, and the emitter built a fallback rule id inside a branch that the
same reasoning made dead. Both now sit in code that every call runs, so what
was unreachable is gone rather than merely excused.

The encoder still emits exactly what it did: fixtures 5141 of 5141, corpus
455,975 of 467,121, marker bench 837 / 145 / 305 / 398.
The spaced ampersand of 제29항 asks whether a Roman word stands on each side,
and the two answers it can reach without finding one had no test: already
encoded output, which proves nothing about what it holds, and the end of the
token stream, which proves nothing at all.
Article 51's body parts a 표제 from its 내용 with a 쌍점 written against the
label and followed by one blank, and the rule already did that — but only when
Korean stood on both sides of the colon. What the 내용 happens to be written in
was never what the article turned on, so `모델명:PN50` and `일시:2006년` were
being run straight on.

The test stays on the left of the colon. Nothing Korean in front of it means
the mark was never a 쌍점 at all but a sign inside a Roman identifier, which is
how `NVH:Noise` and `A:IR` keep running on. [다만 2]'s exceptions are untouched
for the same reason: `오전 10:20` and `요한 3:16` have a figure in front, and
`청군:백군` is still caught by the 대비 쌍 test above.

25 more corpus sentences read correctly, 456,000 of 467,121. The marker bench
reads 838 / 145 / 305 / 398 against 837 / 145 / 305 / 398, one more error over
66 more sentences that now line up word-for-word and enter the comparison at
all (8,572 to 8,638). On the shared sentences the markers are unchanged.
A 가운뎃점 between figures names an event or an issue — 제주4·3, 광주5·18,
통권 제54·55·56호 — and 제5항 writes it ⠐⠆. Print often sets the Korean that
names them against the figures, and the detector judged the whole word, so the
Korean prefix made it fail the numeric test and the word fell through to the
math route, where the dot became a product and 제11항 wrapped it in two blanks.

The prefix is not noise in that judgement; it is the evidence. Reading past it
and judging the figures alone leaves 제주4·3운동 and 제54·55·56호 exactly as they
were, since neither reached the math route to begin with.

25 more corpus sentences read correctly, 456,025 of 467,121, with the marker
bench unmoved at 838 / 145 / 305 / 398.
The tracer answers which rule wrote each braille cell, but sixteen rules
answered with a placeholder rather than an article: two carried a word where a
number belongs, and the rest inherited the trait default that exists precisely
to say nobody has checked yet.

rule_map.json at the repo root settles them. It holds all 448 articles with
their text, so each rule was matched against the one whose wording describes
what it does — 제46항 for the operator spacing that gives rule_math its name,
제53항 for the ellipsis, 제47항 for both fraction detectors, 제60항 for the
asterisk, 제54항 for the quote attachment and for the tortoise-shell gloss that
turns out to be the same bracket rule, 제35항 for the digital notation, 제19항
for the head of the 옛 글자 articles the middle-Korean detector switches into,
제49항 for the two dash and bracket spacing rules added earlier, and RUEB 8.4
for the capitals run, which no Korean article governs.

Two take the honest marker the emitter's inter-word blank already uses: the
space rule encodes the blank cell itself, and the LaTeX merge only joins a
formula split across blanks and writes nothing.

Over 5,160 traced sentences the placeholders fall from 4,331 to 2,378, and what
remains is the math symbol dispatch, which needs an article per branch rather
than one for the whole chain. Behaviour is untouched: fixtures 5141 of 5141,
corpus 456,025 of 467,121, marker bench 838 / 145 / 305 / 398.
The English path wrote its indicators straight into the output without telling
the tracer, so every capital sign and grade-1 sign belonged to no rule at all.
`A1` left one cell unexplained and `Q50 2.2d` left three, and the landing
page had to show those cells as coming from nowhere.

The indicators now carry their RUEB sections: the capital-letter sign is 8.3,
the capitalised-word sign is 8.4, and the grade-1 sign is 5. Spaced numeric
output such as the `2.2` in `Q50 2.2d` takes 6.

Recording them is not simply another attempt. `settle_word_attribution`
credits a spelled-out word to 4.1 only while `attempt_count()` has not moved
since the word began, so booking indicators as attempts would have silently
stripped attribution from the very words they decorate. Indicators are
therefore collected on their own cursor, counted separately, and subtracted
from the broad fallback ranges during alignment so no cell is claimed twice.

The test that pinned the gap open is now three tests that pin it shut, and a
capitalised spelled-out word guards the hazard above.

Attribution only: no output cell changes. Fixtures 5141 of 5141, corpus
456,025 of 467,121, marker bench 838 / 145 / 305 / 398.
The math engine answered the tracer with a placeholder 2,378 times over 5,160
traced sentences. One rule, MathSymbolRule, dispatches more than thirty symbols
to as many different articles, so a single article per rule could never be
honest: the arrow arm follows 제10항, the set arm 제60항, the quantifier arm
제61항, and one number had to stand for all of them.

A rule may now own several registry slots. It declares the extra articles in
variant_metas() and returns the one it actually used through ConsumedWithMeta,
carrying the static itself rather than an index that would change meaning the
moment the list is reordered. Dispatch resolves it by pointer identity within
that rule's own declarations and refuses anything undeclared, so a rule cannot
borrow a neighbour's article by accident.

The flattening forced a second fix. Ids came from the rule's physical position
in the dispatch vector, which stops matching the registry as soon as one rule
occupies several slots; every rule after MathSymbolRule would have resolved to
someone else's metadata. Dispatch now tracks the flattened base instead.

The 221 shortcut characters carry their article in the same record as their
cells, so the two cannot drift apart. Grouping them by article also made the
table say which rule governs what, which it never did before. Seven symbols
keep the honest placeholder because the standard does not name them.

Two comments were simply wrong and are corrected: 제26항 is 행렬, not the
product sign, and @9 is 제61항 1 부정, not 닮음, which is ,' under 제42항.

Every one of the 221 character-to-cell mappings is unchanged, checked pair by
pair against the previous table, and output bytes are identical across the
fixtures and all 467,121 corpus rows. Fixtures 5141 of 5141, corpus 456,025,
marker bench 838 / 145 / 305 / 398. Placeholders fall from 2,378 to 11.
Seven shortcut characters kept the honest placeholder because searching the
standard's text for the character itself found nothing. That was the wrong
search. The standard prints its examples in the internal braille notation, not
in the Unicode characters an editor types, so a symbol is found by matching the
cells it produces against the notation the article prints.

Matched that way, five of the seven name their article plainly. The fraction
slash writes ⠌, which 제7항 1 calls the 분수표 and prints as /. Script R and
the not-similar sign write ⠠⠗ and ⠨⠈⠔, which 제34항 prints as ,R and .@9
for 관계가있다 and 관계가없다. The superscript c writes ⠘⠉, printed as ^c for
여집합 in 제60항 5. The not-implies sign writes ⠨⠒⠒⠕, printed as .33O for
항진명제의 부정 in 제61항 4.

The same reading corrects 제34항's own description: it is the relation-symbol
article, and the negation mark already filed under it belongs to its second
clause, not to some separate rule about negation.

Two symbols still have no article and now say why. ∏ carries the cells of
Greek capital pi, which invites filing it under 제13항 by resemblance; the
standard never mentions it, so it stays unattributed on purpose. ⸩ stands in
for LaTeX's \right., a null delimiter with nothing printed for an article to
govern.

Only the article attached to each entry moved; all 221 character-to-cell
mappings were compared pair by pair and are unchanged. Placeholders over 5,160
traced sentences fall from 11 to 3, all three now the null delimiter.
A \$...\$ span standing alone is a formula and the math engine owns it. The
same span inside English prose stays on the UEB side, where the parser turns it
into a Technical token and rule_11 encodes it. The engine emitted those cells
and told the tracer nothing, so a sentence like \�bc \{D}\$\ explained its
first four cells and left the other six belonging to no rule at all. Where a
whole line was formula and prose together, nothing at all was explained.

RUEB calls this code switching: 14.6.2 is the short inline fragment among
ordinary text, which is exactly this token's shape. 14.6.3 covers the long
terminal passage and is not what the parser produces here.

The cells are recorded on their own channel rather than as a contraction
attempt, because a word is credited to 4.1 only while the attempt count has not
moved since it began; booking these as attempts would have stripped the
attribution from neighbouring spelled-out words. Alignment places the direct
spans first and then searches for word and indicator cells outside them, so no
cell is claimed twice.

Two smaller holes closed with it. The code-switch encoder may try a span, emit
records, then refuse the input and hand it back; those records used to survive
into a trace they no longer described, and are now rolled back to a checkpoint.
And the symbol arms that finish through a shared continue never settled their
pending output, which is why fullwidth = and + went unexplained.

Attribution only: no output cell changes. Over 5,160 traced sentences the
unexplained cells fall from 1,892 to 449 and fully explained sentences rise from
5,119 to 5,127. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398.
Chemistry and engineering papers reach for code points the standard names but
this table did not carry, so eight otherwise ordinary lines failed to transcribe
at all rather than producing a single wrong cell.

Two are the same symbol drawn differently. A reaction arrow is written long,
U+27F6, where geometry writes U+2192; an n-ary product is written U+2A09 where
arithmetic writes U+00D7. Same meaning, so same cells and same article as the
form already present — anything else would report two articles for one role
depending on which glyph an author typed.

The third was a gap rather than a variant: 제27항 defines 나누어떨어진다 as \
and 나누어떨어지지않는다 as .\, but only the negated form was here. The plain
sign is now its undotted counterpart, which a test pins so the pair cannot drift.

Nothing already encodable changes: all 221 previous character-to-cell mappings
were compared pair by pair and are untouched, and no line that transcribed
before transcribes differently. Fixtures 5141 of 5141, corpus 456,025, marker
bench 838 / 145 / 305 / 398. Over the 5,160 traced papers, failures fall from
18 to 10.
Engineering papers write resistance with U+2126 OHM SIGN rather than U+03A9
GREEK CAPITAL OMEGA. Unicode declares the two canonically equivalent — the ohm
sign decomposes to omega and nothing else — so they are one character wearing
two code points, and no transcription may tell them apart.

We told them apart. Five papers carrying \[Ω]\ failed outright while the same
text with capital omega transcribed fine. The English side never saw the problem
because it normalises before it reads, which is why \R[Ω]\ already worked while
a bare \Ω\ did not.

The sign now shares omega's cells and article. Folding it in the pipeline
instead would have meant normalising ahead of the NFD step 제65항 5 relies on for
accented Latin, for a gain of exactly this one character: across the 5,160
papers and all 5,141 fixtures, the ohm sign is the only thing normalisation
would have touched.

Nothing already encodable changes: the previous 224 mappings were compared pair
by pair and are untouched, and no line that transcribed before transcribes
differently. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398. Failures over the traced papers fall from 10 to 5.
제49항 defines one 홑낫표. Unicode spells it twice, 「 」 at full width and
「 」 at half width, and a statute quoted in a survey paper used the half-width
pair, so the line failed to transcribe while the same words in the full-width
pair transcribed fine.

The standard's own prose writes it half width — 「통일영어점자 규정」 appears
that way in the articles we transcribe from — so refusing the form is refusing
the regulation's own typography.

Both spellings now take 제49항's cells, pinned by a test that compares the pair
rather than restating the cells, so the two can never drift apart.

Nothing already encodable changes: no line that transcribed before transcribes
differently. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398. Failures over the 5,160 traced papers fall from 5 to 4.

The remaining four are not gaps we may close by inference. Circled capitals
Ⓐ Ⓑ need 제64항's wrapping combined with 제28항's capital sign, and the article
demonstrates only the lowercase ⓐ, leaving the order of the two indicators
undetermined; that is now a question for 국립국어원 rather than a guess. ≫, ℧
and ℑ appear nowhere in the standard at all.
Every rule the tracer can credit should cite the article it implements, and the
last few commits settled them one family at a time by reading each file. Reading
is how the last one was missed: a search for the placeholders I knew about --
"?", "math", "space" -- cannot find a placeholder nobody thought to look for.

The registry now answers instead. Walking every engine's registered rules and
demanding each section be an article number, a dotted RUEB section, or "-" for
output the standard prescribes without giving it an article, the test found
unicode_fraction_encoding filed under the section "fraction". 한글 제47항 governs
it, and says so with the very character class the rule handles: "분수는 분수표
/을 사용하여 분모, 분수표, 분자 순으로 적고", worked through as ⅔ → #c/#b.

Corpus measurement would never have caught it. The 5,160 traced papers write
their fractions in LaTeX, so this rule never fires there; the defect was real
and silent at the same time.

The one placeholder that remains is listed by name rather than skipped as a
class, so adding a new undeclared rule turns the test red -- and so does
retiring the last placeholder, which should be a deliberate edit rather than a
quiet pass. It stands for the n-ary product sign, which the standard never
mentions, and the right double parenthesis, which represents a LaTeX delimiter
that prints nothing.

No output changes. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398.
수학 제35항 to 제39항 give one mark each and in order: 선분 @c, 호 @[,
직선 [3O, 반직선 3O, 각 ?. Three of our marks cited the article next door.

The overline that spans two points is 제35항's segment bar; it sat under
제36항, which is the arc. The two-headed arrow drawn above a pair is 제37항's
line; it sat under 제38항. The single-headed one is 제38항's ray -- an article
whose 붙임 also lends it to vectors -- and it sat under 제39항, which is the
angle and nothing else.

Nothing caught this because the cells were right either way. A reader checking
the transcription against the standard would have been sent to an article that
does not mention the mark in front of them, which is the whole failure the
tracer exists to prevent.

제23항 was checked and left alone. It gives the bar over a variable, 켤레
복소수 and 평균값, the very same @c cells as the segment bar, so the two are
told apart by code point alone: a combining or spacing macron marks a
variable, the overline spans a pair of points. A test now says so, because the
duplication otherwise looks like something to tidy away.

The 제35항 slot also had to be declared in the rule's variant list. Reporting
an article a rule never declared is refused by design, and it refused this one
-- a fixture failed the moment the article moved without the declaration,
which is the invariant doing exactly its job.

Attribution only: all 225 character-to-cell mappings were compared pair by pair
and are unchanged. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398.
The Linux coverage gate names ten lines nothing reaches. Two of them are
reachable behaviour that simply had no test.

제25항 writes a summation's bounds as a group and then leaves a blank before
the body, but only when the body runs straight into it -- a summation already
followed by a space, or one ending the expression, must not gain a second
blank. The blank itself was never exercised.

제53항 reads a middle dot as the multiplication sign when the same expression
also composes arithmetically, which is how derivative and product formulas are
written. Nothing built a token stream that put a middle dot beside an equals
sign, so the test for that condition never ran.

Both are driven through the rule directly rather than through a written
expression, because the parser reaches these token shapes only from inputs that
would exercise a dozen other rules at the same time and prove nothing about
these two.

The remaining eight lines are not missing tests: they are a blank line, a
closing brace, a method signature, two fields of a placeholder static, and a
step in the middle of an iterator chain. They need the instrumentation looked
at rather than more assertions, and that has to be read off CI because
tarpaulin cannot run here -- the workspace needs a system Python for pyo3, and
this platform's recorder miscounts by design.

No output changes. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398.
제64항 wraps a circled character in 7 7 and works the lowercase case through:
ⓐ is 70a7, the roman sign then the letter. It never shows a capital, so the
order of the roman sign and the capital sign was undetermined and four readings
all fitted the wording. Rather than pick one, the question went to 국립국어원,
who answered on 2026-09-21: the roman sign first, then the capital sign, then
the letter.

Ⓐ is therefore ⠶⠴⠠⠁⠶. A test pins the order with that provenance written
down, because nothing in the article can be used to check it.

The gate that decides whether a character is an enclosed symbol at all listed
only digits and lowercase letters, so the capitals never reached the encoder
even to fail there. They do now.

Exam papers label their choices Ⓐ Ⓑ, and a paper carrying them previously
failed to transcribe as a whole. Fixtures 5141 of 5141, corpus 456,025, marker
bench 838 / 145 / 305 / 398 all unchanged.
The standard names the summation in 수학 제25항 and writes it ,.S, which is
Greek capital sigma. Nothing names the product. We wrote it ,.P anyway -- Greek
capital pi, by analogy -- and marked the article unknown, which reads as an
article we merely have not found yet rather than an extension we invented.

국립국어원 answered on 2026-09-21: the product sign cannot be transcribed. So
the table no longer carries it, and an expression containing it is refused
instead of quietly given cells the standard never granted.

The cells being exactly Greek capital pi's is what made borrowing them look
reasonable, so a test records that and asserts the character is absent, or the
gap invites the same repair again.

The summation is untouched and still cites 수학 제25항. Fixtures 5141 of 5141,
corpus 456,025, marker bench 838 / 145 / 305 / 398.
수학 제38항 writes a ray as 3o,,AB and 제37항 writes a line as [3O,,AB: the
arrow ahead of both capitals, attached. 제10항 writes the arrows that stand
between things -- right, left, up, down and the four diagonals -- and lists the
right arrow among them. We sent every right arrow to 제38항 and every
two-headed arrow to 제37항, so a reaction equation and an ordinary mapping both
claimed to be geometry.

The give-away was in our own table: the left, up, down and diagonal arrows all
sat under 제10항 while the right arrow sat alone under 제38항, which is not a
shape the standard has.

The dispatch now asks what the standard's own notation asks: does the arrow
come ahead of two capitals with nothing before it. If it does, it is drawn over
them and the geometry articles apply. Otherwise it is standing between two
things and 제10항 does.

국립국어원 answered on 2026-09-21 that a chemical reaction arrow follows
과학점자규정 제18항, whose text spells the same cells (+ 5, → 3o, ← {3, ⇄ [7O).
Telling a reaction equation from any other standing arrow needs a chemistry
signal the math engine does not carry, so reaction arrows now reach 제10항
rather than the geometry article they had before -- closer, and honest about
what we can determine. The science article is recorded as remaining work.

Limits are untouched: the arrow in \lim_{n \to \infty} never reaches this rule.

No output changes -- both branches call the same encoder and all 224
character-to-cell mappings were compared pair by pair. Since the cells are
identical either way, only an article assertion can catch a regression here,
and one now does. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398.
Removing the n-ary product from the symbol table left nine integration
snapshots still expecting cells for it. They failed on every platform, and I
did not see it because I ran the suite with --lib for the preceding commits,
which skips tests/ entirely. The project's gate is the whole suite, and it says
so; I narrowed it and this is what that cost.

The snapshots record the Ok/Err shape on purpose -- the module says so at the
top -- so the fix is to let them record the refusal rather than to delete the
cases. Each of the nine now reads err: Invalid character, and nothing else in
the two snapshot files moved.

The Π(a,b) dispatch is untouched. It matches Greek capital pi, U+03A0, not the
product sign, and keeps its own unit test.

Whole suite: 5253 + 20 + 8 + 351 + 162 passing, tests/ included this time.
I set out to delete the cell that `\right.` emits. 국립국어원 ruled that nothing
printed means nothing transcribed, `\right.` draws no delimiter, and the cell
sat under a placeholder article -- every sign pointed one way. Deleting it broke
a regulation fixture.

수학 제6항 1 lists the brackets and includes 연립식 괄호, opening `7'` and
closing `,7`. `7'` is two cells, ⠶⠄, and LaTeX writes that opening half as
`\left\{ ... \right.`. So the cell is the brace's second half, and the sentinel
standing for `\right.` carries it. The matrix code already said as much in a
comment about `\begin{cases}`; I did not read it before cutting.

The sentinel now cites 제6항 instead of the placeholder, which is what it should
have cited all along. A test records why a thing named after a right delimiter
belongs to the bracket article, so the same deletion is not attempted again.

With that, no registered math rule keeps the placeholder and no shortcut
carries an unknown article. Over 5,160 traced papers the tracer reports an
article for every cell it explains -- the count of unknown articles reaches
zero. Output is untouched: all 224 mappings compared pair by pair, no traced
line transcribes differently, fixtures 5141 of 5141, corpus 456,025, marker
bench 838 / 145 / 305 / 398.
한글 제53항 writes the ellipsis in prose. 수학 제12항 [붙임 1] claims it back
inside an expression -- "쉼표는 " 으로 적고, 줄임표는 ,,, 으로 적는다" -- and
국립국어원 confirmed on 2026-09-21 that an ellipsis in a formula follows the
math standard.

Both write the same three cells, so nothing in the output could have shown the
wrong choice. The rule for an ellipsis met inside math cited the Korean
article, which also rendered as 수학 제53항 in the trace, and 수학 제53항 is
the derivative.

Article 12 is titled 로마자 변수 표기, so its 붙임 carrying the ellipsis is not
something a reader would guess; the rule now quotes it. A test pins the section
and the 붙임, because only the article separates the two cases.

Prose is untouched: an ellipsis outside a formula still goes through its own
rule under 한글 제53항. No output changes, all 224 mappings compared pair by
pair. Fixtures 5141 of 5141, corpus 456,025, marker bench 838 / 145 / 305 / 398.
The middle dot met inside an expression was filed under 한글 제50항, the
punctuation mark. The cells say otherwise: 제50항 writes 가운뎃점 as ⠐⠆, two
cells, and this table gives the same character a single ⠐.

수학 제2항 [붙임] is where that single cell comes from -- "점으로 표현된 곱셈
기호는 " 으로 적는다" -- so a dot standing between operands is the
multiplication sign, not the punctuation it resembles.

The article was wrong twice over. It named the wrong rule, and because the rule
runs in the math engine the trace rendered it as 수학 제50항, which is 무한대.

A test compares the two tables' cells for the same character, since that
difference is the whole evidence and a section number alone does not show it.

No output changes, all 224 mappings compared pair by pair. Fixtures 5141 of
5141, corpus 456,025, marker bench 838 / 145 / 305 / 398.
The trace decided the series from the engine that produced the cell: math
engine, therefore 수학 제N항. That is wrong whenever a rule implements an
article from another series, and several do. Circled numbers are 한글 제64항
and the math engine encodes them, so the page read them as 수학 제64항, which
is 햇. The degree sign, the colon, the semicolon and the tortoise-shell gloss
are the same shape of error.

Each rule already records the full citation in standard_ref -- "2024 Korean
Braille Standard, 한글 제64항" -- and nothing exposed it. The span now carries
standard_ref and subsection alongside the number, and the label reads the
series from the citation, falling back to the engine only when the citation
does not name one.

The article data was right the whole time; only the display inferred. No cells
change and no rule's article changes.

Fixtures 5141 of 5141, corpus 456,025, marker bench 838 / 145 / 305 / 398.
The middle-Korean detector counted any CJK ideograph as strong evidence that a
word belonged to a historical text, alongside the old jamo and the private-use
syllables. 국립국어원 answered on 2026-09-21 that the mode follows the 옛글자 --
"옛글자가 들어가면 그 글자에 대하여 그렇게 표기합니다" -- and that a Hanja
cannot be transcribed as itself at all, so it is evidence of nothing.

A modern sentence quoting one in parentheses -- 플로리다주(州)로,
지천명(知天命)의 -- is ordinary Korean, and the corpus has seven such sentences.

The trigger turns out to have been inert: with it gone, the regulation fixtures
still pass 5141 of 5141, the corpus still matches on 456,025, the marker bench
still reads 838 / 145 / 305 / 398, and all seven Hanja sentences still agree
with their reference. It fired only where the old jamo or the private-use
syllables were already firing.

A test now pins the negative, because historical texts are in fact full of
Hanja and the range looks like it belongs.

The judgment unit was already the word with a look at its neighbours, which is
finer than the sentence the reply describes, so nothing there needed changing.
Asked which article to cite when one decision rests on several, 국립국어원
answered on 2026-09-21: 다 적습니다. A section may now carry a list.

The English-context punctuation rule is the case that prompted the question. It
does three things at once and each has its own article: 제33항 keeps a comma
between Roman and Korean in the Korean shape, 제34항 drops the Roman terminator
when brackets or quotes enclose the Roman text, and 제49항 gives the punctuation
its cells. It cited 제49항 alone, and its standard_ref pointed at two chapters
by number -- Ch.4 Sec.10 + Ch.6 Sec.13 -- which no reader could turn into
articles.

The registry guard accepts a list without also accepting a rule that never
chose an article, and both the page and the trace harness render one as
제33항·제34항·제49항 rather than the 제33, 34, 49항 a naive join would give.

Over 5,160 traced papers the rule reports all three articles 2,402 times. No
output changes and no other rule's article moves. Fixtures 5141 of 5141, corpus
456,025, marker bench 838 / 145 / 305 / 398.
The placeholder existed as the trait's default so an unchecked rule reported
itself as unattributed instead of borrowing an article. Every math rule now
names a real one, which left the default unreachable and the placeholder static
never read -- two of the ten lines the coverage gate is holding out for.

Making meta required turns the runtime guard into a compile error: a rule
without a checked article no longer builds. That is the stronger statement, and
it is what the placeholder was standing in for all along.

The three dummy rules in the dispatch tests take a stand-in article, marked as
such, since they exercise dispatch and never reach the registry.

No behaviour changes. Fixtures 5141 of 5141, corpus 456,025, marker bench
838 / 145 / 305 / 398.
devfive and others added 30 commits September 25, 2026 11:02
…ose, and read middle-dot runs as 줄임표

- 제18항 [다만] withholds the abbreviation only when a letter touches it; a bracket or period does not (`유일한(그리고`, `컸다.그러면서도`).
- Digits with brackets (`[3]`, `[장면 2]`, `3-2[6-4,`) stay number notation like `02-799-1000`.
- 문장 부호 제21항 [붙임 1·2]: three or more middle dots are a 줄임표 (`리메이크···‘그땐`).

Corpus 456,494 -> 456,529 (+37, -2). The two lost rows write `···` as three 가운뎃점; seven other rows write it as 줄임표.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
- 제49항 (『문장 부호 해설』 7): a slash is written attached, or spaced on both sides between multi-word phrases. One spaced side (`업그레이드/ 새일센터`, `(28.83㎡ /안내`) is neither, so it joins; a standalone `/` or `//` keeps its spaces.
- The one-sided hyphen join now also covers digits (`0- 2로`, `1877- 0090`).
- Snapshot `detect_bracket_digits`: `[123]` now writes the 제49항 대괄호 ⠦⠆ ⠰⠴ (from 0eb0793).

Corpus 456,529 -> 456,573 (+44, 0 lost).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
제33·34항: a comma or bracket after `http://www.justware.co.kr` or `www.usquare.co.kr` is punctuation between Roman and Korean, so the digital-notation word no longer appends ⠲ before it. A Korean particle or a 제33항 [다만] `~` still takes it.

Corpus 456,573 -> 456,582 (+9, 0 lost).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
… section

- A Korean opening bracket bounds the look-behind like a closing one, so `OPS(.837)` writes the decimal as Korean number text instead of reopening the section. A bracket after a bracket keeps its own rule (`((W)`).
- 제33항 citation years take a distinguishing letter (`1998a, 1998b`); a 제69항 unit letter (`1000m, 1500m,`) is a measurement and continues the section as the unit path does.

Corpus 456,582 -> 456,605 (+23, 0 lost).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
제35항 `MP4 Player`: in `FM 98.1 MHz` the number continues the `FM` section, so the unit that follows takes no second Roman indicator. A unit after a number outside a Roman section (`가 98.1 MHz`) still opens one.

Corpus 456,605 -> 456,609 (+4, 0 lost).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
… a digit-like unit

- 제33항: in `KIA,27개` the comma stands between Roman and a number running into Korean, so 제41항's attached Roman comma steps aside for the Korean one. A bare number (`A,1`) keeps the Roman comma.
- UEB 6.5.2: where a Roman section continues through a number, a unit starting with a-j takes ⠰ in place of its dropped ⠴ (`GS 450h`, `ES 300h`); after a comma it takes neither (`nit, cd/㎡`).

Corpus 456,609 -> 456,620 (+11, 0 lost).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
…before its annotation

- 제33항 treats the comma and colon alike: `(Bq: 1초에` and `(Amir Kabir: 1807-1852)가` take the Korean colon, sharing the Korean-number test with 제41항's comma.
- 제34항: a 2-3 letter grade keeps its UEB minus before an attached Korean annotation (`AA-(안정적)에서`).

Corpus 456,620 -> 456,626 (+6, 0 lost).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The 제33항 colon test now looks at the full digit/colon group after an attached colon: `PM 22:35에` (제51항 time) and every colon of `52:24:4이고` take the Korean cell together, instead of only the last one. A standalone spaced colon (`KDS 14 20 20 : 2021`) keeps the Roman cell.

Corpus unchanged at 456,626 (0 lost).

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The separated-unit branch never saw a numeric previous word (the number is already pre-encoded there), so its UEB 6.5.2 check was dead and left an uncovered line. The ⠰ now lives only in the in-word number-chain step, which `GS 450h` exercises; output is unchanged.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
A text with no Hangul went to the science notation whenever its shape looked
chemical, so `HOCH₂` came out as ⠠⠠⠠⠓⠕⠉⠓⠰⠼⠃⠠⠄ and the UEB 8.8.3 fixture
needed an english context to pass. With no Hangul and no context the text is
English, as it was on main, and UEB 8.8.3 prints its chemical formulae in
English braille.

The science shapes -- diagrams, quantities, chemical and structural formulas,
galvanic cells -- are now read only in the science context or inside Korean
text. A formula in a Korean sentence is written as before. Maths-shaped input
(`$...$`, a short scripted word such as `H₂O`) keeps the maths route, as on main.

HOCH₂ drops the english context it gained in this PR and passes as it did on
main. The 83 science fixtures without Hangul carry the science context (48 of
them fail without it); every one already passed in that context.

Fixtures 5284/5284, science 143/143. Corpus unchanged at 456,626/467,121, row by
row. Workbook: 33 rows change -- 31 without Hangul (mostly chemistry choices)
now take the maths or English route, and 2 Korean rows write `${I}_{3}^{-}$`
with the Roman indicator of Korean article 68. Suite 5671 + 20 + 8 + 351 + 162,
fmt and clippy clean.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Coverage relied on default-mode inputs without Hangul to reach two branches:
`a∥b` ran the cell rule to its end without finding a cell, and `cm³g⁻¹⁰` read
a zero superscript in a unit run. Those inputs now take the English or maths
route, so the cell rule is tested on parallel lines in Korean text and the unit
run on a two-digit exponent directly. No output change.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Every braille cell printed in the 2024 Korean braille rules and in the Rules of
Unified English Braille 2024 was matched against the fixtures. This adds the
395 rows that had none; no existing row changes.

- English (383): the rows of the section symbol and contraction tables that no
  fixture held alone (wordsigns, groupsigns, shortforms, Greek and early
  English letters, arrows, currency, punctuation and more), Appendix 1
  (shortforms) and Appendix 2 (word list) entries found in no other fixture,
  and the missing examples of 2.4.7, 2.6.2, 2.6.3, 4.4.1, 9.1.3 and 15.2.
- Korean (5): the lone symbols of articles 52 (//), 59 (;) and 60 (* and ※),
  and the title line of the article 72 example.
- Science (7): the first line of article 1, the NH3 structural and electron-dot
  formulas of article 4, the pressure and oxidation-reduction reactions of
  article 18 (both forms of the latter), and the electrode of article 21.

A lone arrow, question mark or nondirectional double quote is printed without
the grade 1 indicator in its symbol table and with it where it stands alone in
text, so those rows accept both forms; ※ likewise accepts the form article 60
[다만 1] uses. Left out: the apostrophe of article 61 and the salt bridge ∥ of
science article 21 (their print is another rule's standalone symbol), and
Appendix 2's `Al`, which 10.9.7 prints with a grade 1 indicator.

Rows the engine does not yet write as printed stay failing; the PR description
lists them.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Thirteen English fixtures held the engine's output as their answer while their
internal field held the braille printed in the Rules of Unified English Braille
2024; nine of them said so in a note ("aligned to current encoder output").
Their expected and unicode now come from the printed braille:

- contractions the rules use: afternoon ⠁⠋⠝ and from ⠋ beside the line
  indicator, herself ⠓⠻⠋ (2.6.3), ar ⠜ (4.2.2), ea ⠂ (8.5.6), and ⠯ (9.1.3),
  ed ⠫ and en ⠢ (9.3.2);
- the linear poem format with ⠸ (2.6.3), the capitals word UEB and the full
  stop after the underline terminator (9.1.3), no underline symbol indicator
  before the closing full stop (9.8.1), and one blank where 8.5.6 breaks the
  braille line after DIGEST: (the print leaves three spaces).

Four inputs had lost their print styling and now carry it: the italic passage
with an en dash instead of quotation marks (2.6.3), bold Ž written as bold Z
with a combining caron (4.2.2), the italic X of EXAMINE (8.4.2) and the dotted
underline of "much" and "mine" (9.5.1, U+0323 as the parser reads it).

Five fixtures whose unicode already broke the line where the print does now
break it in internal and expected too (11.6.1, 15.1.2, 15.2.1). The 9.4.4
filename keeps its one-line answer and its internal drops the line
continuation indicator ⠐⠐, which the rules print only at the end of a braille
line. The 9.8.1 examples filed a second time under 9.7.3 are removed there.

Rows the engine does not yet write as printed stay failing.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
math_4 wrote -5<x<-2 without the second minus sign, and math_66 dropped the
second f of f(x+a)f(x-a)=f(x+1)f(x-1) from both input and braille. Both rows
and their LaTeX twins now follow the 2024 Korean braille rules; world and
jeomsarang stay as they were.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The integrity test converted internal to expected and unicode for the Korean,
maths and science fixtures but never for the English ones, so 18 English rows
whose three fields disagreed went unnoticed. English is checked the same way
now.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Where a fixture's input breaks lines, or the printed braille does, the fixture
held one chosen layout. It now accepts each layout the rules support, with the
same cells in every form:

- the braille lines exactly as printed, matched line by line against the PDF
  page;
- the print's own line breaks where these differ from the printed braille
  lines (an attribution or a heading on its own line);
- the same cells on one line.

This covers 11.6.1, 15.1.2 (2 rows), 15.2.1 (3 rows), the linear-format poem
of 2.6.3, 8.5.6 and the CHAPTER 6 example of 9.1.3. For 9.4.4 the printed
form keeps the line continuation indicator ⠐⠐ that ends its first line, and
the one-line form has none.

Word-division examples (10.13, 8.4.3, 8.4.4) keep their single form: the
division exists only at the end of a line, and one line would read as two
words.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The examples of Korean article 42 and its [붙임], article 43, maths article 66
and science articles 19 and 24 are printed divided across braille lines: a
number or formula goes on to the next line after the continuation sign ⠠, or a
formula breaks after an operator (maths article 66). Their fixtures held only
the joined line, so the division these articles describe was in no fixture.

Each now accepts the printed lines as well as the joined line. The printed
form was matched line by line against the PDF text, and the joined form is the
printed lines with the continuation sign dropped.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The extraction kept only single words, so the Appendix 2 entries holding a full
stop, a space or parentheses were left out: alt., IT (Information Technology),
mst file, so la ti and US (United States). No other fixture held them.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Science article 13 prints its hexagonal, pentagonal and triangular ring
drawings in spatial form, and the [부록] lists the structures, weather,
cloud-cover, front and circuit symbols a transcriber's note may use. None of
them had a fixture.

- 제13항 4, 5, 7 (science_spatial): Kekulé benzene, naphthalene, benzene
  with H, m-xylene, cyclopentane and cyclopropane. The input draws the ring
  with / \ | and doubles a stroke where the braille doubles the cell for a
  double bond. The PDF sets its blanks half a cell wide, so the diagonals look
  half a cell apart; the cells are placed so that each diagonal moves one
  cell per line, with the blanks inside a line as printed.
- [부록] (science/appendix, science context):
  - the six structures, answered with the symbol line and the transcriber's
    note as printed and with the note joined on one line;
  - the symbols the PDF prints as text (≡ ∇˙ ☈ ● H L);
  - the picture-only symbols (snow, cloud cover 0-10, fronts, typhoon,
    tropical depression, 23 circuit elements), whose input names the picture
    as [그림: 명칭].
- 제21항: the lone salt bridge ∥ in the science context.
- 제2항: the ion signs ⁺ and ⁻ on their own.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Article 61 writes the apostrophe ’ as ⠄; the same print is article 49's
closing single quotation mark on its own, so one of the two rows cannot pass
until the engine is told which is meant. Article 53 writes both lengths of the
middle-dot ellipsis (…… and …) as ⠠⠠⠠ and both lengths of the full-stop
ellipsis (...... and ...) as ⠲⠲⠲.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
- Appendix 3 (english/appendix_3): 61 symbols with a Unicode code point and
  no standalone fixture, mostly from the technical-material list (∞ ≤ ≥ ∈ ∀
  ∂ √ ± ∴ ∵ and so on). A glyph printed with more than one braille form in the
  rules accepts each of them (“ ⠦/⠘⠦, ” ⠴/⠘⠴, ’ ⠄/⠠⠴).
- 6.1 (english/rule_6_1_1): the ten digits with the numeric indicator.
- 15.3 (english/rule_15_3_1): the tone letters ˦ ˧ ˨; the other tone symbols
  are printed as arrows that already mean arrows, or have no glyph.
- 3.27: [open tn] and [close tn] on their own.
- Appendix 2: Al as printed there, although 10.9.7 writes a lone Al with a
  grade 1 indicator.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Every English fixture's braille now occurs in the Rules of Unified English
Braille 2024. 27 did not, and four more matched the printed cells but not
the printed layout.

- Line mode (16.2.2, 16.3.1, 16.4.2, 16.4.3, 16.5.1): the answers are the
  printed grids, matched cell by cell against the PDF. Two tables were held
  on one line, a box was a cell too narrow, a junction had an extra vertical
  cell, and the diagrams framed by dot locators ended with a bare grade 1
  terminator instead of ⠐⠐⠿⠰⠄. Three inputs change with them: the crossing
  diagram has the printed seven rows, the organisation chart gains its
  missing vertical, and noughts and crosses uses the printed lowercase x o.
- Typeforms: inputs carry the typeform the braille shows and no other.
  Headings and foreign sentences printed bold but brailled plain (8.5.6,
  13.6.4) lose their styling, as do the translations and headings that were
  italicised without a print basis (8.5.7, 13.6.4, 16.5.1); letter d and
  FAHRENHEIT gain the italics their braille shows (5.7.1, 8.5.3).
- Five rows whose print styling cannot be written in Unicode keep the
  printed braille and are marked as a limitation: bold ! ? $ € (9.2.1),
  italic digits (8.5.3) and script digits (9.3.2).
- 11.5.1 opens the radical with one grade 1 indicator, as printed.
- 7.2: the standalone long dash I added as U+2014 is U+2015; U+2014 is the
  dash ⠠⠤, as the 7.2.1 examples and Appendix 3 write it.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
The last icon example of 3.22 ("Recycling in my town.") was left out because
its icon, a question mark in a circle, has no single print character. It is
now written as ? with a combining enclosing circle (U+20DD); the braille is
the circle shape enclosing a question mark, ⠰⠫⠿⠪⠦.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
english/appendix_1 held only the twelve rule 2 examples, because the first
extraction skipped words another fixture already held. It now follows
Appendix 1 in order with every entry printed in braille: the three
exceptions to 10.9.5 (abouts, almosts, hims), the 75 shortforms heading the
list, and the examples of rules 2 to 6 for list construction. The rejected
forms of rule 3 ([not] ⠁⠃⠎ and so on) stay in the notes, not the answers.

The longer words listed under each shortform are printed without braille and
are not added.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Rows that were appended after their section now sit where the PDF prints them, and the science groups follow article order after math, so a reviewer reads each fixture alongside the book.
…r own

Appendix 3 gives each technical sign its grade 1 form, so a sign declared English now reads from that list instead of failing over to the Korean path. A sign standing alone takes its primary use (the prime, the dash, the UEB inverted marks, the grave accent alone). Eng and schwa (4.4), ratio and proportion (3.17), the level tones (15.3), a sign inside an enclosing circle (3.22) and a lone early-English capital (12.2) are read as printed. Text not declared English still leaves the technical signs to the math path.
The parser dropped the minus in -5<x<-2 to match the old fixture; the rules print 9#e99x999#b, with the minus.
The blanks of rule 60 separate the asterisk from the words around it; alone it is the sign itself.
A salt bridge alone (science 21), the weather signs of appendix 5 and a lone capital of appendix 8 are the signs themselves. A paragraph made only of a formula inside Korean text is written in the science reading, as science 1 prints the halogen list.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant