Repository navigation
Conversation
Two things, and the second is the one worth reading. **W6 merged 2026-09-04 as PR #89.** Recorded in the four status homes plus the `## W6` wave heading, which was the only wave heading with no merge marker. Each home now also records what the merge review actually returned — four rounds, three blockers between them — because every other wave's clause records its review findings and W6's said nothing: the text was written before those rounds happened, and the branch shipped its own stale status. **"43 of 48 items" was wrong. It is 39.** The audit traced the drift to a specific mechanism, and the arithmetic closes on its own: at W4's merge the count was 28. Then W5's items were counted TWICE — incrementally as they landed (28→32), and again wholesale when the wave closed (32→38). W6's correct +5 rode the inflated base to 43. 28 + 6 + 5 = 39. The figure lived in five homes, so the error propagated to all five. Verified from git history rather than argued: `git show <sha>:phase-doc | grep 'of 48'` returns 28, 32, 38, 43 at exactly those commits. `CR-63` was the other half of why the number could not be checked: its heading carried no closure mark although its body and its register row both say it closed on 2026-09-03. With that fixed, a `grep` of the headings now reproduces the status line instead of merely disagreeing with it — which is the property that would have caught this four commits ago. The convention is now stated rather than assumed: `CR-95` is counted on its non-deferrable short-term half, and its long-term half keeps the item on the board. That single item is what decides between 39 and 38, so it is named. **And the sharper fact the count was hiding is now next to it:** two of the seventeen NON-DEFERRABLE items remain open — `CR-73` and `CR-80` — so exit criterion 1 is not met today. That pair, not the 9-vs-5 arithmetic, is what gates the phase. Also corrected, all figures repeated across more than one home: - the expression scan's soundness claim was false in EIGHT ways, not seven — an eighth shape (a Unicode identifier escape, where `run` IS `run` to JavaScript and was not to the scanner) was found after the wave was declared complete. ADR-0093 carries it as a dated correction, in the wave's own discipline, rather than a rewrite. - 32 commits, not 27 — the original count predated the four maintainer folds. - ELEVEN residuals, not five. That is the one W6 figure that UNDERSTATED rather than overstated, and it was already nine when the sentence was written. - `current.md` still listed `CR-64` as open; it shipped 2026-09-02 with W6. - the `CR-62` residual's list of what the scan abandons on was missing the escape bail — the list growing again is itself the argument for the rewording it asks for, which `expression-sandbox-spec.md` and `run-plan.md` now carry. - `expression-sandbox-spec.md` still claimed the VM-side follow-up needs ADR-0027 §3/§6 amendments; ADR-0093's own correction says it does not, so two of three homes agreed and the spec was the outlier. - two roadmap documents carried Last-updated stamps older than their own content. README is deliberately untouched. It is the one home that survived six waves without drifting, and it did so by refusing to hold a number. Refs: ADR-0091, ADR-0092, ADR-0093, ADR-0094 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ady shipped, §1 has not ADR-0087 merged unapproved inside a `fix(core)` commit on 2026-08-29 and stayed Proposed. A review refused acceptance until its gaps were written down. Every finding was verified against the code before being folded; two of them were wrong and are recorded as such. What the ADR now states rather than implies: - **Implementation is partial.** §2 and §4 shipped in `W3`. §3 shipped in `W5` as `CR-54` — the pin runs at the dispatch boundary and `#settleCompleted` writes the pinned value into `#states`, and `deInlineMedia` rewrites a `base64` source exactly as it does a `url`. The register called §3 unimplemented for two waves because it landed under a different item number, and a stale `size-bounds.ts` comment kept saying so; one review read that comment and concluded the opposite of the truth. §1 is still unimplemented and the `W3` live blocker stands. - **§1 names one seam for two situations that do not share one** (C1). A host calls `createSessionHandle` directly, so the session half is buildable as written; `createRunHandle` is not exported and only `WorkflowEngine.start()` / `resumeFromCheckpoint()` produce a `RunHandle`, so the run half needs the mode plumbed onto an engine entry point first. An empty iteration is also not "loud". - **"Bounded per consumer" is not true today** (C2), and is now explicitly out of scope: `whenDrained()`'s un-parked fast path grants a slot it does not reserve. - **A terminal event is not measured** (C4). `measureDraft` returns before measuring, so "measured and never refused" described an intention no code held. The word is `exempt`, in the ADR, the code and the canonical spec. - **§4's central premise was false** (C8): an age policy needs no new host seam — `ExecutionHost.clock` has existed since ADR-0036. The decision survives on the reasons that are true. The same false premise is corrected in `engine.ts`. - C3 (a session replays typed rows, not an event log), C5 (a gate payload is counted, not refused — and a run's LAST gate is never checked), C6 (the measured value domain is undefined), C9 (§5 claims movement ADR-0086 caused). Attribution drift fixed: `size-bounds.ts`, `invariant-error.ts`, `engine.ts` and `workflow-yaml-spec.md` all pointed at ADR-0086 for decisions ADR-0087 makes, and `errors.ts` sent a reader to ADR-0086 for `ceiling_exceeded`'s rationale, which appears only in ADR-0087 §5. The three size rows are also split out of the admission-ceilings table: nothing about them is checked at admission. ADR-0036 and ADR-0042 gain the dated amendment notes their refinement requires; ADR-0042 moves from *Related* to *Amends*, since §3 re-places its `deInlineMedia` choke point. Residuals: one closed (media retained as base64), three sharpened, two added. No behaviour change — comments, docs and one ADR status. `pnpm run ci` green. Refs: ADR-0087, ADR-0036, ADR-0042 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…rrable items
Both were shipped claims the code did not keep, and both had the cheap fail-closed
option the phase register says is always available.
**CR-73 — `invoke_agent` was advertised with no delegate anywhere.** `ctx.invokeAgent`
is set nowhere in the tree, so a model granted the tool was offered it, called it, and
got `tool_unavailable` for something the engine had just advertised. `ToolDef` gains an
optional `requiresDelegate`, declared by the tool rather than looked up in a central
id→delegate switch, and one predicate — `delegateAvailable` — answers it for all three
paths that advertise: the CLI filter (`wiredToolIds`, whose new `delegates` option
defaults to NONE, which is the fix and not a detail), `AgentSession.buildLlmTools`, and
`agent-runner.ts`'s `buildLlmTools`.
The first attempt covered only the CLI filter. The run path advertises independently,
and there the defect was worse than the item recorded: `agent-runner.ts`'s
`dispatchContext` literal has no `invokeAgent` or `mediaRead` field at all, so no host
could have supplied one — the tool was structurally guaranteed to fail. The two engine
`buildLlmTools` are near-duplicates in two files, so the predicate has one home; the
first engine mutation reddened the run path and nothing on the session path, which is
the exact shape of this defect, and the session test exists because of that.
`read_media` rides along. Its delegate is equally unwired, and `CR-50`'s own residuals
argue the fail-closed absence is better than a tool answering `unknown media handle` to
every call. `CR-50` stays closed on its delivery half; this changes what the model is
told, not what the tool does.
**CR-80 — a rejected custom base URL fell back to the official API.** The `catch`
swallowed `InvalidBaseUrlError` and left the DEFAULT adapter standing, with a comment
calling that "refuse the custom endpoint" — the default adapter is `api.openai.com`. A
user pointing Relavium at an internal gateway, whose stored row later drifted to
something non-HTTPS, private or credential-bearing, silently sent prompts and their API
key to the official API. It now installs a refusing adapter: `generate` rejects (a
synchronous throw would escape a bare `.catch()`, breaking the contract the refusal is
enforcing) and `stream`'s `next()` rejects at the first pull, where a real adapter's
failure lands. `InvalidBaseUrlError` gains a `reason` field so the refusal composes
without nesting its own prefix; the message names the URL's shape, never its credentials.
Not a throw at construction: the resolver builds for every command, so throwing would
break `relavium provider list` — the command you would use to find the bad row.
**The finding under CR-80 was its test.** It asserted only
`expect(resolver.resolveProvider('openai')).toBeDefined()`, which a refusing adapter
satisfies as well as a fail-open one, so it passed either way. The replacement stubs
GLOBAL fetch, because the default adapter builds its own and an injected recorder would
have been empty under the defect too. Under a mutation restoring the fail-open, the
credential test takes ~470 ms — it really reaches out to `api.openai.com`.
Also closing the remaining pre-`W7` exit criteria:
- **Criterion 5** — the three missing security sittings are recorded with the
adversarial test per control: prompt/trust provenance, hostile MCP, provider/config
trust. Each says it was written after its wave merged.
- **Criterion 7** — `W5` and `W6` closing registers, per item, with the test that fails
if it is reverted. `CR-63`'s "closed by" is *nothing*, and the grep that established
that is named so the suspicion can be settled by repeating it.
- **Criterion 2 (partial)** — deferral records for `CR-93` and `CR-95`'s long half, with
severity, trigger and the product claim each narrows. `CR-94` is deliberately NOT
deferred: deferring a High item about money is a maintainer's call.
- **Criterion 6** — `pnpm run ci` and `pnpm coverage` both exit 0, read from `$?`.
`pnpm run ci | tail` reports tail's status and produced one false green in this
session; the criterion now says so.
Count: 41 of 48, arithmetic shown rather than asserted (40 ✅ headings plus `CR-95`'s
short-term half), because it has been wrong twice here.
Refs: CR-73, CR-80, ADR-0055, ADR-0057, ADR-0065
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…dments Four decisions the maintainer made on 2026-09-12 and 2026-09-13, accepted on 2026-09-14. Each carries an Implementation line: accepted means settled, not shipped (the ADR-0087 lesson). - ADR-0095: a session's tool history is persisted as structure and never as content. Carrying it into the model's context is deferred. The export fills its `tools` union and carries no tool content, and the authored `memory` policy decides the request and every compaction. (CR-70, CR-71, CR-72) - ADR-0096: a request is measured before it is sent. A context overflow is a classified kind in both LlmErrorKind and ErrorCode, recovered once and only before any tool round has run. A pre-content overflow releases its budget admission on positive evidence, and input is priced at every attempt. - ADR-0097: a budget approval grants a dispatch-owned allowance, shown from a field no workflow can forge and frozen durably. Exhausting it fails the step closed rather than pausing again. Model binding and expiry are deferred. (CR-94, and CR-96, a shipping defect: a crash after a budget decision resumed the agent as complete with an invented output) - ADR-0098: a session's effect row holds no result, never replays, never reuses an identity, and discloses an effect from a turn that did not complete. (CR-97, shipping defects: tool output at rest, and stale replay after compaction plus resume) Six review rounds shaped these ADRs: the maintainer's review of the first drafts, four adversarial workflow rounds, and a short sixth round. - From the second round on, most defects were in the previous round's fixes. The branch-order pattern repeated: the allowance loop came back three times, and the session effect bookkeeping was wrong twice. - After the fifth round, the maintainer had the ADRs restated as decisions, invariants and acceptance tests. Mechanism traps now sit in non-normative implementation notes, with the file and line where each was found. Dated notes are appended, not rewritten, on ADR-0062 (§6's reasoning corrected; boundary by turn; §1, §4 and §5 refined), ADR-0026, ADR-0050 (premise narrowed to what the database actually holds), ADR-0080, ADR-0028 and ADR-0074. The index README now says "three sections", which is what it lists. Refs: ADR-0095, ADR-0096, ADR-0097, ADR-0098 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…7, narrow three false claims The register catches up with the decisions accepted in the previous commit. - CR-70 to CR-72 and CR-94 now record what was decided, and link the ADR. - CR-94's section keeps one reversed clause, with its reason. The acceptance said an approved node that exceeds its lease "pauses again"; ADR-0097 makes it fail closed, because three drafts that paused again each found a new path back into the 1.AC loop. - CR-96 and CR-97 are opened as shipping defects found during the ADR reviews. Both meet this phase's own non-deferrable definition, so exit criterion 1 is open again. Recording them as deferrable to keep a criterion green is the drift the register exists to prevent. - The count is 41 of 50, with the arithmetic shown: 40 headings marked done, plus CR-95's short-term half. deferred-tasks.md gains three deferral records, each with its severity, its trigger and the claim it narrows: - carrying tool history (CR-70); - a budget allowance's model binding and expiry (CR-94); - recovering an overflow after a tool round. The note saying no maintainer call had been made on CR-94 is corrected, and two older entries now point at the decisions that superseded them. Three canonical claims were already false and are narrowed to what the code does today, ahead of the code: - agent-session-spec.md, chat-session.md and architecture/agent-sessions.md said a turn appends tool messages, and that the export carries `tools` and the full transcript. Neither is true yet. - agent-yaml-spec.md named `none` as `memory`'s default. The field is inert, and the shipped behaviour is the opposite. - workflow-yaml-spec.md said the cap is checked before each LLM call. An approved budget step runs with no check at all. CLAUDE.md, AGENTS.md, current.md and the phase status agree on one state: 41 of 50, W7 unblocked, four ADRs accepted with their implementation staged. AGENTS.md and current.md's header had been left at 39 of 48. A verification workflow checked these changes against the code and the ADRs before commit. It found 18 real issues: two sibling docs still carrying the false claims, a wrong section citation, a present-tense description of unshipped behaviour, and stale counts. All are folded. `pnpm run ci` exits 0. Refs: CR-70, CR-71, CR-72, CR-94, CR-96, CR-97, ADR-0095, ADR-0096, ADR-0097, ADR-0098 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… 113 findings, 17 decisions, two blockers Before any `W7` code, an eight-dimension review read the four `W7` ADRs, the register, the deferrals and every canonical document they touch against the tree; each finding was verified by an independent adversarial pass. 113 held, and the maintainer settled 17 decisions. ADRs are append-only, so every correction is a dated note of 2026-09-18; no ADR body is rewritten. Three document defects would each have been discovered mid-implementation: - **ADR-0098's join does not exist.** It said the effect journal already records the tool-call id and cited `registry.ts:601`, which is the model-facing `tool_result` part. Every session surface writes a wiring constant (`'session'` / `'home'` / `'agent-run'`), so the join that decides "did this turn complete" cannot be built as described. `W7` threads a per-call, engine-assigned id through `EffectDispatchPort.prepare` — a seam change, now recorded. - **ADR-0098's monotonic turn key could not be derived.** The committed-row sweep deletes the rows the high-water mark would be read from, so the key becomes a durable column on the session row; the Alternatives' "not required" is withdrawn. - **ADR-0096 §1's per-part floor** would have charged an inline 5 MB image about 1.7 M tokens, contradicting the per-modality media ceiling in the same section. The maintainer's decisions that change what ships beyond those ADRs: `history.db` opens with `PRAGMA secure_delete = ON` and checkpoints the WAL after the legacy clear and after each session sweep, so "no tool result at rest" is true of the file's bytes after a successful checkpoint — with a blocked checkpoint and pre-upgrade free pages recorded as residuals rather than claimed; a pre-content HTTP 4xx releases its admission instead of committing an input-sized estimate, which keeps a rate-limit burst from exhausting a session's cap once input is priced; the summariser is bounded by a capped four-pass fold; a budget refusal of the summariser is a budget outcome, not a failed compaction; and `relavium budget resume` moves into `W7` with `--gate` and `--approve-amount`, so an approval has a non-TTY surface and a stale quote is refused. Register: `CR-98` opened for ADR-0096's three prerequisites, which had been in `W7`'s scope with no item of their own (41 of 51); a fifth security sitting, `history.db` at rest; ADR-0087 §1 — the `W3` live blocker that belonged to no wave — scheduled into `W8`; a per-document landing checklist so the wave closes against one list rather than five. Corrected ahead of the code, because they are false today: the SSE schema's "user/assistant/tool messages are persisted", the chat reference's "full transcript" in the export, and the three canonical security claims that `history.db` "holds no credentials" and that the bus masks tool I/O (`keychain-and-secrets.md`, `database-schema.md` twice, and a dated note on ADR-0029). A review of this change itself returned two blockers, both folded: the WAL guarantee was stronger than `secure_delete` plus one checkpoint can give, and three canonical security documents still taught the old at-rest model. Refs: ADR-0095, ADR-0096, ADR-0097, ADR-0098, CR-94, CR-96, CR-97, CR-98 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Record the maintainer-approved compaction and durable budget authorization contracts before implementation. Reconcile the preflight privacy and recovery claims, preserve historical ADR bodies, and retain the independent review record and twelve-step implementation/review plan on development. Validation: pnpm run ci exits 0; relative links and diff whitespace pass. The pre-existing root coverage collection failure is step 1 work. Refs: ADR-0099, ADR-0100, Phase 2.6.5 W7 Co-Authored-By: Claude <noreply@anthropic.com>
Keep private analysis out of root tests and coverage while checking every real suite in root and package modes. Add the maintainer capture runner, typed adapter-zone request/refusal logic in llm, offline boundary checks and the canonical procedure required before W7 overflow classification. Refs: ADR-0096, Phase 2.6.5 W7 step 1 Co-Authored-By: Claude <noreply@anthropic.com>
Inspect every JSON string before saving exact provider evidence, invoke the capture runner directly, refuse replaced destinations, isolate guard probes by invocation, and keep required GitHub CI aligned with local CI. Record seven verified first-round findings and regression validation. Refs: ADR-0096, Phase 2.6.5 W7 step 1 review round 1 Co-Authored-By: Claude <noreply@anthropic.com>
Remove racy existence and directory-read prechecks from empty-parent cleanup. Reproduce the real two-process interleaving in the ownership guard and check completed capture artifacts for supplied-key collisions. Record the independent second-round review and successful validation. Refs: ADR-0096, Phase 2.6.5 W7 step 1 review round 2 Co-Authored-By: Claude <noreply@anthropic.com>
Only remove invocation-owned probes so a finishing guard cannot remove a starting guard’s parent between mkdir and mkdtemp. Replace the cleanup scheduling probe with a passing/starting regression that fails 90bae8a at the real filesystem boundary. Refs: ADR-0096, W7 step 1 corrective review Co-Authored-By: Claude <noreply@anthropic.com>
Close the owned descriptor without unlinking the caller-selected pathname on refusal. Document retained empty or partial destinations and exercise the prior real last-stat race plus safe refusal retention in offline command smoke. Refs: ADR-0096, W7 step 1 corrective review Co-Authored-By: Claude <noreply@anthropic.com>
Record the clean fifth review and qualify capture failure artifacts and inode ownership precisely. Update the fixed refusal message so maintainers inspect a failed destination before reuse. Continue the approved W7 plan with Step 2. Refs: ADR-0096, W7 step 1 Co-Authored-By: Claude <noreply@anthropic.com>
Validate session-only structural tool parts and every denormalized metadata projection on write and read. Refuse raw values without leaking unknown property names or JSON diagnostics. Allocate effect-turn keys transactionally from a session-owned high-water mark, backfilled before cleanup and never clobbered by stale transcript flushes. Shared schemas and core reader types change atomically with the store; run/event/IPC durable content stays intact. Completed-turn production and dispatch wiring remain W7 step 3. Lint, typecheck, tests, build, format, seam/dependency/tool/smoke checks and root coverage pass. The committed-migration CI check runs immediately after this commit because its contract requires clean generated migration history. BREAKING CHANGE: SessionMessage tool parts hold engine ids and byte counts rather than raw args/results. Unknown message/metadata fields and mismatched denormalized projections are now refused at persistence boundaries. Refs: ADR-0095, ADR-0098, W7 step 2 Co-Authored-By: Claude <noreply@anthropic.com>
Refuse allocation inside an active outer transaction so a later rollback cannot erase and reissue an engine effect key. Cap shared session-only media metadata at safe integers without changing generic durable media. Record both independently verified Step 2 round 1 findings and regressions. Refs: ADR-0095, ADR-0098 Co-Authored-By: Claude <noreply@anthropic.com>
Record the second independent privacy and durability review, including real SQLite boundary, cross-process lock and killed-process probes. Close the approved structural storage and key-allocation step, with Step 3 integrations and all six W7 acceptance items still open. Refs: ADR-0095, ADR-0098 Co-Authored-By: Claude <noreply@anthropic.com>
Connect shared/core/db/cli so completed exchanges atomically retain tool structure and explicit empty terminals. Share the whole-turn projection across resume, export and compaction boundaries. Forward engine-owned call identities into actual per-call effect preparation, independent of max_turns. Require host allocation and persister durability attachment. One-shot runs retain only an atomically tombstoned identity record, never a resumable chat. BREAKING CHANGE: SessionDeps requires reserveEffectTurnKey and persisters require allocator/probe attachment; AgentTurnResult includes toolHistory. Refs: ADR-0095, ADR-0098; Phase 2.6.5 W7 step 3 Co-Authored-By: Claude <noreply@anthropic.com>
Deep valid model JSON must remain correctable even when native recursive serialization exhausts its stack. Count plain JSON iteratively on that path and preserve the original policy-denial classification in completed history. Regression tests fail the reviewed implementation and pass the correction. Native escaping, omitted values, cycles and duplicate references are pinned. Clarify that durable key uniqueness does not confer session ownership, and record the first independent review round with its verified findings. Refs: ADR-0095, ADR-0098, W7 step 3 Co-Authored-By: Claude <noreply@anthropic.com>
Host scope errors intentionally omit the tool id. Verify the exact call name against registry membership before recording it so recovered read denials remain present in structural history and the exported tool union. Keep the fixed unknown marker for unresolved calls. Core and real jailed filesystem/SQLite regressions pin identity, classification and export privacy. Refs: ADR-0095, W7 step 3 round 2 Co-Authored-By: Claude <noreply@anthropic.com>
A cached idle-command key bypassed the allocator's persistence failure check. Check the live attached probe at effect preparation, independently of key allocation and provider admission. Refuse without another journal row or key, while preserving settlement and discard of effects already dispatched. Real host/SQLite regressions stop a second bang after failed trim persistence and stop an MCP effect after attempt-cost persistence fails. Record both fresh round-two findings and the independent parent verification. Update canonical contracts and keep the step open for a fresh corrective round. Refs: ADR-0095, ADR-0098, W7 step 3 round 2 Co-Authored-By: Claude <noreply@anthropic.com>
Run trusted synchronous host admission after approval and preparation, including unjournaled tools. Discard prepared claims on proven non-dispatch and preserve settlement of effects already started and retained-result run replay. Pin asynchronous admission races and unjournaled reads with regressions. Refs: ADR-0080, ADR-0095, W7 step 3 review round 3 Co-Authored-By: Claude <noreply@anthropic.com>
Check the live persister probe immediately before actual registry dispatch, including write_file overwrites that intentionally have no journal tier. Keep the prepare guard and allow settlement of effects already started. The real CLI/SQLite filesystem regression failed before correction. Reconcile the canonical export introduction and record independent round-three findings. Required checks and root coverage pass with 6,415 passing tests. Refs: ADR-0095, W7 step 3 review round 3 Co-Authored-By: Claude <noreply@anthropic.com>
Preserve nontrailing bare-user text written by the pre-W7 atomic persister. Keep it in resumed context and the next completed export prompt without fabricating terminal rows or completed-turn counts. Interrupted structural exchanges and trailing bare users still roll back. Expose matching boundary sequence slots so compaction and trimming do not resurrect legacy context. Record the dated ADR-0095 read-compatibility correction. Refs: ADR-0095, W7 step 3 review round 4 Co-Authored-By: Claude <noreply@anthropic.com>
Use the shared boundary projection to preserve old nontrailing bare-user text on reseat and to drop it correctly during compaction or trim. Real CLI/SQLite regressions pin the durable marker, resumed context and historical export. Record the independently verified round-four upgrade finding and corrections. All required checks and coverage pass with 6,422 passing tests. Refs: ADR-0095, W7 step 3 review round 4 Co-Authored-By: Claude <noreply@anthropic.com>
Record the clean fifth independent contracts and durability review over the complete step and all corrections. Retain the explicit legacy read ambiguity and validation limits. Proceed to the approved Step 4 privacy/retention plan. Required checks, root coverage and full CI passed at 6d5b3f0. Refs: ADR-0095, ADR-0098, W7 step 3 Co-Authored-By: Claude <noreply@anthropic.com>
Explain the engine-owned session call id as completion evidence. Deduplication still uses the effect identity, never its occurrence id. Refs: ADR-0098, ADR-0095, ADR-0080 Co-Authored-By: Claude <noreply@anthropic.com>
Never inspect or serialize a session result in the reference journal. Matching session prepares refuse; run result retention and replay remain. Refs: ADR-0098, ADR-0095, ADR-0080 Co-Authored-By: Claude <noreply@anthropic.com>
Clear legacy session results after high-water initialization; secure-delete and checkpoint their freed bytes. Read all historical completion evidence in one owned transaction and sweep exact captured identities atomically. Breaking-Change: sweepCommittedForSession takes captured id/address pairs instead of beforeTurn and returns deletion count plus checkpoint status. Refs: ADR-0098, ADR-0095, ADR-0080 Co-Authored-By: Claude <noreply@anthropic.com>
Reconcile only after active chat/Home publication and successful disclosure. Preserve evidence on read, sink, activation and replacement races. Sweep one-shot committed rows only after its own identity reservation succeeds. Refs: ADR-0098, ADR-0095, ADR-0080 Co-Authored-By: Claude <noreply@anthropic.com>
Document the all-history disclosure join, exact captured sweep and conditional main-file/WAL erasure guarantee. Append landing notes without rewriting ADRs and retain the accepted busy-reader and pre-upgrade freed-page residuals. Refs: ADR-0098, ADR-0095, ADR-0080 Co-Authored-By: Claude <noreply@anthropic.com>
Record both independent findings, exact final physical audits, causal controls and forced Original validation. Keep historical review records intact and round twelve, remaining groups, ADR-0102 approval and live-provider steps explicitly open. Refs: ADR-0055, ADR-0082, ADR-0085; W7 systematic group 2a Co-Authored-By: Claude <noreply@anthropic.com>
Own every LlmRequest data field, specify null-prototype working records with actual SDK generate/stream acceptance, and explicitly scope separate-endpoint media outside this proposal. Reconcile intake probe figures, name UnsupportedRequestDataError and add forward-compatibility landing notes. Keep ADR-0102 Proposed pending independent review and maintainer approval. No dependent implementation or Accepted ADR body changes. Validation: relative targets, git diff --check, Prettier. Refs: ADR-0102; W7 post-PR maintainer review Co-Authored-By: Claude <noreply@anthropic.com>
Capture typed output before pricing callbacks can rewrite provider-owned text. Price known quantities even if projection fails and preserve accounting failure precedence, exactly-once settlement and existing private non-retryable diagnostics. Whole-request ownership under Proposed ADR-0102 remains unimplemented. Refs: ADR-0082, W7 Co-Authored-By: Claude <noreply@anthropic.com>
Preserve terminal cancellation raised by a reconstructed session provider resolver before controller setup. Add independent cold-plan, actual workflow pricing and current money-barrier controls, plus pricing/projection failure precedence tests. Update older ordering controls with distinct error identities. Refs: ADR-0055, ADR-0082, W7 Co-Authored-By: Claude <noreply@anthropic.com>
Exercise the public session and native SQLite effect-journal ports with an actual registered synthetic command. Require no stored result bytes and refusal to replay the same committed session identity. No provider or paid call is involved. Refs: ADR-0098, W7 Co-Authored-By: Claude <noreply@anthropic.com>
Record two recovered interrupted reviews without full-scope acceptance, their verified cancellation/pricing repairs and current checks. Record two fresh static ADR-0102 clarification reviews while retaining Proposed and maintainer approval. Fresh complete group 2a round 13 and the remaining W7 work stay open. Refs: W7, ADR-0102 Co-Authored-By: Claude <noreply@anthropic.com>
Recheck the existing idle precondition after cold provider resolution so a nested send retains its cancellation controller. Preserve terminal cancellation, idle abort and warm memoization, with real reconstructed-session and event-handle tests. Add generated projection and host-pricing contrasts. Advance W7 systematic group 2a under ADR-0055; ADR-0102 remains Proposed and unimplemented. Co-Authored-By: Claude <noreply@anthropic.com>
Record both complete qualified round-13 reviews and independently audited seals, the proven ownership correction and 8,195 passing Original tests. Keep round 14, ADR-0102 approval, remaining systematic groups and whole-wave acceptance open. Co-Authored-By: Claude <noreply@anthropic.com>
Record complete independent round 14 reviews and audited evidence at 207b271. Keep ADR-0102, remaining systematic groups and final W7 gates open. Refs: ADR-0095, ADR-0097, W7 Co-Authored-By: Claude <noreply@anthropic.com>
Expose a refusal-only comparison using frozen priced quantities. Preserve full engine quote validation as the approval authority. Retain resolved budget-kind evidence in shared suspension checkpoints and fail refused deadline auto-approval through run_timeout. Refs: ADR-0097, ADR-0100, ADR-0028, W7 Co-Authored-By: Claude <noreply@anthropic.com>
Render safely projected excluded candidates on approval and discovery surfaces. Refuse stale rates before secrets and MCP, reuse the open resume database, and return a refused inline decision as paused. Prune settled notices and handle synchronous or asynchronous warning-sink failure. Update canonical contracts and append the ADR-0097 correction. Refs: ADR-0097, ADR-0100, ADR-0098, W7 Co-Authored-By: Claude <noreply@anthropic.com>
A synchronous human decision owns the gate while its durable authorization acknowledgement is held. The timeout failure backstop now checks that claim instead of treating the still-pending row as available. Six independent actual-runner controls preserve approved/rejected outcomes and genuine timeout/cancel behaviour. Validation: forced lint/typecheck/test (23 tasks, 362 files, 8239 pass, 12 existing skips), six forced builds, test isolation and format. Original old-code baseline: 7 expected failures / 17 passes; corrected controls: 24 passes. Refs: ADR-0097, ADR-0100; W7 systematic group 3 round 1 Co-Authored-By: Claude <noreply@anthropic.com>
An invalid inline decision stops prompting but drains the actual aggregate pause acknowledgement before releasing native SQLite and cancellation resources. A concurrent terminal retains its own result. Promote thirteen native/current-price and held-write controls; the existing driver tests now require the actual pause rather than forbidding its delivery. Validation: forced lint/typecheck/test (23 tasks, 362 files, 8239 pass, 12 existing skips), six forced builds, test isolation and format. Original old-code baseline: 7 expected failures / 17 passes; corrected controls: 24 passes. Refs: ADR-0097, ADR-0100; W7 systematic group 3 round 1 Co-Authored-By: Claude <noreply@anthropic.com>
Record both complete independent Group 3 reports, the two verified High findings and scoped correction commits, exhaustive sealed evidence verification and qualifications. Canonical commands/execution and the dated ADR-0097 note describe the claim and actual pause acknowledgement barriers. Fresh complete round 2 remains required; Proposed ADR-0102 and whole W7 remain open. Validation: 167 canonical and 692 complete changed-document relative file targets resolve; ADR-0097 historical bytes preserved; format and diff check pass. Production checks: 23 forced tasks, 8239 pass / 12 existing skips, six builds and isolation. Refs: ADR-0097, ADR-0100; W7 systematic group 3 round 1 Co-Authored-By: Claude <noreply@anthropic.com>
An invalid inline budget decision must not let an acknowledged aggregate pause outrank a later cooperative cancellation. Keep the native database and signal handler alive until the cancellation terminal is acknowledged. Eleven native regression cases cover invalid/stale decisions, held pause and cancellation writes, genuine approval/rejection, ordinary gates and human claims racing late timeout preparation. Old production fails the two cancellation cases; corrected full checks pass 8250 tests. Refs: ADR-0097, ADR-0078, ADR-0079, W7 systematic group 3 round 2 Co-Authored-By: Claude <noreply@anthropic.com>
Record both complete Group3 round2 reviews, their verified cancellation finding, exact evidence qualifications and the native regression checks. Clarify the existing command contract and append the ADR0097 correction. Keep Group3 open for fresh complete round3 and retain all remaining W7 and PR gates, including Proposed ADR0102. Refs: ADR-0097, W7 systematic group 3 round 2 Co-Authored-By: Claude <noreply@anthropic.com>
Latch prompt cancellation and suppress subsequent queued gate cards. Keep the primary event iterator through renderer unmount and consume the acknowledged terminal before releasing native SQLite or signal ownership. Ink settles cancellation before publishing its single persistent summary; renderers without that barrier still retain command resources. Eleven native and actual Ink regressions have seven causal failures on prior production. Forced lint, typecheck, test, build and isolation pass. Genuine self-SIGINT controls are platform-qualified on Windows. Refs: ADR-0097, ADR-0100 Co-Authored-By: Claude <noreply@anthropic.com>
Record both independently verified Group 3 round 3 findings, causal regressions, native durability and actual Ink summary controls. Clarify the canonical command contract and append the ADR-0097 correction. Fresh complete round 4 remains required; other W7 and PR gates stay open. Refs: ADR-0097, ADR-0100 Co-Authored-By: Claude <noreply@anthropic.com>
Suppress queued cards from emitted rejected budget authority, recheck prompt validity after Ink suspension and abort outstanding Clack input. Keep the primary iterator authoritative for terminal outcome, durable acknowledgement and teardown. A losing idempotent rejection cannot fabricate terminal expectation or block the next genuine gate. Native command/SQLite/Ink controls and actual Clack synthetic-input controls fail against pinned production and pass after correction. Full forced lint/typecheck/test passes 8,286 tests; build, isolation and format pass. Fresh whole-scope independent review remains required. Refs: ADR-0097, ADR-0100 Co-Authored-By: Claude <noreply@anthropic.com>
Record both complete round-four findings, causal native and actual Clack controls, full toolchain results and exhaustive sealed evidence audits. Keep inherited helper/startup/read/platform qualifications and require fresh whole-scope round five before accepting Group 3. Update the canonical CLI contract and append an ADR-0097 correction without changing the accepted decision. W7, Proposed ADR-0102, pending capture work, remaining groups and final PR gates stay open. Refs: ADR-0097, ADR-0100 Co-Authored-By: Claude <noreply@anthropic.com>
Advance W7 and ADR-0097 review findings by awaiting actual command completion, controlling real async cleanup, and specifying the existing safe excluded-entry JSON projection. Co-Authored-By: Claude <noreply@anthropic.com>
Record complete W7 Group 3 round 5 findings and evidence audit; qualify voluntary pause teardown separately from verified cancellation and permit independently reviewed Groups 4–6 to advance. Co-Authored-By: Claude <noreply@anthropic.com>
Record two fresh independent correction rounds and qualified evidence audits for ADR-0097. Keep the voluntary paused-departure High and ADR-0102/0103 approval gates open. Co-Authored-By: Claude <noreply@anthropic.com>
Isolate malformed legacy session effect keys at startup while refusing their reservation, preserve fail-closed aggregate discovery with a shared typed corruption boundary, and reconcile recorded budget rejection accurately. Require keptTurnCount at current projection producers while retaining legacy replay compatibility. Advances W7 review M1/M5/L8/L9/L11/L12/L16 and ADR-0095/0098/0100; does not resolve the Group 3 departure race. Co-Authored-By: Claude <noreply@anthropic.com>
Preserve ADR-0098 legacy identity evidence or its trustworthy floor in the same transaction as committed-row retention. Refuse corrupt sweeps without deleting evidence, and diagnose the blocked allocation with a fixed warning. Address the independently reproduced Group-4 round-1 High. Native activation and chunk rollback controls pass; fresh complete review remains required. Co-Authored-By: Claude <noreply@anthropic.com>
Complete ADR-0100 aggregate corruption refusal before SQLite recovery can misclassify a known suspension and append an interruption terminal. Add native read-surface and engine regressions, preserve matching unknown compatibility, and record Group 4 round-2 evidence and remaining gates. Co-Authored-By: Claude <noreply@anthropic.com>
Close the Group 4 interruption/state-reader projection bypass before any streaming filter. Preserve genuine streaming folds and unknown-event compatibility; pin native refusal, retention and causal controls for ADR-0100. Co-Authored-By: Claude <noreply@anthropic.com>
Record the complete Group 4 fourth review and verified evidence limits for ADR-0095, ADR-0098 and ADR-0100. Correct three literal newline escapes without changing the prior result. Other W7 gates remain open. Co-Authored-By: Claude <noreply@anthropic.com>
Close the systematic review CI, provenance, dependency closure, evidence retention and readiness gaps for ADR-0075 and ADR-0100. Pin the cooled Turbo 2.11.4 toolchain without changing task definitions. Fresh independent Group 5 reviews remain required. Co-Authored-By: Claude <noreply@anthropic.com>
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.




What
Preserve completed session tool structure without storing raw tool payloads, apply one memory/request projection, and make workflow budget approval a durable, bounded authorization. Approval retains its frozen allowance through dispatch and replay; CLI resume uses the same engine protocol.
Draft — W7 is incomplete and must not merge. Published head:
f3dd7246,development→main. Accepted ADR-0095–ADR-0101 govern the implemented work. ADR-0102 remains Proposed and requires explicit maintainer approval. ADR-0103 is an unlanded, unapproved lifecycle draft; neither dependent implementation is authorized.Implementation and open gates
Rendering corrections are accepted after two fresh independent rounds. Group 2a has qualified scoped acceptance after round 14; it does not implement deep request ownership. All six maintainer clarifications to ADR-0102 are committed, with two qualified clean static draft reviews. The ownership High remains open until approval, implementation and its reviews.
Group 3 remains changes requested. The full round-5 review and independent native controls verify that an authored deadline can advance durable state during voluntary paused renderer finalization, allowing a stale paused result or SQLite closure beneath a writer. Two narrow correction rounds are accepted separately and do not close this High. Both seventh ADR-0103 reports and complete Parent audits now verify another static design gap: successful admitted effect dispatch followed by cancellation can bypass settlement and every proposed attention observer. The next unapproved candidate tracks correlated pending obligations through transferred receipt scopes and the final join. Fresh complete draft review remains required; no live scheduling reproduction or approval is claimed.
Group 4 is accepted within scope in
593a9ba3after fresh complete round 4. All 33 corrected paths / 61 hunks and ten full test files were reviewed. The sole validator checks every stored row before either discovery or state-reader streaming exclusion; corruption refuses without invented terminals, changed materialized state, leases or healthy-neighbour changes. Genuine streaming rows remain excluded from the returned fold; their validation increases parsing work. Independent runtime passes 29 files / 753 tests and five fresh builds. Discovery/state-only causal removals fail 43 / 31 controls; exact restoration passes. Both reports and all 3,145 physical entries / 2,781 regular files are Parent-audited, with historical read/capture/preparation qualifications preserved.Group 5 is implemented in
f3dd7246and remains changes requested pending reviews/fixes. Mandatory GitHub CI now runs the offline budget-replay smoke with baseline history available. The actual predecessor's portable byte archive is shipped independently of the live installation; all 143 source files are checked against actual baseline Git blobs. Native SQLite remains the explicit invocation-hashed ABI exception. Completed evidence retention, readiness controls and exact cooled Turbo 2.11.4/local schema are implemented. Initial independent runtime/static reviews verify failure-path gaps in secondary evidence writes, stderr EPIPE, early child-close joins and retirement ownership; fixes and a fresh second review pair remain required. The roadmap's IPC wording also needs correction: isolation uses file-marker readiness; retention children use IPC.Group 6's remaining documentation and individually adjudicated Sonar corrections are prepared and unlanded. No new runtime dependency is introduced. Migration
0017adds the session effect-turn high-water mark. Earlier W6 status reconciliation, ADR-0087 clarification and CR-73/CR-80 fixes remain included.Validation
The committed Group-5 implementation's code candidate passed forced, cache-free
pnpm run ciunderCI=true: 23 lint/typecheck/test tasks, 373 files, 8,397 passing tests / 12 existing skips, six builds, formatting, migration sync, fences, CLI smoke, offline overflow-capture smoke and budget-replay smoke. Subsequent changes before commit were documentation-only and passed formatting, relative file-target and diff checks. Initial tool-lint and formatting failures are retained and are not described as green runs.typecheck:toolschecks its configured TypeScript programme, not every MJS helper.The actual replay smoke verifies 143 predecessor blobs, seven provenance controls, 17 closure controls and eight retention controls plus seven finalization cases in three result groups. Actual current/legacy/predecessor workers pass with zero provider calls, key resolutions or fetches. Independent reviewers are challenging uncovered error paths; passing smoke does not accept Group 5.
Earlier coverage and remote results are historical, not final-head acceptance. The exact
843f0adbSonar analysis enumerated 179 labels (173 code smells, five vulnerabilities, one bug). Individual remediation, current-head GitHub CI/coverage/Sonar and whole-wave acceptance remain open. No blanket dismissal or suppression is applied.Review records and acceptance
W7 execution plan and live gates.
Systematic intake and review index.
Group-4 round-4 scoped acceptance.
Group-3 full round-5 open High and two-round narrow correction acceptance.
ADR-0102 clarification reviews and Proposed request-ownership decision.
Three-of-five provider artifacts and capture runbook.
Rendering corrective group accepted.
Group 2a bounded correction accepted with ownership exclusions.
Both narrow Group-3 correction rounds completed; whole-group High remains open.
Group 4 accepted after complete fresh round 4.
Approve ADR-0102, implement ownership and complete reviews.
Finish ADR-0103 draft reviews, obtain approval, implement lifecycle fixes and complete reviews.
Fix and review Group 5; complete Group 6.
Obtain separately authorized missing provider evidence; implement/review Steps 7–8.
Verify final-head required CI, coverage, replay and individually resolved Sonar gate.
Reconcile canonical documents/registers and complete Step 12.
Strict TypeScript, engine purity, vendor seam, no-new-runtime-dependency and append-only accepted ADR rules remain in force. Keys stay in the OS keychain; arbitrary persisted history/event content is not guaranteed secret-free. Merge readiness is not claimed.