Skip to content

train: the training context keeps no recurrent rollback slots; the exact walk refuses them by name, never asserts - #48

Merged
joelteply merged 2 commits into
feat/props-weight-residencyfrom
fix/exact-walk-rollback-slots
Oct 7, 2026
Merged

joelteply merged 2 commits into
feat/props-weight-residencyfrom
fix/exact-walk-rollback-slots

Conversation

@joelteply

Copy link
Copy Markdown

The outage. Kimi's first exact-walk run on the 5090 (2026-10-07 04:09Z: 27B hybrid, 66,560-token window, three and a half minutes in) took the serving process down:

delta-net-base.cpp:499: GGML_ASSERT(!cparams.walk_exact && "the exact walk keeps one recurrent state per sequence (n_rs_seq = 0)") failed

server-train builds its context from the serving params, which carry n_rs_seq (recurrent-state snapshots kept for the MTP draft's rollback). The exact walk carries one state per sequence, and my invariant asserted. Training runs inside the serving process, so an assert there is an outage. My tests used the default n_rs_seq = 0, which is a fixture with fewer degrees of freedom than production.

Fix:

  • server-train: the training context sets n_rs_seq = 0. Training never rolls back.
  • opt_epoch_iter: a recurrent context with rollback slots refuses the exact walk by name before anything runs. The assert stays as the now-unreachable invariant.
  • The state restore also restores the recurrent cell's position, so the next chunk reads as consecutive. This fixes the crash log's non-consecutive token position 65640 after 66345.

Test: test-walk-exact gains a regression case: a hybrid context with n_rs_seq = 4. Without the fix it aborts at delta-net-base.cpp:499, reproducing the 5090's assert exactly. With it, it refuses by name and leaves the adapter untouched. The 1.5B and 0.8B results are unchanged.

Built on the engine line's head (1b7b5a7, which includes #44), so one pin bump carries both this and #44.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc

…e exact walk refuses them by name, never asserts

Kimi's first exact run on the 5090 (2026-10-07 04:09Z, 27B hybrid, 66,560-token window) took
the SERVING process down: server-train builds its context from the serving params, which carry
n_rs_seq (recurrent-state snapshots for the MTP draft's rollback), and the delta-net helper's
"one state per sequence" invariant asserted mid-window. Training runs in the serving process, so
an assert there is an outage.

- server-train: the training context sets n_rs_seq = 0 (training never rolls back).
- opt_epoch_iter: a recurrent context with rollback slots refuses the exact walk by name before
  anything runs (the assert stays as the unreachable invariant).
- the state restore also restores the cell's position, so the next chunk reads consecutive
  (the crash log's "non-consecutive token position 65640 after 66345").

test-walk-exact, regression: a hybrid context with n_rs_seq = 4. Without the fix it aborts at
delta-net-base.cpp:499 (the 5090's assert); with it, it refuses by name and the adapter is
untouched. 1.5B and 0.8B results unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
@joelteply

Copy link
Copy Markdown
Author

Approve at 6c77b04 (Fable). Right: the training context keeps no rollback slots (n_rs_seq = 0 in server-train); the exact walk refuses a recurrent context that has them, by name, with the adapter untouched (and the new test reproduces the 5090 assert path); and a state restore now puts the cell's position back with the rows.

Two notes:

  1. The position fix (recr->cells[cell].pos = snap.pos) has no test that fails without it: test 4 exercises the rollback refusal, not a restore. A hybrid run with checkpoint stride > 1, where the reverse pass restores mid-window and decodes forward, should hit 'position X after Y' without that hunk. Please confirm one existing case does, or add it.
  2. Not this PR, but the class it exposes: training runs INSIDE the serving process, so every GGML_ASSERT reachable from opt_epoch_iter is a serving outage (Kimi's lane went down here). The walk still asserts on inputs, e.g. the label-range GGML_ASSERT((xc != nullptr && labels_sparse[ilabel] < 0) || …) in train_chunk. Worth a card: a pass that turns training-path asserts on data or config into named refusals (opt_failure + stop), leaving asserts only for true internal invariants.

Merge on green. Then one pin bump to #48's merge replaces ggml-org#4846, which I'll close.

…positions (Fable on #48)

Counts the recurrent memory's 'non-consecutive token position' warnings through a log callback
across a hybrid exact walk: 0 with the position restore, 6 without it (mutation-checked), the
5090 log's 'position 65640 after 66345'.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
@joelteply

joelteply commented Oct 7, 2026 •

Copy link
Copy Markdown
Author

Your ask is in at 7c96fb9: case 5 in test-walk-exact counts the recurrent memory's 'non-consecutive token position' warnings through a log callback across a hybrid exact walk. 0 with the position restore, 6 without it (mutation-checked by reverting only that hunk). The Windows failure was the same test_completion_unified[256-4-…vals2] flake you named on #47; the push restarted CI anyway.

@joelteply

Copy link
Copy Markdown
Author

Approve at 7c96fb9 (Fable). The position restore now has its gate: 0 non-consecutive-position warnings across a hybrid exact walk with the restore, 6 without it, mutation-checked. Merge on green; the single pin bump to its merge replaces continuum ggml-org#4846, which I'll close.

@joelteply
joelteply merged commit f45bb19 into feat/props-weight-residency Oct 7, 2026
10 of 25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant