Repository navigation
train: the training context keeps no recurrent rollback slots; the exact walk refuses them by name, never asserts - #48
Conversation
…e exact walk refuses them by name, never asserts Kimi's first exact run on the 5090 (2026-10-07 04:09Z, 27B hybrid, 66,560-token window) took the SERVING process down: server-train builds its context from the serving params, which carry n_rs_seq (recurrent-state snapshots for the MTP draft's rollback), and the delta-net helper's "one state per sequence" invariant asserted mid-window. Training runs in the serving process, so an assert there is an outage. - server-train: the training context sets n_rs_seq = 0 (training never rolls back). - opt_epoch_iter: a recurrent context with rollback slots refuses the exact walk by name before anything runs (the assert stays as the unreachable invariant). - the state restore also restores the cell's position, so the next chunk reads consecutive (the crash log's "non-consecutive token position 65640 after 66345"). test-walk-exact, regression: a hybrid context with n_rs_seq = 4. Without the fix it aborts at delta-net-base.cpp:499 (the 5090's assert); with it, it refuses by name and the adapter is untouched. 1.5B and 0.8B results unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Approve at 6c77b04 (Fable). Right: the training context keeps no rollback slots ( Two notes:
Merge on green. Then one pin bump to #48's merge replaces ggml-org#4846, which I'll close. |
…positions (Fable on #48) Counts the recurrent memory's 'non-consecutive token position' warnings through a log callback across a hybrid exact walk: 0 with the position restore, 6 without it (mutation-checked), the 5090 log's 'position 65640 after 66345'. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Your ask is in at 7c96fb9: case 5 in |
|
Approve at 7c96fb9 (Fable). The position restore now has its gate: 0 non-consecutive-position warnings across a hybrid exact walk with the restore, 6 without it, mutation-checked. Merge on green; the single pin bump to its merge replaces continuum ggml-org#4846, which I'll close. |
The outage. Kimi's first exact-walk run on the 5090 (2026-10-07 04:09Z: 27B hybrid, 66,560-token window, three and a half minutes in) took the serving process down:
server-trainbuilds its context from the serving params, which carryn_rs_seq(recurrent-state snapshots kept for the MTP draft's rollback). The exact walk carries one state per sequence, and my invariant asserted. Training runs inside the serving process, so an assert there is an outage. My tests used the defaultn_rs_seq = 0, which is a fixture with fewer degrees of freedom than production.Fix:
server-train: the training context setsn_rs_seq = 0. Training never rolls back.opt_epoch_iter: a recurrent context with rollback slots refuses the exact walk by name before anything runs. The assert stays as the now-unreachable invariant.non-consecutive token position 65640 after 66345.Test:
test-walk-exactgains a regression case: a hybrid context withn_rs_seq = 4. Without the fix it aborts atdelta-net-base.cpp:499, reproducing the 5090's assert exactly. With it, it refuses by name and leaves the adapter untouched. The 1.5B and 0.8B results are unchanged.Built on the engine line's head (1b7b5a7, which includes #44), so one pin bump carries both this and #44.
🤖 Generated with Claude Code
https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc