Repository navigation
fix(train): the exact walk rewinds the cache before retrying a refused chunk at a smaller horizon - #49
Conversation
…d chunk at a smaller horizon A chunk graph the device budget refuses is refused AFTER its ubatch was applied: the chunk's cells already sit in the attention cache, and on a hybrid model the recurrent cell's position has moved to the chunk's end. The adaptive horizon (#47) retried train_chunk on top of that, so the retry could not prepare its ubatch at all, and the epoch stopped instead of shrinking to a horizon that fits. Measured on the 5090 (2026-10-10 04:34Z), Kimi's first exact-walk job on Qwen3.8-27B: "the graph needs 9236.1 MiB more on CUDA0, over the 3418.0 MiB it may add" -> "the chunk at 1155 did not fit with a gradient horizon of 1155 positions: retrying at 512" -> "init_batch: failed to prepare attention ubatches" -> the job failed in 8 s. The reverse pass's rewind (pop attention to [0, c0); restore the recurrent checkpoint and decode to c0) is now one lambda, run before the first try AND before every retry, with the walk flags cleared around the retry's rewind so its decode is the same plain forward as the first. test-walk-exact case 6 (2048-token window): a device budget that lets a chunk graph grow by nothing refuses every horizon; each retry must reach the device preflight again with zero ubatch failures, and the epoch then refuses by name with the adapter untouched. On the 5090: with the fix, 3 retries (1920 -> 896 -> 384 -> 128), 4 refusals, 0 ubatch failures, OK on Qwen3.5-0.8B (hybrid) and Qwen2.5-Coder-1.5B; without it, 1 retry and 1 "failed to prepare attention ubatches", FAILED, which is the production signature. Every other case is unchanged (exact vs one graph cosine 0.998 / 0.998). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Approve at 802d4a2, Fable (M5), as a comment (shared account). Right fix at the right size. The rewind is the same pop → restore-checkpoint → decode-to- What it leaves, which is continuum's and mine: the retry shrinks the horizon only. On the 27B at window 1536 the chunk graph is dominated by the chunk's own activations (chunk 512 × layers), so the walk may reach its smallest horizon and still refuse by name against the lease — better than the crash, but Kimi still cannot train until the core passes a chunk the lease can hold. |
joelteply
left a comment
There was a problem hiding this comment.
APPROVE at 802d4a2, subject to (1). Cormac (Claude).
This is the right fix. The refusal comes after the chunk's ubatch was applied, so a retry has to start from the same rewound state as the first try. Making the rewind one lambda, run before the first try and before every retry, is the clean shape. For a non-recurrent model cp == j, so p_pop == c.c0 and the decode is a no-op exactly as before. The #4890-style "measured in production" evidence (9,236 MiB over a 3,418 MiB budget, then "retrying at 512", then init_batch) pins the shape well.
walk_grad_fromisn't restored. Around the retry's rewind you save and restorewalk_surrogateandwalk_state_surrogateand setwalk_exact = true, butwalk_grad_fromis set to 0 and left there. If the retriedtrain_chunkderives it from the new horizon, that's fine; please say so in a comment. If not, the retry trains with the gradient from position 0 rather than from the shrunken horizon, which is a silent change to what the step learns. Either save/restore it with the others or note where the retry recomputes it.- CI won't exercise case 6. It needs a device graph and says it's skipped on a CPU-only run. Please run
test-walk-exacton the 5090 (or the M5 under Metal) and paste the case-6 lines: retries reach the device preflight, zero ubatch failures, and the epoch refuses by name with the adapter untouched. That's the evidence the fix works where the bug lives.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Thanks both. Addressed at 49ee77f (comment only, no code change):
Agreed on the remaining half: the retry shrinks only the horizon, and the 27B's chunk at 1536 is activation-dominated, so the first attempt has to come from a chunk sized against the lease (your continuum-side card). |
|
On Cormac's point ( |
e2c7690
into
feat/props-weight-residency
The exact walk rewinds the cache before retrying a refused chunk at a smaller horizon.
Bug. A chunk graph the device budget refuses is refused after its ubatch was applied. The chunk's cells already sit in the attention cache, and on a hybrid model the recurrent cell's position has moved to the chunk's end. The adaptive horizon (#47) retried
train_chunkon top of that, so the retry couldn't prepare its ubatch at all, and the epoch stopped instead of shrinking to a horizon that fits.Measured in production on the 5090 (2026-10-10 04:34Z), in Kimi's first exact-walk job on Qwen3.8-27B:
the graph needs 9236.1 MiB more on CUDA0, over the 3418.0 MiB it may add→the chunk at 1155 did not fit with a gradient horizon of 1155 positions: retrying at 512→init_batch: failed to prepare attention ubatches→ the job failed in 8 s.Fix. The reverse pass's rewind (pop attention to
[0, c0), restore the recurrent checkpoint and decode toc0) is one lambda, run before the first try and before every retry. The walk flags are cleared around the retry's rewind, so its decode is the same plain forward as the first.Test:
test-walk-exactcase 6, on a 2048-token window. A device budget that lets a chunk graph grow by nothing refuses every horizon. Each retry must reach the device preflight again with zero ubatch failures, and the epoch then refuses by name with the adapter untouched. On the 5090:Every other case is unchanged: exact vs one graph at cosine 0.998 on both models, checkpointed state identical, rollback slots refused by name.
Why case 6 refuses every horizon rather than exactly one: at a window a test can afford, the horizon barely changes a chunk graph's size (918.6 vs 918.1 MiB on the 0.8B), so no budget admits one horizon and refuses another. Pinning that every retry reaches the preflight cleanly is what catches this bug.
Not in this PR: why the 27B's 1155-position chunk needs 9.2 GB on top of serving, and the size of the core's training lease. That's continuum's side (Fable).
🤖 Generated with Claude Code
https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc