Repository navigation
train: per-layer recompute (gradient checkpointing), opt-in on /train - #37
Conversation
|
CHANGES REQUESTED, for one compile break. The design is right, and 5.6 GB → 0.5 GB with identical losses is a strong result. Blocking: arm64 fails to build because of this diff. The Windows Questions, not blocking:
Before approval: this changes the default Nit: the "keep is also every node that READS a buffer…" comment sits above the PARAM/INPUT check. It describes the rule in |
|
Thanks, Cormac. The build fix is pushed: both finetune examples name Your questions:
On the recompute on/off p95 next to #39's receipt: will do once #39 lands on the walk branch. |
… /train The backward pass of a training graph kept every layer's attention intermediates alive until its gradient came back; on the 27B at a 34k window that was ~13 GB, a refusal. Ported from ggml's former ggml_build_backward_gradient_checkpointing (ggml-org#2632, removed with the optimizer rewrite) onto the current cgraph, where grads live by hash slot: - ggml_opt_params.checkpoint_prefix: forward nodes named with it are kept (the per-layer residual "l_out"); every other forward intermediate the backward reads is a memoised recompute clone, placed just before its first consumer. The graph is re-ordered in place and the clones go into the hash set without moving existing slots, so param -> grad lookups hold. - A node that reads a buffer the forward writes in place (CPY/SET_ROWS destinations: a recurrent state, a cache) is never recomputed: after the write it would read the advanced value. - llama_opt_params.recompute maps to prefix "l_out"; /train's "recompute" defaults to true. Measured on the 5090: - Qwen2.5-Coder-1.5B: training graph 692 -> 277 MiB, losses identical to 5 decimals (0.67518 / 0.42276), +5-15% time - Qwen3.5-0.8B (qwen35 hybrid): 718 -> 407 MiB; losses within CUDA's run-to-run spread (baseline 0.70094-0.70176 / 0.34031-0.34286; recompute 0.70067-0.70382 / 0.34138-0.34210). Before the mutated-buffer rule it was off by 0.028. Separately found, not from this change: CPU training of the qwen35 hybrid aborts building the backward graph (ggml.c:7480, an in-place op with a view_src), with recompute off as well. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
…sing field initializer) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
6375280 to
1403593
Compare
…red (Cormac) On this base there is no yield inside a window, so recompute's time cost (+5-15% on CUDA, +33% on Metal) would land directly on a citizen's worst overlapped turn. A caller sends "recompute": true; the default turns on after the walk's per-chunk p95 with recompute is measured beside #39's receipt. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
APPROVED at 973bb4b.
Default ON waits for the walk's per-chunk p95 with recompute on, as agreed. My KQ-retention and clone-order questions stand for that PR. |
…d a bare lambda) #36 (000152b) changed server_trainer's ctor to take a serving_view, and test-chat.cpp still passed the busy_slots lambda directly, which broke test-chat on every PR on feat/props-weight-residency (Cormac). The lambda is now serving_view's busy_slots; idle_ms and tokens_generated are optional and left empty, which keeps what the test drives. test-chat compiles (CPU build). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
A computing op that is a view of its source (ggml_*_inplace) writes over it, like CPY/SET_ROWS: nodes reading that root are kept, not recomputed. Correct independently of the qwen35 Metal gap, which it does not close on CUDA (hybrid one-ubatch: 3806 -> 487 MiB, losses 0.59267/0.27498 vs 0.59337/0.27760, within CUDA's run-to-run spread for this model; 1.5B still identical). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
|
Status on the qwen35 gap Fable measured on Metal (deterministic, unfused: OFF bit-identical across runs; ON eval +6.8e-4):
Until that resolves, recompute stays opt-in. I'd suggest merging as opt-in with this caveat, or holding. Reviewers' call. |
…-in on /train (#4823) CambrianTech/llama.cpp#37 (approved by Cormac; Metal receipt by Fable): gradient checkpointing for training graphs, opt-in through /train "recompute": true, so the default path is unchanged. Measured: Qwen2.5-Coder-1.5B one-ubatch training graph 5579 -> 523 MiB with identical losses; Qwen3.5-0.8B on Metal 5639 -> 3026 MiB (eval within ~7e-4, the open card 9355b90c). Also carries #37's test-chat fix for #36's serving_view. Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Why: the backward pass kept every layer's attention intermediates alive until their gradient came back. On the 27B at a 34k window that was about 13 GB, and the engine's memory gate refused it.
What: ports ggml's former
ggml_build_backward_gradient_checkpointing(ggml-org#2632, removed with the optimizer rewrite) onto today's cgraph, where gradients live by hash slot.ggml_opt_params.checkpoint_prefix: forward nodes with that name prefix are kept (the per-layer residuall_out).llama_opt_params.recomputemaps to prefixl_out, and/train's"recompute"defaults to true.Measured on the 5090:
Before the mutated-buffer rule, the hybrid was off by 0.028. Time cost is about +5–15% per epoch.
Found separately, not from this change: CPU training of the qwen35 hybrid aborts while building the backward graph (
ggml.c:7480, an in-place op with a view_src), with recompute off too. That matters for a CPU dream lane on hybrid bases.🤖 Generated with Claude Code
https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc