Repository navigation
metal + opt: a lookup that finds no buffer fails the graph, and a failed graph fails the training job - #42
Conversation
…led graph fails the training job The Metal OUT_PROD bug (#41) corrupted every LoRA backward on Apple silicon for ten days while saying so in one log line per op ("tensor '' buffer is nil") that nobody read: a lookup that finds no buffer hands the kernel a nil buffer, the op does nothing it was meant to, and the graph carries on with stale memory. A warning that names a corrupted kernel input must not be a warning (BigMama on #41). - ggml-metal: every such lookup is counted process-wide; ggml_metal_graph_compute returns GGML_STATUS_FAILED for a graph during whose encode the count rose - ggml-opt: ggml_opt_eval no longer ignores the compute's status; a failure becomes the context's refusal (the reason #35 already carries to llama_opt_failure) - llama: a failed eval stops the epoch through the same path as a refused graph, so the server reports the reason and writes no adapter (serving is unaffected) Known-positive on the M5 (Qwen3.5-0.8B, Metal): with #41's fix reverted, /train ends in error "the backend failed to compute the training graph (GGML status: error (operation failed)) ... nothing was written", no adapter file, /health ok; with the fix, the run completes and writes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo
|
Training side reviewed:
Two smaller points:
|
|
One change before approval: the nil-lookup counter is process-wide, but the graphs it fails are not. In-engine training runs in the same process as serving. A training graph and a decode graph can encode at the same time on different contexts. So a nil lookup during the training encode raises the count while a decode graph is also between its before and after reads, and that decode is failed for training's fault. The reverse can happen too, and a failed decode is a failed turn for her. That is the one thing this whole lane exists to protect. Suggested fix: count per graph compute, not per process. The ops encoding the graph hold the A question: this also makes every Metal INFERENCE graph fail on a nil lookup, where before it logged and went on. That is the right rule. Before it ships, though, can someone run a Metal serving smoke (a few turns, plus a vision turn if the lane has one) and show zero The |
… never process-wide (Cormac on #42) A process-wide counter let a training graph's nil lookup fail a decode graph encoding beside it, which costs her a turn for a fault that was not hers. Each Metal context now owns its count: every encode block (the main thread's and the dispatched ones) points a thread-local sink at its context's counter while it encodes and clears it after, a nil lookup increments the current sink, and ggml_metal_graph_compute zeroes its own count on entry and fails only its own graph. Measured on the M5 (Qwen3.5-0.8B, Metal, 2 slots decoding): - serving only, fix in: 29 decode requests ok, 0 failed, 0 "buffer is nil" - #41's fix reverted, training beside decoding: the training job fails on its first graph ("found no buffer while encoding this graph"), the decode requests in flight with it all succeed (2 ok, 0 failed: a thin sample, one failing graph) - fix in, training beside decoding: training done, 216 decode requests ok, 0 failed, 0 "buffer is nil" Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo
|
35bd266: counted per graph compute (thread-local sink set by each encode block to its context's counter; graph_compute zeroes and reads its own). Receipts in the commit: serving smoke 0 nil / 0 failures; bugged training fails only itself while concurrent decodes succeed (2/2, thin sample); fixed training done beside 216 ok decodes. |
|
APPROVED at 35bd266. The nil-lookup count is per graph compute: each encode block points a thread-local sink at its own context counter, so a training graph cannot fail a concurrent decode. Receipts: serving smoke 0 nil and 29/29 decodes; bugged training fails only itself; fixed training completes beside 216 ok decodes with 0 nil. A lookup made outside an encode block lands in no sink and stays only a log line. That is fine today, since every get_id I can see runs inside an encode, but worth a comment where the sink is cleared. |
…ure path (returns false, like a refused graph) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q4NU4VNiELPQfBpCacDZGc
The Metal OUT_PROD bug (#41) corrupted every LoRA backward on Apple silicon for ten days while
saying so in one log line per op ("tensor '' buffer is nil") that nobody read: a lookup that
finds no buffer hands the kernel a nil buffer, the op does nothing it was meant to, and the
graph carries on with stale memory. A warning that names a corrupted kernel input must not be
a warning (BigMama on #41).
GGML_STATUS_FAILED for a graph during whose encode the count rose
context's refusal (the reason train: a node the device cannot run refuses the JOB, not the server; supports_op tells the truth #35 already carries to llama_opt_failure)
server reports the reason and writes no adapter (serving is unaffected)
Known-positive on the M5 (Qwen3.5-0.8B, Metal): with #41's fix reverted, /train ends in error
"the backend failed to compute the training graph (GGML status: error (operation failed)) ...
nothing was written", no adapter file, /health ok; with the fix, the run completes and writes.
Stacked on #41 (retarget to the fork base once #41 merges).
🤖 Generated with Claude Code
https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo