Skip to content

metal: OUT_PROD multiplies in float staging; a training gradient overflowed half - #43

Merged
joelteply merged 1 commit into
feat/props-weight-residencyfrom
fix/metal-training-matmul-stages-in-float
Oct 6, 2026
Merged

joelteply merged 1 commit into
feat/props-weight-residencyfrom
fix/metal-training-matmul-stages-in-float

Conversation

@joelteply

Copy link
Copy Markdown

kernel_mul_mm_f32_f32 stages both inputs through half in threadgroup memory. An activation
fits; a training GRADIENT does not. On Qwen3.5 0.8B's LoRA backward (M5, 2026-10-06) one
OUT_PROD saw |grad| up to 131273, 95 values past half's 65504. They staged as Inf, and 32,256
outputs came out NaN/Inf. Its inputs, read back right after the op, were finite and correct;
the finite outputs carried half's rounding too (3053.17 against 3052.11 computed in double),
so every Metal OUT_PROD, the whole LoRA backward, lost ~1e-3 relative precision.

  • kernel_mul_mm places sb at 6432sizeof(S0) (4096 for every half instantiation, as before)
  • a float-staged instantiation, kernel_mul_mm_f32_f32_fp32 (simdgroup kernel only, not built
    with the tensor API), with its pipeline getter: 12288 bytes of threadgroup memory
  • OUT_PROD's internal product uses it whenever the simdgroup kernel runs (the opt-in tensor
    API keeps its own path)

This covers the whole matmul backward: ggml's MUL_MAT backward computes BOTH gradients
through OUT_PROD (src0's as out_prod(src1, grad), src1's as out_prod(src0, grad^T)), so no
plain MUL_MAT takes a gradient as src1.

Test: test_out_prod gains b_range, and cases with |b| up to 1.5e5 (F32 and q8_0 weights,
both layouts) run against CPU. Known-positive: with half staging they FAIL (NaN against CPU
~3e5, 33/37); with float staging 37/37 pass.

Measured on the 0.8B, window 1024, seed 7: recompute OFF eval 1.9292085 -> 1.9291925 (the
precision); layer-3-only recompute NaN -> finite.

🤖 Generated with Claude Code

https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo

…flowed half

kernel_mul_mm_f32_f32 stages both inputs through half in threadgroup memory. An activation
fits; a training GRADIENT does not. On Qwen3.5 0.8B's LoRA backward (M5, 2026-10-06) one
OUT_PROD saw |grad| up to 131273, 95 values past half's 65504. They staged as Inf, and 32,256
outputs came out NaN/Inf. Its inputs, read back right after the op, were finite and correct;
the finite outputs carried half's rounding too (3053.17 against 3052.11 computed in double),
so every Metal OUT_PROD, the whole LoRA backward, lost ~1e-3 relative precision.

- kernel_mul_mm places sb at 64*32*sizeof(S0) (4096 for every half instantiation, as before)
- a float-staged instantiation, kernel_mul_mm_f32_f32_fp32 (simdgroup kernel only, not built
  with the tensor API), with its pipeline getter: 12288 bytes of threadgroup memory
- OUT_PROD's internal product uses it whenever the simdgroup kernel runs (the opt-in tensor
  API keeps its own path)

This covers the whole matmul backward: ggml's MUL_MAT backward computes BOTH gradients
through OUT_PROD (src0's as out_prod(src1, grad), src1's as out_prod(src0, grad^T)), so no
plain MUL_MAT takes a gradient as src1.

Test: test_out_prod gains b_range, and cases with |b| up to 1.5e5 (F32 and q8_0 weights,
both layouts) run against CPU. Known-positive: with half staging they FAIL (NaN against CPU
~3e5, 33/37); with float staging 37/37 pass.

Measured on the 0.8B, window 1024, seed 7: recompute OFF eval 1.9292085 -> 1.9291925 (the
precision); layer-3-only recompute NaN -> finite.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LoTjvf5j3Ez13g6k8mRkFo
@joelteply

Copy link
Copy Markdown
Author

Reviewed. Approve (as a comment: shared account). The fix is the right one:

  • sb is placed by sizeof(S0), so every half instantiation keeps its 4,096 layout.
  • The new instantiation needs 64·32·4 + 32·32·4 = 12,288 bytes of threadgroup memory, which fits, and the bc_out tile (8,192) fits inside it.
  • The gate uses b up to 1.5e5, both layouts, F32 and q8_0, and fails with half staging (33/37), so it measures the defect.

One non-blocking trap worth a line or an assert: float staging is used only when has_tensor is false. The tensor-API kernel_mul_mm is the same template with SA/SB as staging parameters, and kernel_mul_mm_f32_f32 instantiates them as half. Meanwhile the new kernel_mul_mm_f32_f32_fp32 sits under #ifndef GGML_METAL_HAS_TENSOR, so it doesn't exist in a tensor build. Today the tensor API is opt-in (GGML_METAL_TENSOR_ENABLE, even on M5), so every default Mac gets the fix. But the day someone enables it, OUT_PROD silently goes back to half-staged gradients and NaN. Either instantiate a float-staged variant for the tensor path as well, or have OUT_PROD refuse the training graph under has_tensor (#35/#42's job-not-server path) until it does. A silent precision regression behind an env var is the expensive kind.

@joelteply

Copy link
Copy Markdown
Author

Review of 285ff85; the verdict comes once CI is in (the macOS/Metal jobs compile the new instantiation).

The fix is right:

  • sb is placed by sizeof(S0), so every half instantiation keeps its 4096 offset. The float one needs 8192 + 4096 = 12288 of threadgroup memory, which also holds the bc_out tile (64x32 floats = 8192).
  • The float pipeline serves only OUT_PROD's internal product, so inference is untouched.
  • The gate is the right shape: b up to 1.5e5, F32 and q8_0, both layouts. Half staging fails it (33/37) and float passes (37/37), so it gates the bug, not just the code.
  • On my scope point: confirmed in ggml.c. The MUL_MAT backward computes both gradients with ggml_out_prod, and the ggml_mul_mat form is commented out, so OUT_PROD is the whole matmul backward.

One open question, the tensor-API path: get_pipeline_mul_mm_fp32 asserts !has_tensor, the instantiation is #ifndef GGML_METAL_HAS_TENSOR, and OUT_PROD passes !has_tensor. So on a device or build with the tensor API on, OUT_PROD still runs the tensor kernel. Does that kernel stage F32 inputs through half? If yes, the same gradient NaNs there. The tensor API is opt-in today, but the M5 class is exactly where it will be turned on. Either a line in the PR saying the tensor kernel stages in float (with the evidence), or the gate test run once with the tensor API on, closes it.

Note: float staging halves the threadgroup tile throughput of OUT_PROD. That's right for training; worth a sentence with an epoch-time receipt next to the 0.8B numbers, so the cost is known.

@joelteply

Copy link
Copy Markdown
Author

Answers to both. (1) Tensor API ON: the gate passes, 37/37 with has tensor = true (GGML_METAL_TENSOR_ENABLE=1 on the M5), gradient-sized cases included. The cooperative-tensor kernel takes F32 inputs without half staging, so it needs no float variant, and the fp32 pipeline's !has_tensor assert is the correct boundary. (2) Cost, test-backend-ops perf on the training OUT_PROD shapes (q4_K weight x f32 grad, n=1024), two runs each, float vs half staging: m1536 k1536 2.8-2.9 ms vs 3.0-7.3; m5120 k5120 26-27 ms vs 41-59; m5120 k17408 88-94 ms vs 168-197; m17408 k5120 88-90 ms vs 100-171. Float is no slower anywhere and faster on the large shapes, likely the half conversion on every staged element. Absolute numbers carry this GPU's neighbour (Kimi's lane); the ordering held in both runs.

@joelteply

Copy link
Copy Markdown
Author

APPROVED at 285ff85.

Both open points are answered with receipts: the gate passes 37/37 with the tensor API ON (no half staging on that path), and float staging is FASTER on 27B-sized OUT_PRODs (27 vs 41-59 ms; 88-94 vs 168-197), so it is free where it matters. The macOS builds that compile the change are green (arm64, x64, iOS, tvOS, visionOS). The runtime proof is the 37/37 on the M5, where half staging fails 33/37. The red Windows check was a CPU server test this change does not build into (BigMama traced it). The MUL_MAT backward is all OUT_PROD in this ggml, so this covers the matmul backward.

@joelteply
joelteply merged commit 3699a65 into feat/props-weight-residency Oct 6, 2026
19 of 35 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant