perf: keep the value stack's height in a register across tail-call handlers - #79
Open
matthargett wants to merge 4 commits into
Open
matthargett wants to merge 4 commits into
matthargett wants to merge 4 commits into
Conversation
matthargett
force-pushed
the
perf/stack-height-arg
branch
from
September 30, 2026 21:32
7bb7a50 to
29679a7
Compare
explodingcamera
self-requested a review
October 1, 2026 16:09
explodingcamera
added a commit
that referenced
this pull request
Oct 3, 2026
originally reported in #79 Signed-off-by: Henry <mail@henrygressmann.de>
explodingcamera
added a commit
that referenced
this pull request
Oct 3, 2026
based on the fix in #79 Signed-off-by: Henry <mail@henrygressmann.de>
Owner
|
I just pushed a fix for the function call panics from this to next. I do see a good performance improvement overall, and I’d definitely like to bring in more of these changes. I’m just a bit hesitant to add more nightly features right now, so I might shuffle the flags around to separate tail-call dispatch from a more general nightly feature. I’ll probably do some more testing in this direction and pull the changes in incrementally rather than all at once. |
explodingcamera
removed their request for review
October 3, 2026 14:32
On arm64_32 (watchOS) the Rust ABI passes the 8-byte Instruction by reference, so every dispatch stores it and the next handler loads it back. The C convention passes it in a register, and `C-unwind` still lets a host function's panic unwind. Other targets keep the Rust ABI.
The height of a value-stack lane is its own field, over a Vec that holds every slot the lane has reached. Slots above the height hold stale values that are written before they are read again. The tail-call handlers can then pass the height between them in a register and write it back without unsafe code. Entering a function writes its locals, and its operand-stack reservation if that is at most 64 slots, so later entries at that height compare once. A larger reservation stays capacity, and a push writes each slot the first time it reaches it, so a function whose deep branch rarely runs keeps no pages for it. With 64 stores of a module whose untaken branch declares a 30000-deep i64 stack, each store keeps next's footprint (35 KB on macOS, 16-19 KB on glibc) instead of 273 KB and 250 KB. v128.any_true and i8x16.all_true test the vector with integer arithmetic: LLVM scalarized their iterator forms once pushes gained that path. Only builds with debug assertions check indices against the height. CI also runs the tests optimized with them, keeping the release profile's wrapping arithmetic. Entering a function zeroes one or two locals with direct stores, and only the ones in between with memset.
The executor takes the store's value stack when it is created and gives it back when it is dropped, and around host calls, which read their arguments from the store and may call back into wasm. Every handler receives the executor as a `&mut`, so a stack access reached through it needs one load fewer (the store pointer), and the compiler knows that no store to a slot or to memory changes a stack's height, so it can keep the heights in registers within a handler.
…lers Each tail-call handler receives the height of the 32-bit value stack, which most instructions push to or pop from, as its last argument. It writes the height to the stack before its body and reads it back before dispatching, and with the stack in the executor the compiler can keep it in a register in between, as it does in most handlers. A handler no longer loads the height its predecessor stored, a store-to-load round trip through memory on every dispatch. The Unbudgeted handlers now take six integer arguments and the Bounded ones five, which stay in registers on arm64 (eight) and x86-64 System V (six). Windows x64 passes four, so there the height goes on the stack along with the Instruction, which already does without this change.
matthargett
force-pushed
the
perf/stack-height-arg
branch
from
October 4, 2026 01:23
29679a7 to
fc1f817
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The stack-height half of #78, which I separate into four commits that can I measured one by one. None of them adds a nightly feature, and each one builds and passes the tests on its own, so they can go in one at a time. The panic fix that was the fifth commit is on
nextas 9d810e1.Instructionin a register on arm64_32 (see below). Only what the convention passes in registers stays out of memory between two handlers.Vec. A lane's height is its own field over aVecthat only grows. Slots above the height are stale and written before they are read. This lets the handlers write the height back withoutunsafe. A function's operand-stack reservation is written along with its locals only when it's at most 64 slots. A larger one stays capacity, and each slot is written the first time a push reaches it, so a function whose deep branch rarely runs keeps no pages for it: 64 stores of a module whose untaken branch declares a 30,000-deep i64 stack keepnext's footprint (35 KB each on macOS, 16–19 KB on glibc) instead of 273 KB and 250 KB. Only builds with debug assertions check indices against the height, and a new CI job runs the tests optimized with them. A push that reaches a new slot checks the capacity and then usesVec::push, which makes no call; the nightlypush_within_capacitywould save about 1% of instructions per call on arm64.tests/host_nested_panic.rs). Every handler receives the executor as a&mut, so a stack access needs one fewer load, and the compiler knows that no store to a slot or to memory changes a height (verified this in disassembly and with low-level CPU counters).Apple devices
Change against
next(9d810e1) withnightly-tail-calls, median of five interleaved launches per build. Energy is the kernel's estimate of the CPU energy each call used (task_power_info_v2). Commit 1 does not change the code on 64-bit arm64.The iPhone 12 and the Apple TV 4K weren't available for this version. On the earlier version (an older
next, withpush_within_capacity, built for the A10), all four commits were −5.2% cycles on the iPhone 12 (A14) and −8.0% on the Apple TV 4K (A10X). On the A10X (Apple TV 4K first gen) the kernel doesn't reports energy or which cores ran the benchmark, which makes sense since it plugs into a wall :D I have an A10X iPad Pro, but that device didn't get updated past iPadOS 17 so the performance profiling is limited in different ways. So far I've found that what's good for A12 effiency cores maps well to A10X's high-performance (Hurricane) cores in terms of wallclock times.On Apple Watch SE (S8, arm64_32), 13 benchmarks (the watch leaves out audio DSP and the GC trees due to peak memory usage constraints, something to optimize for later), measured on the earlier version; commit 1 hasn't changed since:
extern "C-unwind")Between some launches the watch changed its performance state, which moves energy per cycle between about 46 and 73 pJ, too often to give commit 1's energy on its own. Profiling on the physical watch is extremely annoying, and very touchy (literally!), but I feel like I got enough on-device data for it to be defensible.
All four commits against
next, per benchmark:Commit 4 pays where the height's store-to-load round trip costs: on the iPhone 12 it removed 88–97% of the memory-order flushes on every benchmark profiled, and on audio DSP back-end stalls fell from 14.5% to 3.0% of pipeline slots. On the iPhone SE and XS Max it's neutral to about 2% slower, while commits 1–3 save 6.3% and 4.6% there, so commits 1–3 also stand on their own.
The call-heavy rows are bound by the call path instead: on the iPhone 12's tail-call FSM, about a third of the samples were the
Shared<WasmFunction>refcount updates on each call and return, which #74 removes.With the default dispatch, the series no longer changes instructions per call (within 0.05% on all 15 benchmarks):
next's own match loop got 1.8–5.7% cheaper since 017780e, which is what commits 2 and 3 used to give it.Calling convention (commit 1)
Instructionon the stack today, and with commit 4 the height too.extern "rust-preserve-none"would pass twelve on every OS, but it's a nightly feature (rust_preserve_none_cc), so it isn't in this PR.extern "C-unwind". The Rust ABI passes the 8-byteInstructionby reference there, so onnextevery dispatch stores it and the next handler loads it back. The C convention passes it in a register, andC-unwindstill lets a host function's panic unwind (super cool!)Hot path per dispatch, weighted by xmrsplayer's dispatch mix, from the compiled handlers (instructions / stack operands / pushes and pops):
nextHandlers that push or pop on the hot path: 258 → 273 on System V and 547 → 614 on Windows. Without
rust-preserve-none, commit 4 costs on x64: the height takes a register the handler bodies would otherwise use, and on Windows it goes on the stack. These are counts from the compiled handlers; they have not been timed on x64. I have a Surface Book 2 here I can resurrect and test on if need be, but I'd love to outsource to other people's deployment targets that I surely don't have :)