Skip to content

perf: keep the value stack's height in a register across tail-call handlers - #79

Open
matthargett wants to merge 4 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/stack-height-arg
Open

matthargett wants to merge 4 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/stack-height-arg

Conversation

@matthargett

@matthargett matthargett commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

The stack-height half of #78, which I separate into four commits that can I measured one by one. None of them adds a nightly feature, and each one builds and passes the tests on its own, so they can go in one at a time. The panic fix that was the fifth commit is on next as 9d810e1.

  1. Pass the tail-call handlers' Instruction in a register on arm64_32 (see below). Only what the convention passes in registers stays out of memory between two handlers.
  2. Keep each value stack's height outside its Vec. A lane's height is its own field over a Vec that only grows. Slots above the height are stale and written before they are read. This lets the handlers write the height back without unsafe. A function's operand-stack reservation is written along with its locals only when it's at most 64 slots. A larger one stays capacity, and each slot is written the first time a push reaches it, so a function whose deep branch rarely runs keeps no pages for it: 64 stores of a module whose untaken branch declares a 30,000-deep i64 stack keep next's footprint (35 KB each on macOS, 16–19 KB on glibc) instead of 273 KB and 250 KB. Only builds with debug assertions check indices against the height, and a new CI job runs the tests optimized with them. A push that reaches a new slot checks the capacity and then uses Vec::push, which makes no call; the nightly push_within_capacity would save about 1% of instructions per call on arm64.
  3. Hold the value stack in the executor while it runs. The executor takes the store's value stack when it's created and gives it back when it's dropped and around host calls, also when a host function catches a panic from a nested call (tested in tests/host_nested_panic.rs). Every handler receives the executor as a &mut, so a stack access needs one fewer load, and the compiler knows that no store to a slot or to memory changes a height (verified this in disassembly and with low-level CPU counters).
  4. Pass the 32-bit value stack's height through the tail-call handlers. A handler no longer loads the height its predecessor stored. The Unbudgeted handlers take six integer arguments and the Bounded ones five.

Apple devices

Change against next (9d810e1) with nightly-tail-calls, median of five interleaved launches per build. Energy is the kernel's estimate of the CPU energy each call used (task_power_info_v2). Commit 1 does not change the code on 64-bit arm64.

iPhone SE (A13), efficiency cores iPhone XS Max (A12), efficiency cores
commits 1–3, cycles −6.3% −4.6%
commits 1–3, energy per call −3.5% −5.0%
commit 4 on top of 3, cycles +1.6% −0.1%
commit 4 on top of 3, energy per call +2.3% +0.6%
all, cycles −4.7% (15/15 faster) −4.7% (15/15 faster)
all, energy per call −1.2% −4.4%

The iPhone 12 and the Apple TV 4K weren't available for this version. On the earlier version (an older next, with push_within_capacity, built for the A10), all four commits were −5.2% cycles on the iPhone 12 (A14) and −8.0% on the Apple TV 4K (A10X). On the A10X (Apple TV 4K first gen) the kernel doesn't reports energy or which cores ran the benchmark, which makes sense since it plugs into a wall :D I have an A10X iPad Pro, but that device didn't get updated past iPadOS 17 so the performance profiling is limited in different ways. So far I've found that what's good for A12 effiency cores maps well to A10X's high-performance (Hurricane) cores in terms of wallclock times.

On Apple Watch SE (S8, arm64_32), 13 benchmarks (the watch leaves out audio DSP and the GC trees due to peak memory usage constraints, something to optimize for later), measured on the earlier version; commit 1 hasn't changed since:

cycles wall time energy per call
commit 1 (extern "C-unwind") −13.2% (13/13 faster) −14.5%
all four commits −19.7% (13/13 faster) −19.8% −7.5%

Between some launches the watch changed its performance state, which moves energy per cycle between about 46 and 73 pJ, too often to give commit 1's energy on its own. Profiling on the physical watch is extremely annoying, and very touchy (literally!), but I feel like I got enough on-device data for it to be defensible.

All four commits against next, per benchmark:

benchmark A13 cycles A13 energy A12 cycles A12 energy
crc32 (scalar) −8.8% −2.3% −7.5% −6.6%
audio DSP −7.6% +0.0% −9.3% −8.5%
convolution (scalar) −6.5% −2.3% −6.2% −9.1%
bulk_memory (scalar) −5.7% −1.3% −4.2% −10.1%
EH parser, exnref −5.7% −4.7% −4.6% −5.3%
matmul relaxed-simd FMA −5.4% −1.1% −2.6% +3.7%
tail-call FSM −4.5% −2.2% −4.4% +2.1%
call_indirect −4.2% −2.5% −4.7% −7.7%
fib(30) −4.1% −1.3% −3.9% −4.1%
vtable_poly4 −4.1% −2.8% −3.9% −3.1%
call_ref twin: call_indirect −4.1% −2.2% −3.1% −2.0%
sieve (scalar) −3.7% +1.8% −8.3% −8.5%
graphql-validation −2.9% +0.9% −1.7% −4.2%
xmrsplayer −2.3% +2.5% −4.6% −2.4%
GC binary trees −1.1% −0.6% −1.3% +1.0%

Commit 4 pays where the height's store-to-load round trip costs: on the iPhone 12 it removed 88–97% of the memory-order flushes on every benchmark profiled, and on audio DSP back-end stalls fell from 14.5% to 3.0% of pipeline slots. On the iPhone SE and XS Max it's neutral to about 2% slower, while commits 1–3 save 6.3% and 4.6% there, so commits 1–3 also stand on their own.

The call-heavy rows are bound by the call path instead: on the iPhone 12's tail-call FSM, about a third of the samples were the Shared<WasmFunction> refcount updates on each call and return, which #74 removes.

With the default dispatch, the series no longer changes instructions per call (within 0.05% on all 15 benchmarks): next's own match loop got 1.8–5.7% cheaper since 017780e, which is what commits 2 and 3 used to give it.

Calling convention (commit 1)

  • x64: the handlers keep the Rust ABI. It passes six integer arguments in registers on System V, as many as the Unbudgeted handlers take with commit 4, but four on Windows, so on Windows every dispatch puts the Instruction on the stack today, and with commit 4 the height too. extern "rust-preserve-none" would pass twelve on every OS, but it's a nightly feature (rust_preserve_none_cc), so it isn't in this PR.
  • arm64_32 (watchOS 11): the handlers are extern "C-unwind". The Rust ABI passes the 8-byte Instruction by reference there, so on next every dispatch stores it and the next handler loads it back. The C convention passes it in a register, and C-unwind still lets a host function's panic unwind (super cool!)
  • Everything else keeps the Rust ABI.

Hot path per dispatch, weighted by xmrsplayer's dispatch mix, from the compiled handlers (instructions / stack operands / pushes and pops):

x86-64 System V Windows x64
next 33.8 / 0.47 / 1.77 47.4 / 5.44 / 5.37
commits 1–3 34.1 / 0.47 / 1.88 48.2 / 5.44 / 5.47
all four commits 35.3 / 0.47 / 2.06 51.4 / 7.21 / 6.36

Handlers that push or pop on the hot path: 258 → 273 on System V and 547 → 614 on Windows. Without rust-preserve-none, commit 4 costs on x64: the height takes a register the handler bodies would otherwise use, and on Windows it goes on the stack. These are counts from the compiled handlers; they have not been timed on x64. I have a Surface Book 2 here I can resurrect and test on if need be, but I'd love to outsource to other people's deployment targets that I surely don't have :)

@explodingcamera
explodingcamera self-requested a review October 1, 2026 16:09
explodingcamera added a commit that referenced this pull request Oct 3, 2026
originally reported in #79

Signed-off-by: Henry <mail@henrygressmann.de>
explodingcamera added a commit that referenced this pull request Oct 3, 2026
based on the fix in #79

Signed-off-by: Henry <mail@henrygressmann.de>
@explodingcamera

explodingcamera commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

I just pushed a fix for the function call panics from this to next. I do see a good performance improvement overall, and I’d definitely like to bring in more of these changes. I’m just a bit hesitant to add more nightly features right now, so I might shuffle the flags around to separate tail-call dispatch from a more general nightly feature.

I’ll probably do some more testing in this direction and pull the changes in incrementally rather than all at once.

@explodingcamera
explodingcamera removed their request for review October 3, 2026 14:32
On arm64_32 (watchOS) the Rust ABI passes the 8-byte Instruction by
reference, so every dispatch stores it and the next handler loads it back.
The C convention passes it in a register, and `C-unwind` still lets a
host function's panic unwind.

Other targets keep the Rust ABI.
The height of a value-stack lane is its own field, over a Vec that holds
every slot the lane has reached. Slots above the height hold stale values
that are written before they are read again.

The tail-call handlers can then pass the height between them in a register
and write it back without unsafe code.

Entering a function writes its locals, and its operand-stack reservation
if that is at most 64 slots, so later entries at that height compare once.
A larger reservation stays capacity, and a push writes each slot the first
time it reaches it, so a function whose deep branch rarely runs keeps no
pages for it. With 64 stores of a module whose untaken branch declares a
30000-deep i64 stack, each store keeps next's footprint (35 KB on macOS,
16-19 KB on glibc) instead of 273 KB and 250 KB.

v128.any_true and i8x16.all_true test the vector with integer arithmetic:
LLVM scalarized their iterator forms once pushes gained that path.

Only builds with debug assertions check indices against the height. CI
also runs the tests optimized with them, keeping the release profile's
wrapping arithmetic.

Entering a function zeroes one or two locals with direct stores, and only
the ones in between with memset.
The executor takes the store's value stack when it is created and gives it
back when it is dropped, and around host calls, which read their arguments
from the store and may call back into wasm.

Every handler receives the executor as a `&mut`, so a stack access reached
through it needs one load fewer (the store pointer), and the compiler knows
that no store to a slot or to memory changes a stack's height, so it can
keep the heights in registers within a handler.
…lers

Each tail-call handler receives the height of the 32-bit value stack, which
most instructions push to or pop from, as its last argument. It writes the
height to the stack before its body and reads it back before dispatching,
and with the stack in the executor the compiler can keep it in a register
in between, as it does in most handlers. A handler no longer loads the
height its predecessor stored, a store-to-load round trip through memory
on every dispatch.

The Unbudgeted handlers now take six integer arguments and the Bounded ones
five, which stay in registers on arm64 (eight) and x86-64 System V (six).
Windows x64 passes four, so there the height goes on the stack along with
the Instruction, which already does without this change.
@matthargett
matthargett force-pushed the perf/stack-height-arg branch from 29679a7 to fc1f817 Compare October 4, 2026 01:23

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants