Skip to content

perf: borrow the executing function from its module instance - #74

Open
matthargett wants to merge 4 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/cheaper-calls
Open

matthargett wants to merge 4 commits into
explodingcamera:nextfrom
rebeckerspecialties:perf/cheaper-calls

Conversation

@matthargett

@matthargett matthargett commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Every call or return between two functions cloned the callee's Shared<WasmFunction> and dropped the previous one (two atomic refcount updates each way) and looked the function up in the store.

A module instance now keeps a reference to its module, whose functions the store allocates contiguously (helping with locality for cache and prefetch). InterpreterRuntime holds the instance for the whole run and the executor borrows the executing function from it, so a call or return inside the instance switches a reference (I think this is what makes cache evict more often than I would expect). A direct call to one of the module's own functions also skips the address table, the host check and the owner check. Execution that moves into another instance's function (through an import, a table, a function reference, a return or an unwinding exception) ends the run and resumes with an executor for that instance. Fuel and time budgets carry over, so budgeted runs should suspend at the same points as before.

Entering a function also skips the value-stack lanes it never uses (most functions only touch the 32-bit lane), and a single-result return moves its result down once instead of popping and pushing it.

Because the executing function is borrowed from the instance rather than owned by the executor, the default (non-tail-call) dispatch loop keeps the function's instructions in a local and reloads them only when a call or return switches functions, instead of reloading the function and the slice's pointer and length on every step.

tests/cross_instance_calls.rs covers calls, tail calls, table calls, callbacks and exceptions across an instance boundary in both directions, and fuel- and time-budgeted runs that cross it.

Microbenchmark explains the uplift in the larger integrated benchmarks: a loop calling a one-line function drops from 410 to 319 instructions per iteration with the default dispatch, and from 441 to 349 with nightly-tail-calls.

Change in time against 017780e (#77) on an M4's performance cores, with cargo bench-suite (median of five interleaved runs) and cargo coremark:

benchmark default nightly-tail-calls
execute/argon2id −4.7% +0.4%
execute/compress +4.8% +0.2%
execute/decompress +0.4% −0.1%
execute/json −4.3% −1.9%
execute/nested_tinywasm −1.5% −1.7%
modes/json/fuel_per_instruction −7.8% −0.1%
modes/json/fuel_weighted −8.0% +0.4%
startup/nested_tinywasm/instantiate +0.2% −0.6%
CoreMark −1.3% −3.2%

For CoreMark the clock advances a fixed step per read, so both builds run the same iterations; its time is the median of ten interleaved launches, each the best of five runs. Parsing, encoding and decoding don't run this code.

Change in cycles per call on WasmBench (nightly-tail-calls) against 017780e, on the efficiency cores of an iPhone XS Max (A12) and iPhone SE (A13), median of five interleaved launches per build:

benchmark A12 A13
xmrsplayer (1024-frame buffer) −5.6% −5.4%
audio DSP (1000 frames × 512) −0.8% +0.1%
graphql-validation (AS) −5.2% −5.7%
crc32 (64 KB) +0.4% −2.8%
convolution 256×256 +0.3% −1.1%
sieve (10000) −1.6% +0.0%
bulk_memory (memory.copy/fill) +0.2% −1.6%
matmul relaxed-simd FMA +3.0% −3.6%
GC binary trees (~130K struct.new) −1.0% −1.9%
fib(30) −6.5% −4.4%
tail-call FSM (65536 return_call) −13.8% −10.6%
call_indirect (200K) −16.9% −8.8%
call_ref (200K) −24.9% −12.1%
vtable_poly4 (200K) −16.8% −9.6%
EH parser, exnref (4096 stmts, 25% throw) −29.4% −18.6%
geomean, cycles −8.5% −5.9%
geomean, instructions −4.7% −5.3%
geomean, wall time −8.7% −5.7%
rows faster (cycles) 11/15 14/15

Entering a function reserved and zero-filled all three value-stack lanes
even when the function never touches the 64- or 128-bit lane, which is
the common case. Skip a lane whose locals and operand-stack depth are
both zero.

A single-result return now moves the result down and truncates once,
instead of popping, truncating and pushing it back.
Every call and return between two functions cloned the callee's
`Shared<WasmFunction>` and dropped the previous one, two refcount updates
each way, and looked the function up in the store.

A module instance now keeps its own functions, which the store allocates
contiguously. The interpreter holds the instance for the whole run, so
the executor borrows the executing function and its module from it: a
call or return within the instance switches a reference, and a direct
call to one of the module's own functions skips the address table, the
host check and the owner check.

Execution that continues in another instance's function (a call through
an import, table or reference, a return, or an exception unwinding into
it) ends the run, and `InterpreterRuntime` resumes that frame with an
executor for its instance. Fuel and time budgets carry over, so a run
suspends at the same points as before.
@explodingcamera

explodingcamera commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Interesting change! In my preliminary tests I don't see reliable gains / some regressions so I'm a bit hesitant to merge this as-is due to the large amount of changes / complication of the execution machinery.
Improving call performance would definitely be great, I think I'll keep it around as a experiment (for now).

@explodingcamera
explodingcamera removed their request for review October 3, 2026 14:32
The match loop spilled the executor's function and reloaded it, then the
instruction slice's pointer and length, on every step. The executing
function is borrowed from the module instance, not from the executor, so
the loop can hold its instructions in a local and reload them only when a
step switches functions.
The instance kept its own copy of the module's function list for the
executor to borrow from, which cost an allocation and a reference count
per function on every instantiation, and as many again on drop. Hold the
module instead, which costs one reference count per instance.
@matthargett

Copy link
Copy Markdown
Contributor Author

This was meant to stack on some other PRs in my fork, but I can pull it in this one piece so the individual PR is more of a slam-dunk win on both of our sets of benchmarks. Let me know what you think, and if future PRs should be a little less incremental so the ultimate win is more obvious.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants