perf: borrow the executing function from its module instance - #74
Open
matthargett wants to merge 4 commits into
Open
matthargett wants to merge 4 commits into
matthargett wants to merge 4 commits into
Conversation
matthargett
force-pushed
the
perf/cheaper-calls
branch
from
September 27, 2026 22:49
3e11936 to
8370a09
Compare
Entering a function reserved and zero-filled all three value-stack lanes even when the function never touches the 64- or 128-bit lane, which is the common case. Skip a lane whose locals and operand-stack depth are both zero. A single-result return now moves the result down and truncates once, instead of popping, truncating and pushing it back.
Every call and return between two functions cloned the callee's `Shared<WasmFunction>` and dropped the previous one, two refcount updates each way, and looked the function up in the store. A module instance now keeps its own functions, which the store allocates contiguously. The interpreter holds the instance for the whole run, so the executor borrows the executing function and its module from it: a call or return within the instance switches a reference, and a direct call to one of the module's own functions skips the address table, the host check and the owner check. Execution that continues in another instance's function (a call through an import, table or reference, a return, or an exception unwinding into it) ends the run, and `InterpreterRuntime` resumes that frame with an executor for its instance. Fuel and time budgets carry over, so a run suspends at the same points as before.
matthargett
force-pushed
the
perf/cheaper-calls
branch
from
September 29, 2026 23:58
8370a09 to
2bdebab
Compare
explodingcamera
self-requested a review
October 1, 2026 16:09
Owner
|
Interesting change! In my preliminary tests I don't see reliable gains / some regressions so I'm a bit hesitant to merge this as-is due to the large amount of changes / complication of the execution machinery. |
explodingcamera
removed their request for review
October 3, 2026 14:32
The match loop spilled the executor's function and reloaded it, then the instruction slice's pointer and length, on every step. The executing function is borrowed from the module instance, not from the executor, so the loop can hold its instructions in a local and reload them only when a step switches functions.
The instance kept its own copy of the module's function list for the executor to borrow from, which cost an allocation and a reference count per function on every instantiation, and as many again on drop. Hold the module instead, which costs one reference count per instance.
Contributor
Author
|
This was meant to stack on some other PRs in my fork, but I can pull it in this one piece so the individual PR is more of a slam-dunk win on both of our sets of benchmarks. Let me know what you think, and if future PRs should be a little less incremental so the ultimate win is more obvious. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every call or return between two functions cloned the callee's
Shared<WasmFunction>and dropped the previous one (two atomic refcount updates each way) and looked the function up in the store.A module instance now keeps a reference to its module, whose functions the store allocates contiguously (helping with locality for cache and prefetch).
InterpreterRuntimeholds the instance for the whole run and the executor borrows the executing function from it, so a call or return inside the instance switches a reference (I think this is what makes cache evict more often than I would expect). A direct call to one of the module's own functions also skips the address table, the host check and the owner check. Execution that moves into another instance's function (through an import, a table, a function reference, a return or an unwinding exception) ends the run and resumes with an executor for that instance. Fuel and time budgets carry over, so budgeted runs should suspend at the same points as before.Entering a function also skips the value-stack lanes it never uses (most functions only touch the 32-bit lane), and a single-result return moves its result down once instead of popping and pushing it.
Because the executing function is borrowed from the instance rather than owned by the executor, the default (non-tail-call) dispatch loop keeps the function's instructions in a local and reloads them only when a call or return switches functions, instead of reloading the function and the slice's pointer and length on every step.
tests/cross_instance_calls.rscovers calls, tail calls, table calls, callbacks and exceptions across an instance boundary in both directions, and fuel- and time-budgeted runs that cross it.Microbenchmark explains the uplift in the larger integrated benchmarks: a loop calling a one-line function drops from 410 to 319 instructions per iteration with the default dispatch, and from 441 to 349 with
nightly-tail-calls.Change in time against
017780e(#77) on an M4's performance cores, withcargo bench-suite(median of five interleaved runs) andcargo coremark:nightly-tail-callsFor CoreMark the clock advances a fixed step per read, so both builds run the same iterations; its time is the median of ten interleaved launches, each the best of five runs. Parsing, encoding and decoding don't run this code.
Change in cycles per call on WasmBench (
nightly-tail-calls) against017780e, on the efficiency cores of an iPhone XS Max (A12) and iPhone SE (A13), median of five interleaved launches per build: