Skip to content

perf: inline common memory-zero load offsets - #63

Closed
matthargett wants to merge 1 commit into
explodingcamera:nextfrom
rebeckerspecialties:perf/inline-wasm32-memory-loads
Closed

matthargett wants to merge 1 commit into
explodingcamera:nextfrom
rebeckerspecialties:perf/inline-wasm32-memory-loads

Conversation

@matthargett

Copy link
Copy Markdown
Contributor

First of a two-PR interpreter performance stack on next.

The common memory-0 i32.load, i32.load8_u, and i32.load16_s paths currently fetch their static offset from the function operand pool on every dispatch. After normal instruction selection, this embeds a 32-bit offset in the existing 8-byte instruction. Other memories and wider offsets retain the original path. Existing opcode numbers stay fixed; the versioned archive format advances to 07. No unsafe code is added.

Three interleaved Release runs per variant at .utility QoS, with the benchmark thread on efficiency cores:

Device Cycles/call geomean Retired instructions Audio DSP wall time Scalar convolution Scalar CRC32
iPhone 12 (A14), 7 cases −2.0% −2.0% −3.5% −5.5% −2.6%
iPhone XS Max (A12), 7 cases −2.6% −1.9% −1.9% −3.4% −11.7%

xmrsplayer was neutral (+0.3% A14, +0.2% A12). A broader load-and-store variant regressed A14 wall time, so stores are excluded here.

Validation: full tinywasm tests with nightly-tail-calls; focused optimized/unoptimized, multi-memory, and memory64/wide-offset tests; no-std parser/validate/archive check; formatting and library Clippy with warnings denied.

@matthargett
matthargett force-pushed the perf/inline-wasm32-memory-loads branch from aefdda3 to f0bcaba Compare September 25, 2026 04:21
@explodingcamera

explodingcamera commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Thanks for the PR! The idea seems sound, but the extra pass doesn’t fit the single-pass lowering approach I want to keep. Moving the rewrite into instruction selection could also prevent later load fusions though. I’m planning to add accumulator registers soon, so I’d rather revisit this alongside that work. The second part seems interesting though, I'll have to try that one out once it's ready 👍 There should also be space for a u16 memory index in the instruction too so I'd add that as well together with the registers.

@matthargett

Copy link
Copy Markdown
Contributor Author

Thanks for the PR! The idea seems sound, but the extra pass doesn’t fit the single-pass lowering approach I want to keep. Moving the rewrite into instruction selection could also prevent later load fusions though. I’m planning to add accumulator registers soon, so I’d rather revisit this alongside that work. The second part seems interesting though, I'll have to try that one out once it's ready 👍 There should also be space for a u16 memory index in the instruction too so I'd add that as well together with the registers.

That makes sense to me, given the current design intent (translate to IR instead of execute directly like Hermes and JSC IPint, single-pass, etc). I'm keeping an eye on your acc branch and trying to eat around the edges with uplift that won't conflict there. I just wanted to mention the framework I'm evaluating against on my local device profiling so that you can cross-check your approach as you iterate in that branch:

  • in PR pref: borrow instructions in tail call dispatch #75 Instructions drop only 3%, but back-end stall cycles drop 36% and branch mispredicts drop by 15%. On the old next head commit, finding the next handler took four dependent loads before the dispatch branch. With the instruction slice in registers it takes two. Counting instructions would have missed this almost entirely.
  • My PR stack is not improving on every front in terms of low-level CPU counters. Nearly all of its gain is fewer instructions. Per call, front-end stalls rise 7%, discarded work 9%, and memory-order flushes 74%. The flushes are not from app launch or parsing: across a whole capture they go from 0.4M (next) to 2.1M (pref: borrow instructions in tail call dispatch #75) to 4.1M (my fork PR stack tht I'm peeling off one at a time for contribution after I have hard data).
  • My educated guess: the value stack's length and top value are stored by one handler and loaded by the next, and the faster the dispatch (fewer instruction), the more often that load runs ahead of the store. They cost roughly 1% of cycles now and will grow as handlers get faster.
  • The next win of pref: borrow instructions in tail call dispatch #75's kind is keeping the value-stack pointer in the handler arguments, like pref: borrow instructions in tail call dispatch #75 did for the instruction slice. In the stack's i32.add, 7 of 23 instructions and a three-load chain are stack bookkeeping, and it's also the likely fix for the flushes. That seems in your exp/acc branch territory, but lmk if you think there's a slice I can work on and contribute in parallel, let me know.
  • The CPU pipeline front-end share grows as handlers shrink, because every dispatch ends in a taken indirect branch. Fusion or specialization and register operands attack that next ceiling, but making handlers faster will probably run head-first into that ceiling for now. This is in line with my performance work in WAMR, Luau, JavaScriptCore, and WasmEdge over the last several years.
  • I'm using XCode Instruments, but if you're on x64 you can try to find this same balance where the whole CPU pipeline moves forward in efficiency with Intel vTune or AMD uProf, which should give the same info.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants