perf: inline common memory-zero load offsets - #63
matthargett wants to merge 1 commit into
Conversation
aefdda3 to
f0bcaba
Compare
|
Thanks for the PR! The idea seems sound, but the extra pass doesn’t fit the single-pass lowering approach I want to keep. Moving the rewrite into instruction selection could also prevent later load fusions though. I’m planning to add accumulator registers soon, so I’d rather revisit this alongside that work. The second part seems interesting though, I'll have to try that one out once it's ready 👍 There should also be space for a u16 memory index in the instruction too so I'd add that as well together with the registers. |
That makes sense to me, given the current design intent (translate to IR instead of execute directly like Hermes and JSC IPint, single-pass, etc). I'm keeping an eye on your
|
First of a two-PR interpreter performance stack on
next.The common memory-0
i32.load,i32.load8_u, andi32.load16_spaths currently fetch their static offset from the function operand pool on every dispatch. After normal instruction selection, this embeds a 32-bit offset in the existing 8-byte instruction. Other memories and wider offsets retain the original path. Existing opcode numbers stay fixed; the versioned archive format advances to 07. No unsafe code is added.Three interleaved Release runs per variant at
.utilityQoS, with the benchmark thread on efficiency cores:xmrsplayer was neutral (+0.3% A14, +0.2% A12). A broader load-and-store variant regressed A14 wall time, so stores are excluded here.
Validation: full
tinywasmtests withnightly-tail-calls; focused optimized/unoptimized, multi-memory, and memory64/wide-offset tests; no-std parser/validate/archive check; formatting and library Clippy with warnings denied.