pref: borrow instructions in tail call dispatch - #75
Conversation
Signed-off-by: Henry <mail@henrygressmann.de>
|
measuring locally, this PR made each call and return about 16 instructions larger (397 → 413 per loop iteration), because every function switch goes back to its run loop and clones the function handle. looking at PMU counters on iPhone 12, CPU back-end pipelne stalls go down -36% and mispredicts go down -30%, but memory order flushes go up 4x and front-end stalls go up ~10% I revised PR #74 a bit, and on top of this PR we get back down to 323 instructions per loop iteration while not making most other CPU perf counters regress. it's probable that it's less of a wash on laptop/desktop x64, but on mobile efficiency cores there's enough sensitivity where more parts of the CPU pipeline need to get better together to get a purer end-to-end/walltime win. |
Based on the idea from #64 - moves the instruction borrow into dispatch handlers. This version runs even lightly better on my machine, taking coremark from around 1500 to 1630 and improves all benches in the local suite around 5-15%.