feat(precompiles): read logs across a block range - #349
Conversation
`eth_getLogs` is capped per call, so reading any history longer than the cap means walking it in chunks. That loop is short but has two failure modes that both look like a working indexer, and every project needing logs writes it again. The inclusive boundary is the first. The public endpoints allow 2000 blocks and apply the check as `toBlock - fromBlock + 1 <= 2000`, so a range built as `from + 2000` asks for 2001 and is rejected on every chunk with "block range too large (2001), maximum allowed is 2000 blocks". Measured on both networks: 2000 succeeds, 2001 does not. Halving the chunk to be safe works but doubles the round trips a backfill needs. The confirmation depth is the second. Sei finalises a block as it is produced, so there is no reorg window to wait out, and a default lag copied from an Ethereum-shaped library is latency with nothing behind it. `getLogsInRange` reads to head; a caller wanting to lag passes an explicit `toBlock`. `blockRanges` exposes the same arithmetic as a generator without making requests, so a backfill can be planned or driven by a bounded worker pool rather than one sequential loop. `onChunk` reports progress, because a backfill over long history is thousands of requests and is otherwise indistinguishable from a hang. Takes a `PublicClient` rather than constructing one, so it works with whatever transport and chain the caller has configured. No dependency or peer range changes -- this uses the `viem` peer already declared. Verified against a live endpoint as well as the unit tests: a 4501-block span issued three requests of 2000, 2000 and 501 blocks, none over the cap.
PR SummaryMedium Risk Overview The walk uses viem’s Documentation, a minor changeset, export checks in Reviewed by Cursor Bugbot for commit ba4ae4f. Bugbot is set up for automated code reviews on this repo. Configure here. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #349 +/- ##
==========================================
+ Coverage 97.17% 97.82% +0.65%
==========================================
Files 80 81 +1
Lines 5410 5652 +242
==========================================
+ Hits 5257 5529 +272
+ Misses 153 123 -30
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@seidroid review |
|
Claude finished @masih's task in 2m 52s —— View job Review:
|
|
@seidroid review |
|
Claude finished @alexander-sei's task in 2m 41s —— View job Review:
|
The first cut walked a range at a fixed width and broke on the limits it was meant to hide. One dense chunk failed the whole read and took every chunk already fetched with it, and only `event` was forwarded, so reads filtered on indexed `args` or covering several events couldn't use it at all. streamLogsInRange now does the walk and yields each chunk with its logs, and getLogsInRange collects it. A span too heavy to answer is halved and asked again: from sei-chain v6.7 the node refuses one matching more than max_log_no_block logs, before v6.7 it serves a bounded request whole and a busy range passes viem's 10 MiB response limit instead, and a span can also time out. The walk then holds below the span that failed and only tries it again after eight successes, since growing straight back made every other request on a dense range a refusal. "block range too large" drops the walk to the maximum the node names. Sei's busy refusals and its large query rate limit are retried with backoff, as viem won't retry a -32000, and a rate limit that outlasts the retries steps the walk down to 100 blocks, below which the limit doesn't apply. Refusals are matched on the node's wording along the cause chain, because viem's short message for all of them is "Missing or invalid parameters". The filter is viem's own (address, event with args, events, strict), and a whole contract ABI works as `events`, keeping only its events the way getContractEvents does. The getLogs request is typed with no cast. Any viem Client works through getAction, including a wallet client carrying an account. onChunk gets the chunk's logs and is awaited, and the head is read with cacheTime 0 so a read to head doesn't trail by the polling interval. A number, a tag, a blockHash or an ABI with no events is refused with a TypeError up front rather than failing oddly or reading the wrong blocks. The README example no longer starts at a testnet block the public node has already pruned, and imports from the viem entry point, which doesn't load ethers.
There was a problem hiding this comment.
This is a well-built, well-tested addition. getLogsInRange, streamLogsInRange and blockRanges get the inclusive-span arithmetic right, and the halving, retry and range-refusal logic keeps chunks contiguous with no gaps or overlaps. It has a minor changeset and the export checks are updated. I found no blockers, only a few robustness and ergonomics suggestions.
Findings: 0 blocking | 7 non-blocking | 2 posted inline
Blockers
- None at the file/PR level.
Non-blocking
- The Cursor second-opinion pass produced no output (
cursor-review.mdis empty). Codex reported no material issues, and noted it could not run the tests because Bun was not available. - There is no way to cancel a walk (no
AbortSignal/signaloption). A backoff can sleep for up to 30s per retry, andstream.return()or a caller timeout can't interrupt a pendingsleepor in-flight request. Consider accepting an optionalsignaland passing it to the sleep. It would be worth adding before the API is published, since adding it later is harder to do cleanly. - The PR description lists only
getLogsInRange,blockRangesandMAX_GET_LOGS_BLOCK_RANGE. The changeset, README and code also addstreamLogsInRangeand the adaptive halving and retry behaviour. Please update the description so reviewers and release notes match what ships. - The behaviour claims (v6.7
max_log_no_blockrefusals, the error wording, the 100-block rate-limit threshold, about 30 req/s) are cited asevmrpc/filter.goand similar, but there are no links. Because the classifier matches on the node's exact wording, please link the sei-chain source (commit or tag) in the code comments or the PR. A future wording change would otherwise silently turn a retryable refusal into a hard failure. - Positive:
bun run typecheckusestsconfig.test.json, which includes__tests__/**, so the@ts-expect-errortype assertions inlogs.spec.tsare actually checked in CI. - 2 suggestion(s)/nit(s) flagged inline on specific lines.
| continue; | ||
| } | ||
| if (kind === 'rate-limited' && span > RATE_LIMIT_FREE_SPAN) { | ||
| heavy = RATE_LIMIT_FREE_SPAN + 1n; |
There was a problem hiding this comment.
[suggestion] After stepping down to 100 blocks for the rate limit, heavy = 101 is cleared after PROBE_AFTER (8) successes. The walk then probes 200 again, and on an endpoint that is still rate limited it spends the full retryCount backoff (500+1000+2000ms by default, plus 3 extra requests) before dropping back to 100. On a long backfill against a public node this repeats every 9 chunks and adds load to a node that is already throttling. Consider not probing above RATE_LIMIT_FREE_SPAN after a rate-limit step-down (for example, a separate flag or a longer probe interval), or skipping the retries when a probe is what triggered the rate limit.
| if (names.includes('TimeoutError')) return 'client-timeout'; | ||
| // "query matches too many logs" and "query matches too many log bytes", | ||
| // viem's response size cap, and a node that timed out on the span. | ||
| if (names.includes('ResponseBodyTooLargeError') || /query matches too many log|request timed out/i.test(text)) return 'heavy'; |
There was a problem hiding this comment.
[nit] /request timed out/i is matched against the whole cause-chain text and classified as heavy whether or not the endpoint has answered. The answered guard only protects TimeoutError. Any wrapped error whose text contains "request timed out" (for example a proxy or gateway message on a down endpoint) will halve down to one block, about 11 requests, before it throws. Consider matching only the node's -32002 code or error name for the node-timeout case, or applying the same answered guard.
…it without waiting A long backfill had no way to stop short of killing the process, and a wait between retries can run to 30 seconds. `signal` now stops the walk before the next request or part way through a wait. A request already sent still finishes, since viem's getLogs takes no signal. Once the rate limit has pushed the walk under 100 blocks, each probe back over it used to pay the whole backoff again on a node that was still throttling, every nine chunks. A probe that meets the limit again now steps straight back down, so the backoff is paid once per walk. A node timeout now needs the node's own -32002 to count as a heavy span. The same words from a gateway in front of an endpoint that is down used to halve all the way to one block before throwing. The classifier's comment links the sei-chain lines whose wording it matches.
There was a problem hiding this comment.
This PR adds streamLogsInRange, getLogsInRange, blockRanges and MAX_GET_LOGS_BLOCK_RANGE to @sei-js/precompiles. The chunking, backoff and error handling are careful and heavily tested, and a minor changeset is included with no dependency changes. No blocking issues; the notes below are small doc and behaviour inconsistencies.
Findings: 0 blocking | 8 non-blocking | 3 posted inline
Blockers
- None at the file/PR level.
Non-blocking
- The Cursor second-opinion pass produced no output (
cursor-review.mdwas empty). Codex reported no material issues but could not run tests. - The PR description is out of date. It lists only
getLogsInRange/blockRangesand mentions '19 unit tests', but the diff also addsstreamLogsInRange, adaptive halving, retry/backoff, abort support and far more tests. Please update the body so reviewers and release notes match what ships (the changeset text is already accurate). - Error handling depends on matching the node's exact error wording (
query matches too many log,block range too large … maximum allowed is N blocks,server too busy, and so on). The code links the sei-chain source it relies on, which is good. If sei-chain rewords a message, these cases fail safe by throwing rather than silently dropping logs, but consider noting in the changeset or README that this relies on sei-chain wording up to v6.7. - Some README/JSDoc figures are operational facts that will go stale: the public mainnet endpoint 'keeps well under a day of history', 'about 30 requests a second over 100 blocks', and the 10,000
max_log_no_blockdefault. Link a source for each or soften the wording. - Root index now re-exports
./viem/logs, which importsviem/actionsandviem/utilsat runtime. That's fine because the root already importsviemthrough./viem/chainandsideEffects: falsekeeps it tree-shakeable. Worth confirming that ethers-only consumers who import from the package root are unaffected. - 3 suggestion(s)/nit(s) flagged inline on specific lines.
| // back over it steps straight down again rather than waiting out the | ||
| // backoff on a node that is still throttling. | ||
| const stepDown = kind === 'rate-limited' && span > RATE_LIMIT_FREE_SPAN && (limited || retries >= retryCount); | ||
| const transient = kind === 'busy' || kind === 'rate-limited' || (kind === 'behind' && toBlock === undefined); |
There was a problem hiding this comment.
[nit] The JSDoc (and changeset) say only a final chunk refused as after latest available block is retried when no toBlock was given. This condition retries behind on any chunk whenever toBlock === undefined. Harmless in practice, but either restrict it to to === endBlock or update the doc so the two agree.
| // Once the rate limit has pushed the walk under 100 blocks, a later probe | ||
| // back over it steps straight down again rather than waiting out the | ||
| // backoff on a node that is still throttling. | ||
| const stepDown = kind === 'rate-limited' && span > RATE_LIMIT_FREE_SPAN && (limited || retries >= retryCount); |
There was a problem hiding this comment.
[nit] limited never resets once set. For the rest of the walk, any rate-limit refusal on a span over 100 blocks steps down straight away with no backoff, even hours into a long backfill after the node has stopped throttling. That matches the stated intent (one wasted request per probe), but consider resetting limited after a successful probe over 100 blocks, so a later spell of throttling gets the normal retry/backoff again.
| throw nodeRefusal('server too busy, rejecting new request (pending: 900, threshold: 800)'); | ||
| }); | ||
| const controller = new AbortController(); | ||
| setTimeout(() => controller.abort(new Error('shutting down')), 20); |
There was a problem hiding this comment.
[nit] This test uses real timers with a 20ms abort and asserts < 1_000 ms of wall-clock time. The margin is generous, but it's the one timing-dependent test in the suite and could flake on a heavily loaded CI runner. Consider driving it with the withInstantTimers-style mock or fake timers instead.
…rt test Once the rate limit had pushed a walk under 100 blocks, every later throttle on a probe stepped straight back down for the rest of the walk, even hours on when a short wait would have done. A probe over 100 that gets through now resets that, so the next spell is waited out first again. The abort test waited on a real 20 ms timer against a one second budget, which a loaded CI runner could miss. It now mocks the timer so the walk can only leave the wait through the signal, and races it against a real one second timer, because with setTimeout mocked bun's own test timeout can't fire and a regression would hang CI rather than fail it. The docs say a chunk refused as "after latest available block" is retried without a toBlock, which is what the code does, rather than only the last one. The README links the sei-chain defaults it quotes, and no longer states how much history one provider's public endpoint keeps.
There was a problem hiding this comment.
This PR adds a well-tested chunked eth_getLogs walker (streamLogsInRange / getLogsInRange / blockRanges) to @sei-js/precompiles, with a minor changeset, README docs and export checks. The logic is careful and I found no correctness or security blockers, only small suggestions.
Findings: 0 blocking | 6 non-blocking | 1 posted inline
Blockers
- None at the file/PR level.
Non-blocking
- Cursor's review file (
cursor-review.md) was empty, so that pass gave no output. Codex reported no material issues. - Classifying refusals depends on the exact wording of sei-chain's error strings, such as
query matches too many log,block range too large (...)andserver too busy. The source is documented and the spec pins the wording, but if sei-chain rewords a message, a refusal the walk used to handle will be thrown to callers without warning. Consider mentioning this in the README, or tracking sei-chain releases. - Runtime imports from
viem/actionsandviem/utils(getAction) are new. The root entry already pulls in viem at runtime throughviem/chain, so ethers-only users aren't newly affected. It's still worth checking thatgetActionis exported fromviem/utilsacross the whole^2.55.16peer range. getLogsInRangepassesrestthrough a cast (as StreamLogsInRangeParameters<...>) to droponChunk. That's harmless, but a wrong option can get past the type check there without notice. Consider narrowing the destructure's type instead of casting.- The changeset, the README section and the export checks in
check-precompile-exports.tsare all present, which meets the guidelines' changeset requirement. - 1 suggestion(s)/nit(s) flagged inline on specific lines.
| } | ||
| if (stepDown) { | ||
| limited = true; | ||
| heavy = RATE_LIMIT_FREE_SPAN + 1n; |
There was a problem hiding this comment.
[nit] The rate-limit step-down overwrites heavy with 101. If a smaller heavy span (for example 500 after a log-cap refusal) was being held, that memory is lost. After PROBE_AFTER successes the walk grows back past it and pays for another heavy refusal. Consider heavy = heavy !== undefined && heavy < RATE_LIMIT_FREE_SPAN + 1n ? heavy : RATE_LIMIT_FREE_SPAN + 1n, or keep the two holds in separate variables. It's minor, since the walk still converges correctly.
Summary
Adds
streamLogsInRange,getLogsInRange,blockRangesandMAX_GET_LOGS_BLOCK_RANGEto@sei-js/precompiles, for reading logs across more blocks than oneeth_getLogsrequest allows.Every project that reads history writes this loop, and the ways it breaks on Sei aren't obvious. The node counts a span inclusively, so
from + 2000asks for 2001 blocks and is refused on every request. A dense range is refused for matching more thanmax_log_no_blocklogs from sei-chain v6.7, and before v6.7 it's served whole and passes viem's 10 MiB response limit instead. Sei's busy and rate limit refusals arrive as-32000, which viem doesn't retry.The walk sizes every request so the node answers it. A span too heavy to answer (log cap, response size or timeout) is halved, and the walk holds below it until a run of successes. A
block range too largerefusal drops it to the maximum the node names. Busy refusals are retried with backoff, and a rate limit that outlasts the retries steps it under the 100 block threshold the limiter ignores. Anything else is thrown untouched, and asignalstops the walk between requests or part way through a wait.streamLogsInRangeyields each chunk with its logs, so a backfill can store as it goes and resume from the lasttoBlock.getLogsInRangecollects the walk and awaits an optionalonChunk. Both take viem'sgetLogsfilter (address,eventwithargs,events,strict), a whole contract ABI asevents, and any viemClient.Critical notes
toBlock. Nodes before v6.7 silently cut an open-ended request off atmax_log_no_blocklogs. On public testnet an open-ended request came back with exactly 10,000 logs and was missing the last ~600 blocks.toBlockthe walk reads to a fresh head (cacheTime: 0), since Sei finalises a block as it's produced.viem/actionsandviem/utils, forgetAction, which is what lets a client carrying an account work. No dependency or peer range changes.Related issue
Part of PLT-834. Written for Sei Street's indexer, which carries its own copy of this loop today.
Test plan
bun run checkbun run buildbun run test: every package green on a clean checkout merged withmain, registry submodules at their pinned commits, bun 1.3.14 (117 in precompiles, 64 of them for this)bun run lint:pack:allThe unit tests assert on the requests rather than a node's answers, against a recording client and a real viem client on a custom transport. They cover inclusive spans, gapless and non-overlapping coverage, parity with
blockRanges, halving on each heavy refusal and the hold below it, the node's named maximum, retries and their backoff, the rate limit step down, cause chain classification, filter forwarding (topics on the wire) and input validation.logs.tsis at 100% line coverage, and 28 handmade mutants of the walk each fail at least one test.Against the public endpoints, read only:
chunkSize: 2500nwas refused withmaximum allowed is 2000 blocksonce, and the rest walked at 2000maxResponseBodySizelowered to 2 MB: 4000 dense blocks walked as 2000, 1000, 500, 250 and back up, returning the same 24,238 logs as uncapped raw requestsAgainst Sei Street's localnet (sei-chain with the v6.7 log cap,
max_log_no_block = 10000), with Sei Street's ledger and API moved onto this package and every window compared with its current hand rolled loop and with the trades its indexer wrote to Postgres:argsacross 4,930 blocks: the 3 in Postgres, in 3 requests, where a singlegetLogsis refused asblock range too large (4930)In ten 9 minute runs of the full game on Sei Street's Giga localnet, alternating its current indexer with one on this package, the indexer sent 28% fewer
eth_getLogsrequests with 62% fewer refused (every run on this package below every run without it), chain TPS and CPU were unchanged, and the trades it wrote matched the chain exactly.Checklist