Unsorted merge of address-chunked eth_getLogs rewinds the fetch cursor and re-indexes logs

perfloop/rindexer · OVERFETCHING

https://perfloop.ai/t/oss/case_4pnamepnjb

Verdict

VERIFIED · settled 2026-09-30 · pull request opened as joshstevens19/rindexer#474

What happened: The assertion is violated on the comparison and satisfied with this change.

Hypothesis

Source facts (upstream master 64cc742 = fork): 1. provider.rs:636-658 sends any filter with addresses to get_logs_for_address_in_batches unless address_filtering is in-memory; default chunk is 1000 addresses (provider.rs:117, :654). 2. provider.rs:677-687 chunks the address HashSet (helpers/array.rs:3), runs one eth_getLogs per chunk via try_join_all, and returns `chunked_logs.concat()`. Each chunk is block-ordered; the concatenation is not. 3. Factory filters pass every known child address as one HashSet (event/rindexer_event_filter.rs:103-123), so a factory with >1000 children (e.g. the repo's Uniswap V3 factory examples) takes this path by default. 4. Every fetch loop dispatches the whole batch, then sets next from_block = logs.last().block_number + 1: historic fetch_logs.rs:434 and :494-497, parallel fetch_logs_once :952 and :1008-1011, live :1527 and :1599-1604. If the last chunk ends at a lower block than an earlier chunk reached, the next request re-covers blocks whose logs were already dispatched, and they are dispatched again. 5. Raw Postgres event tables have only `rindexer_id SERIAL PRIMARY KEY` (database/generate.rs:36-47), no uniqueness on (tx_hash, log_index), so those logs become duplicate rows; they also re-enter streams and custom-table operations. Callbacks also get out-of-order logs within a batch. No upstream issue/PR covers this. Correctness only; no performance claim.

Change to test: In get_logs_for_address_in_batches (provider.rs:671-688), when more than one chunk was fetched, order the merged logs by (block_number, log_index) instead of the plain concat. Leave the single-chunk and in-memory paths and the fetch loops unchanged. Tests that fail before and pass after: a JsonRpcCachedProvider on alloy's mocked Asserter with MaxAddressPerGetLogsRequest(1) and two addresses whose canned responses interleave (chunk A: blocks 100 and 500, chunk B: block 200); assert get_logs returns them in (block, logIndex) order, and that fetch_historic_logs_stream's next from_block is 501, not 201. Add a `- fix:` changelog bullet per AGENTS.md. Run cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, cargo nextest run --exclude rindexer_rust_playground --workspace.

Where it lives

perfloop/rindexer · core/src/provider.rs

Evidence

The assertion is violated on the comparison and satisfied with this change: `Given a two-address eth_getLogs filter split into one-address requests whose successful responses contain logs at blocks 100 and 500 in one response and block 200 in the other, the merged JsonRpcCachedProvider result and timestamp-enriched parallel historic worker batch are ordered by (block_number, log_index), and the historic fetch cursor advances to from_block 501.`

The assertion is violated on the comparison and satisfied with this change: `For BlockFetcher::attach_log_timestamps, with two block-100 logs in input order (timestamps Some(1000) and None) and a provider timestamp of 2000, the output timestamps are [Some(1000), Some(2000)] in input order.`

Checks: 5 of 5 passed. Verification: no defect found.

Timeline