Avoid copying ordered pending blocks through tmpBlock

perfloop/victoriametrics · AVOIDABLE COPY

https://perfloop.ai/t/oss/case_6rqaezys13

Verdict

PROPOSED · opened 2026-08-26

Hypothesis

The source-trace check followed `partition.inmemoryPartsMerger` through `mergeParts`, `mergePartsInternal`, and `mergeBlockStreams` to `mergeBlockStreamsInternal`. While a merge batch succeeds, the internal loop handles one `bsm.NextBlock()` result at a time. A same-MetricID current block reaches `mergeBlocks` unless the pending block is already too big and is ordered before it. For the strict ordered cases, `mergeBlocks` appends both remaining row slices into `tmpBlock`; this can repeat for consecutive blocks of one series during a compaction.

The source shows a real materialization rather than a slice-header transfer. `tmpBlock.Reset()` gives the destination zero length, and the two append paths write the retained pending rows P and current rows B into its timestamp and value arrays. When P+B exceeds `maxRowsPerBlock`, which is 8,192 rows, the caller then uses `pendingBlock.CopyFrom(tmpBlock)` to materialize the tail again before writing the first block. An output-sized pair of int64 arrays is 128 KiB before headers, so the ordered route can move meaningful data even if its backing capacity is already warm. `blockStreamReader.NextBlock` resets its source Block, so B still needs an owned copy; P already belongs to `pendingBlock`. A specialized pending-before-current route could fill the existing pending block, copy B's remainder into the next pending block, and remove the P transfer plus the branch-specific tail transfer. The runtime share of these copies is unknown from source alone.

Confirm this with a compaction benchmark that sweeps P and B around the 8,192-row boundary for same-series strictly ordered blocks, including both P+B at or below and above the boundary. Collect CPU and allocation profiles that attribute bytes and time to `mergeBlocks` and `Block.CopyFrom`, then measure end-to-end in-memory compaction CPU and latency. Differential output tests must cover retention-trimmed P, the exact block boundary, deduplication enabled and disabled, and equal timestamps so the new route preserves the existing observable output.

Change to test: Add a pending-before-current non-overlap path in mergeBlockStreamsInternal that retains the owned pending rows, copies only current rows needed for the full block and its successor, and writes at the same maxRowsPerBlock boundary. Preserve retention filtering, timestamp order, raw block boundaries, and deduplication behavior, and do not alias the reader block because NextBlock resets it.

Where it lives

perfloop/victoriametrics · lib/storage/partition.go

Evidence

No usable result yet.

Timeline