Defer bitmap materialization for uniform validity slices

perfloop/btrblocks · ALLOCATION HOT LOOP

https://perfloop.ai/t/oss/case_3neh8480ph

Verdict

VERIFIED · settled 2026-07-24 · merged as axiomhq/btrblocks#4

What happened: Validated after narrowing the focused correctness fixture to its required 19-row boundary range. On the uniform 4,096-row Strings.Slice workload, allocs/op fell from 3 to 2 (33.3%, p=0.0000108); B/op fell from 19,040 to 18,528 (2.69%, p=0.0000108), while ns/op held from 17,130.5 to 16,841 (p=0.4243). On the mixed first-valid 4,096-row control, ns/op held from 22,309.5 to 21,681.5 (p=0.3150), with B/op unchanged at 19,040 and allocs/op unchanged at 3. All recorded correctness, full-suite, vet, and formatting checks passed. These are component-level Strings.Slice shapes only, not a claim about production call frequency or end-to-end impact.

Hypothesis

For a nonempty slice of a globally mixed source, array.Validity.Slice allocates ceil(length/8) bytes before the scan, then discards that bitmap when the scan's nullCount selects AllValid or AllNull. The incoming-call check found array.Strings.Slice calling this method at array/string.go:149, so the real cadence is one such speculative materialization per Strings.Slice invocation, not once per row. A temporary reproducible probe, `go test ./array -run '^$' -bench '^BenchmarkValiditySlicePerfProbe' -benchmem -benchtime=100ms -count=1`, built an 8,192-row mixed bitmap whose first and second 4,096-row halves were respectively all valid and all null. Slicing either uniform half measured one allocation and 512 B/op (13,232 ns/op valid; 10,923 ns/op null), confirming that this allocation currently materializes even though the returned representation is canonical and bitmap-free. This probe does not establish production slice widths, call rate, or end-to-end dominance. The proposed change removes one ceil(rangeLen/8)-byte allocation and its bitmap writes for each qualifying uniform range; it does not claim to remove the classification scan. A case session should benchmark array.Strings.Slice and array.Validity.Slice with globally mixed inputs, uniform selected ranges, and a width sweep using -benchmem plus an allocation profile. The confirming signal is one fewer allocation/op and roughly ceil(rangeLen/8) fewer B/op on triggered slices, with correctness differentials covering all start-bit offsets and both canonical outcomes.

Change to test: Track the first validity state while scanning, and allocate and rebase a bitmap only after the selected range is known to contain both valid and null values. Backfill prior valid bits at that transition while retaining the existing canonical AllValid and AllNull results.

Where it lives

perfloop/btrblocks · array/string.go

Evidence

Timeline