Combined one-pass AVX-512 Int64 bounds kernel at default page size
perfloop/parquet-go · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_x674012fms
Verdict
VERIFIED · settled 2026-07-21 · merged as parquet-go/parquet-go#579
What happened: Validated. The lower-boundary predicate change removed the previously significant small-fallback cost while retaining the interval kernel’s win. At the default 32,113-value INT64 Page.Bounds shape, ns/op fell from 3983.5 to 2915 (26.823145475084725% lower; Mann–Whitney U p=0.00001082508822446903; n=10 per arm), exceeding the 20% goal with B/op and allocs/op remaining zero. The newly dispatched upper-window shape also improved from 18856 to 12349.5 ns/op (34.506257955027575% lower; p=0.00001082508822446903), while the below-window and above-window guards held. The small-fallback median was effectively unchanged at 84.495 to 84.99 ns/op; its 0.5858334812710696% regression was unconfirmed (p=0.6842105263157897). All six focused differential, page-statistics, AVX-512 lane, and disabled-runtime checks passed. The retained implementation already loads each 32-value vector block once and uses independent min/max accumulation chains, while the existing paths remain outside the dispatch interval; no smaller further production change is justified by the measured structural ceiling.
Hypothesis
The accepted Parquet Write Workload reaches parquet.boundsInt64 on path_tztwh2qk7w: Writer.WriteRows → ColumnWriter.writeDataPage → makePageStatistics → Page.Bounds → int64Page.bounds → boundsInt64. The indexed locality is page_bounds_amd64.go:76-83. The submitter reports that, around the default 32,113-value (~256 KiB) INT64 page, this helper performs separate AVX-512 min and max reductions, and reports a 27.4% median Page.Bounds reduction from a combined four-accumulator kernel with unchanged allocations; this performance result remains to be independently verified. Proof target: benchmark INT64 Page.Bounds at the default page size and sweep on both sides of the proposed combined-kernel threshold, comparing the current two scans with the combined four-accumulator kernel, plus differential correctness checks with default, AVX-512-disabled, and pure-Go paths.
Change to test: Add a one-pass four-accumulator AVX-512 INT64 min/max kernel in page_bounds_amd64.s and dispatch boundsInt64 to it for the default-page-size window when AVX-512VL is available; retain the existing scalar combined path outside that window. In writeDataPage, calculate the page bound once and share it between makePageStatistics and the column index rather than issuing two Page.Bounds calls.
Where it lives
perfloop/parquet-go · writer.go