Avoid dynamic reflect.ValueOf allocations per row inside writeRowsFuncOfTime loop

perfloop/parquet-go · ALLOCATION HOT LOOP

https://perfloop.ai/t/oss/case_fxbwsnnmrj

Verdict

VERIFIED · settled 2026-06-19 · merged as parquet-go/parquet-go#542

What happened: Successfully optimized parquet.writeRowsFuncOfTime inside column_buffer_write.go by redesigning the row writing loop to eliminate all dynamic reflect.ValueOf heap allocations per row, drastically reducing latency and memory allocations.

Specifically: 1. Replaced the row-by-row Unix time conversion loop with a pre-pass batch conversion that translates the entire sparse.TimeArray into a pre-allocated int64 slice. 2. Reused wideIntBufPool to pool the converted []int64 slice, completely avoiding slice allocation overhead. 3. Hoisted the TimeUnit unit switch entirely out of the loop so it evaluates once per batch instead of once per row. 4. Changed the leaf writer calls to write contiguous non-zero or null slices at once via a single sparse.MakeInt64Array(buf.values[i:j]).UnsafeArray() call, rather than calling writeRows individually with single-element arrays and dynamic reflection pointers.

The optimization was evaluated using the Go BenchmarkGenericWriter/timeColumn microbenchmark at 1000 rows. Results: - allocs/op: decreased from 1001 to exactly 1 (a 99.9% reduction in heap allocations). - ns/op: decreased from 50002 ns/op to 13224 ns/op (a 73.6% latency reduction, or 3.78x throughput speedup). All correctness tests and column-writing unit tests pass without any regressions.

Hypothesis

The function parquet.writeRowsFuncOfTime writes a time column to a parquet file by processing rows individually. Under the Parquet Write Workload, for every row, it converts the time value to int64, then boxes that value through reflect.ValueOf and makeArray to pass a length-1 sparse array to the leaf writer. We structurally hypothesize that this reflect.ValueOf boxing allocates on the heap once per row, resulting in ~1000 allocations to write 1000 rows, creating high GC pressure and CPU overhead, while repeating all level bookkeeping per row. Converting the batch to a pooled []int64 once and calling the leaf writer with a single sparse.MakeInt64Array is expected to avoid dynamic reflection boxing entirely. The proof targets a case session would measure to verify this improvement are allocs/op and write latency using a Go microbenchmark on a required time column.

Change to test: In parquet.writeRowsFuncOfTime (inside column_buffer_write.go), edit the row writing loop to convert the batch of time.Time to a pooled []int64 slice first, hoisting the unit switch out of the loop, and hand required or levelled runs to the leaf writer using a single sparse.MakeInt64Array call instead of calling writeRow per row with individual dynamic reflect.ValueOf boxing.

Where it lives

perfloop/parquet-go · parquet.go

Evidence

Timeline