Avoid assembly call overhead for small slices in SumUint16

perfloop/asm · INEFFICIENT ALGORITHM

https://perfloop.ai/t/oss/case_117zqcpg3b

Verdict

VERIFIED · settled 2026-07-30

What happened: The paired measurements met the required improvement.

Hypothesis

Calling the assembly function sumUint16 incurs a function-call overhead of several nanoseconds because Go assembly functions cannot be inlined by the compiler. Under the high-frequency 'Analytical Vector Byte Operations' workload, these summations run inside hot per-item loops. When slice sizes are small (len < 64), the non-inlinable call overhead outweighs the AVX2 vectorization benefits, leading to significant CPU performance degradation. Implementing an inline Go check and calling sumUint16Generic directly for small lengths avoids this overhead entirely. The uncertainty is the exact slice size distribution in production, which can be grounded via microbenchmarks. Anchor loop context: anchored at the enclosing function's first modeled call; the loop context of the inline mechanism is not modeled.

Change to test: Add a fast-path check in SumUint16 to execute sumUint16Generic inline when len(x) < 64, avoiding non-inlinable assembly call overhead for small inputs. Verify the improvement via go test -bench.

Where it lives

perfloop/asm · slices/sums.go

Evidence

SumUint16 on a 63-element slice · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 77.05 36.47 −52.8% (−40.68) −41.35 to −39.32 < 0 PASSED
B/op 0 0 0 0 to 0 ≤ 0 PASSED
allocs/op 0 0 0 0 to 0 ≤ 0 PASSED

Checks: 3 of 3 passed. Verification: no defect found.

Timeline