Avoid assembly call overhead for small slices in SumUint16
perfloop/asm · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_117zqcpg3b
Verdict
VERIFIED · settled 2026-07-30
What happened: The paired measurements met the required improvement.
Hypothesis
Calling the assembly function sumUint16 incurs a function-call overhead of several nanoseconds because Go assembly functions cannot be inlined by the compiler. Under the high-frequency 'Analytical Vector Byte Operations' workload, these summations run inside hot per-item loops. When slice sizes are small (len < 64), the non-inlinable call overhead outweighs the AVX2 vectorization benefits, leading to significant CPU performance degradation. Implementing an inline Go check and calling sumUint16Generic directly for small lengths avoids this overhead entirely. The uncertainty is the exact slice size distribution in production, which can be grounded via microbenchmarks. Anchor loop context: anchored at the enclosing function's first modeled call; the loop context of the inline mechanism is not modeled.
Change to test: Add a fast-path check in SumUint16 to execute sumUint16Generic inline when len(x) < 64, avoiding non-inlinable assembly call overhead for small inputs. Verify the improvement via go test -bench.
Where it lives
perfloop/asm · slices/sums.go
Evidence
SumUint16 on a 63-element slice · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
77.05 |
36.47 |
−52.8% (−40.68) |
−41.35 to −39.32 |
< 0 |
PASSED |
B/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
allocs/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
Checks: 3 of 3 passed. Verification: no defect found.
Timeline
2026-06-13· Case opened