Counting scatter for dictionary index sorts
perfloop/arrow-rs · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_a0kjk42gzp
Verdict
VERIFIED · settled 2026-08-07
What happened: The paired measurements met the required improvement.
Hypothesis
Source inspection at arrow-ord/src/sort.rs:556-576 shows `sort_dictionary` is called by `sort_to_indices` at line 306 on the modeled Column sort invocation path and performs `sort_unstable_by` over index/rank pairs. The proposed counting-scatter shape is a structural hypothesis: source establishes the comparison-sort work and that dictionary values provide a cardinality K, but it does not establish production N/K distributions or that this function dominates elapsed time. The proof target is a benchmark sweeping realistic N and K/N ratios on dictionary arrays, including both sides of the proposed K-versus-N fallback threshold, with a CPU profile attributing time to this function; verify output ordering, null placement, and sort options against the existing implementation.
Change to test: Replace comparison sorting of `(u32 index, rank)` pairs with a guarded counting scatter: count dense ranks, prefix-sum counts, then scatter indices; use the exact dictionary cardinality as K and retain comparison sort when K is not favorable relative to N.
Where it lives
perfloop/arrow-rs · arrow-ord/src/sort.rs
Evidence
permuted dictionary string sort to indices, n=4096 k=256 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
48731 |
12412 |
−74.5% (−36327) |
−36500 to −35827 |
< 0 |
PASSED |
permuted dictionary string sort to indices, n=65536 k=4096 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
1006540 |
539478 |
−46.5% (−467948) |
−483070 to −446952 |
< 0 |
PASSED |
permuted dictionary string sort to indices, n=65536 k=131072 (K=2N) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
2091726 |
1454605 |
−30.2% (−632116) |
−657761 to −610537 |
< 0 |
PASSED |
permuted dictionary string sort to indices, n=65536 k=524288 (K=8N) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
5397514 |
5145035 |
−4.7% (−254849) |
−406003 to −131290 |
< 0 |
PASSED |
permuted dictionary string sort to indices, n=65536 k=1048576 (K=16N comparison fallback) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
10742640 |
10747997 |
+0.2% (+25066) |
−291333 to +241720 |
≤ 500000 |
PASSED |
sorted-key dictionary string sort to indices, n=4096 k=32768 (K=8N) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
220071 |
224927 |
+2.2% (+4914) |
+2951 to +6376 |
≤ 20000 |
PASSED |
sorted-key dictionary string sort to indices, n=65536 k=524288 (K=8N) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
4203584 |
4244693 |
+1.9% (+79180) |
−73875 to +130998 |
≤ 200000 |
PASSED |
sparse-use dictionary string sort to indices, n=4096 k=32768 used=4 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
229094 |
234029 |
+2.2% (+4963) |
−2970 to +7596 |
≤ 10000 |
PASSED |
null-heavy integer dictionary sort to indices, n=65536 k=65536 valid=1/8 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
269178 |
282507 |
+5.1% (+13806) |
+9633 to +16516 |
≤ 30000 |
PASSED |
nearly ordered dictionary string sort to indices, n=65536 k=16 with one inversion · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
337971 |
343550 |
+1.8% (+5971) |
+3170 to +8682 |
≤ 50000 |
PASSED |
dictionary sorting CPU profile, permuted n=65536 k=524288 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
dictionary_pair_sort_cpu_pct |
29.16 |
0 |
−100% (−29.16) |
−32.88 to −26.01 |
< 0 |
PASSED |
count_scatter_cpu_pct |
0 |
28.47 |
+28.47 |
+26.18 to +32.14 |
> 0 |
PASSED |
Checks: 6 of 6 passed. Verification: no defect found.