Exact-case triple Shufti tables on Sapphire Rapids
perfloop/casei · DATA PARALLEL GAP
https://perfloop.ai/t/oss/case_qwm2h29mxw
Verdict
VERIFIED · settled 2026-09-04 · merged as tsenart/casei#17
What happened: The paired measurements met the required improvement.
Hypothesis
The current generic triple predicate folds only positions proved to be ASCII letters, while the dormant table projection folds every input byte and can therefore admit non-letter aliases before the common decoded plan rejects them. This proposal replaces that coarse representation with exact per-slot case membership while retaining the existing six-table SIMD scan and the common plan's authority for Unicode simple folding, malformed bytes, leftmost order, lowest-pattern ties, and returned widths. Proof target: enumerate all 256^3 input triples against the generic predicate, then on Intel Sapphire Rapids model 143 require an explicit-table route/stop counter and order-rotated generic-versus-normalized-versus-explicit `Matcher.Each` comparison; any predicate mismatch, unchanged alias stops, or no repeatable explicit-over-normalized time reduction falsifies it.
Change to test: For Intel Sapphire Rapids model 143, replace the normalized complete-root triple Shufti representation with one per-triple-slot table that records each fold-marked ASCII byte's two exact cases and every unmarked byte raw, and have its AVX-512 loop consume raw input bytes without the three unconditional bit-five ORs.
Where it lives
perfloop/casei · matcher.go
Evidence
1 MiB alias-heavy ASCII Matcher.Each with eight triple forms, a same-prefix longer form, and a Unicode pattern on Intel Sapphire Rapids model 143 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
349904 |
124428 |
−64.4% (−225331) |
−228891 to −223411 |
< −17495 |
PASSED |
MB/s |
2997 |
8427 |
+181.2% (+5429) |
+5366 to +5465 |
> 149.8 |
PASSED |
B/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
allocs/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
Complete 38-row native BenchmarkBar field comparison, including UTF-8 and complete-triple miss rows, on Intel Sapphire Rapids model 143 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
benchmarkbar_all_rows_winning |
0 |
1 |
+1 |
+1 to +1 |
> 0 |
PASSED |
benchmarkbar_worst_x_vs_best |
2.073 |
0.9707 |
−53% (−1.099) |
−1.103 to −1.086 |
< −0.1036 |
PASSED |
benchmarkbar_rows_below_one |
37 |
38 |
+2.7% (+1) |
+1 to +1 |
≥ 0 |
PASSED |
benchmarkbar_min_entrants |
5 |
5 |
0 |
0 to 0 |
≥ 0 |
PASSED |
benchmarkbar_candidate_vector_bits |
512 |
512 |
0 |
0 to 0 |
≥ 0 |
PASSED |
benchmarkbar_vectorscan_vector_bits |
512 |
512 |
0 |
0 to 0 |
≥ 0 |
PASSED |
benchmarkbar_vectorscan_vbmi |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
Checks: 12 of 12 passed. Verification: no defect found.