Exact-case triple Shufti tables on Sapphire Rapids

perfloop/casei · DATA PARALLEL GAP

https://perfloop.ai/t/oss/case_qwm2h29mxw

Verdict

VERIFIED · settled 2026-09-04 · merged as tsenart/casei#17

What happened: The paired measurements met the required improvement.

Hypothesis

The current generic triple predicate folds only positions proved to be ASCII letters, while the dormant table projection folds every input byte and can therefore admit non-letter aliases before the common decoded plan rejects them. This proposal replaces that coarse representation with exact per-slot case membership while retaining the existing six-table SIMD scan and the common plan's authority for Unicode simple folding, malformed bytes, leftmost order, lowest-pattern ties, and returned widths. Proof target: enumerate all 256^3 input triples against the generic predicate, then on Intel Sapphire Rapids model 143 require an explicit-table route/stop counter and order-rotated generic-versus-normalized-versus-explicit `Matcher.Each` comparison; any predicate mismatch, unchanged alias stops, or no repeatable explicit-over-normalized time reduction falsifies it.

Change to test: For Intel Sapphire Rapids model 143, replace the normalized complete-root triple Shufti representation with one per-triple-slot table that records each fold-marked ASCII byte's two exact cases and every unmarked byte raw, and have its AVX-512 loop consume raw input bytes without the three unconditional bit-five ORs.

Where it lives

perfloop/casei · matcher.go

Evidence

1 MiB alias-heavy ASCII Matcher.Each with eight triple forms, a same-prefix longer form, and a Unicode pattern on Intel Sapphire Rapids model 143 · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 349904 124428 −64.4% (−225331) −228891 to −223411 < −17495 PASSED
MB/s 2997 8427 +181.2% (+5429) +5366 to +5465 > 149.8 PASSED
B/op 0 0 0 0 to 0 ≤ 0 PASSED
allocs/op 0 0 0 0 to 0 ≤ 0 PASSED

Complete 38-row native BenchmarkBar field comparison, including UTF-8 and complete-triple miss rows, on Intel Sapphire Rapids model 143 · 10 sample pairs

metric baseline candidate paired median change confidence range required result
benchmarkbar_all_rows_winning 0 1 +1 +1 to +1 > 0 PASSED
benchmarkbar_worst_x_vs_best 2.073 0.9707 −53% (−1.099) −1.103 to −1.086 < −0.1036 PASSED
benchmarkbar_rows_below_one 37 38 +2.7% (+1) +1 to +1 ≥ 0 PASSED
benchmarkbar_min_entrants 5 5 0 0 to 0 ≥ 0 PASSED
benchmarkbar_candidate_vector_bits 512 512 0 0 to 0 ≥ 0 PASSED
benchmarkbar_vectorscan_vector_bits 512 512 0 0 to 0 ≥ 0 PASSED
benchmarkbar_vectorscan_vbmi 1 1 0 0 to 0 ≥ 0 PASSED

Checks: 12 of 12 passed. Verification: no defect found.

Timeline