Sapphire Rapids: Return ASCII singleton matches and widths without decoding
perfloop-oss/casei · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_9maksxty15
Verdict
VERIFIED · settled 2026-10-04 · merged as tsenart/casei#24
What happened: The paired measurements met the required improvement.
Hypothesis
Make four required singleton Rebar operations faster on Sapphire Rapids by confirming ASCII matches and widths together. Remove two decoded walks without changing the probe or Unicode contract. Official savings remain unproved.
Required rows are curated/01 English count and imported Sherlock, Holmes and Sherlock Holmes count-spans: full899232/594933-byte fixtures, totals522 matches and816/2802/1440 source bytes. At514b165 Each reaches eachASCIIProbe: token confirmation returns only Boolean, then matcherMatchEnd walks again. Curated census records538 ASCII confirmations,522 matches and7830 width-decoded units.
Confirmation/end cost0.77+0.50 of3.33 sampled seconds beside1.67 scanning. Curated exploration69.5-69.9us->48.2-48.9us suggests30%; imported arms9-15%,19-21%,15%. Provisional required removals60.6%,30.6%,29.2%,40.0% are larger: this is a confirmation/end prerequisite. Masks must represent exact plan classes, not weaken folding. No public API or new screen; report setup/memory.
Complete-ASCII candidates can prove match and width with word checks cheaper than two token/source walks.
Pair all four required complete BenchmarkRebar operations on model143 against pristine baseline; each must repeatably reduce ns/op with independent positions/IDs/order/widths and stated totals. Wrong fold/end, overread, neutral target or guard slowdown rejects it. Preserve recovery/malformed-byte handling, exercise fallback, preserve BenchmarkBar field wins and report setup/retained costs.
Independent flag-selected arms kept the screen unchanged, but were not pristine official baseline builds. Imported timings are separate windows; no Sherlock-only CPU profile. This does not revive the bundled one-load proposal or claim every singleton benefits. Integrated/production evidence remains absent.
Change to test: On Sapphire Rapids, certify complete bounded ASCII singleton candidates from the plan's token classes and return their exact widths. Preserve the probe/start recovery and decoded confirmation plus endpoint recovery for unknown windows.
Where it lives
perfloop-oss/casei · audit/rebar/runner/main.go
Evidence
BenchmarkRebar/curated-01-literal-sherlock-casei-en count, 899232-byte fixture · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
70452 |
48060 |
−31.8% (−22392) |
−22620 to −21867 |
< −3523 |
PASSED |
MB/s |
12764 |
18711 |
+46.6% (+5943) |
+5759 to +6032 |
≥ −638.2 |
PASSED |
B/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
allocs/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
BenchmarkRebar/imported-sherlock-name-sherlock-casei count-spans, 594933-byte fixture · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
22747 |
20784 |
−8.7% (−1970) |
−2172 to −1784 |
< −1137 |
PASSED |
MB/s |
26154 |
28625 |
+9.5% (+2487) |
+2226 to +2760 |
≥ −1308 |
PASSED |
B/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
allocs/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
BenchmarkRebar/imported-sherlock-name-holmes-casei count-spans, 594933-byte fixture · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
35802 |
29997 |
−16% (−5712) |
−6072 to −5560 |
< −1790 |
PASSED |
MB/s |
16618 |
19834 |
+19% (+3160) |
+3047 to +3365 |
≥ −830.9 |
PASSED |
B/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
allocs/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
BenchmarkRebar/imported-sherlock-name-sherlock-holmes-casei count-spans, 594933-byte fixture · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
25102 |
21438 |
−14.6% (−3660) |
−3807 to −3469 |
< −1255 |
PASSED |
MB/s |
23701 |
27751 |
+17.1% (+4051) |
+3824 to +4226 |
≥ −1185 |
PASSED |
B/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
allocs/op |
0 |
0 |
0 |
0 to 0 |
≤ 0 |
PASSED |
BenchmarkBar complete 38-row field-win guard on model 143 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
benchmarkbar_all_rows_winning |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
benchmarkbar_worst_x_vs_best |
0.9691 |
0.9706 |
+0.5% (+0.0047) |
−0.0084 to +0.0147 |
≤ 0.04845 |
PASSED |
benchmarkbar_rows_below_one |
38 |
38 |
0 |
0 to 0 |
≥ 0 |
PASSED |
benchmarkbar_min_entrants |
5 |
5 |
0 |
0 to 0 |
≥ −0.25 |
PASSED |
benchmarkbar_candidate_vector_bits |
512 |
512 |
0 |
0 to 0 |
≥ −25.6 |
PASSED |
benchmarkbar_vectorscan_vector_bits |
512 |
512 |
0 |
0 to 0 |
≥ −25.6 |
PASSED |
benchmarkbar_vectorscan_vbmi |
1 |
1 |
0 |
0 to 0 |
≥ −0.05 |
PASSED |
Checks: 10 of 10 passed. Verification: no defect found.