Sapphire Rapids: Stream three-name counts through root candidates
perfloop-oss/casei · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_qte7kjr85q
Verdict
VERIFIED · settled 2026-10-06 · merged as tsenart/casei#26
What happened: The paired measurements met the required improvement.
Hypothesis
Make the required Sherlock/Holmes/Watson count-spans faster on Sapphire Rapids by selecting the existing root iterator. Results and widths must stay identical; the gain remains unproved.
The594933-byte fixture expects4104 source bytes, not matches. Its bucket already compiles; the four-pattern minimum alone keeps full Each on generic suffix searches. Baseline findFiltered has3.18/3.52 cumulative sampled seconds including1.76 scanning; the change removes restarts, not that whole cost.
Gate-only exploration changed122-124us to88-89us, about28%. The provisional2.063 ratio needs51.5% removal, leaving roughly1.49x: an iterator prerequisite for later confirmation/screen work. case_ttc9g9ze4t/cand_vk3vwyqvzx Verification establishes shipped four/five root Each, not this extension or a fused screen. Its gains are already baseline. No new filter or public API.
Three-pattern anchored replay is selective enough to cost less than repeated generic searches on this fixture.
Pair the complete required operation on model143 against pristine baseline; require repeatable ns/op reduction and independent matches/IDs/order/widths totaling4104. Neutral/slower timing or discrepancy rejects it. Guard existing four/five English/Russian counts, early stop, nested/concurrent use, tails and exercised fallback; preserve existing BenchmarkBar field wins.
The failing eligibility assertion encoded the old exclusion and was revised/rerun; it was not a semantic discrepancy or independent proof. Official pairs and integrated field evidence remain absent. Report setup/retained costs; guest affinity lacks physical exclusivity.
Change to test: On Sapphire Rapids, lower root-bucket Each's minimum from four patterns to three. Retain its upper bound, other guards, source-order iterator and exact trie confirmation.
Where it lives
perfloop-oss/casei · audit/rebar/runner/main.go
Evidence
Rebar imported/sherlock/name-alt5-casei count-spans on a retained Matcher over the 594,933-byte Sherlock fixture (Sherlock|Holmes|Watson) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
122281 |
92425 |
−24.2% (−29653) |
−31129 to −28778 |
< −6114 |
PASSED |
MB/s |
4865 |
6437 |
+32.3% (+1570) |
+1513 to +1623 |
≥ −243.3 |
PASSED |
BenchmarkBar per-row x_vs_best win-status guard across all 38 Sapphire Rapids model-143 rows · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
BenchmarkBar/multi/multi_N1_unicode_pair_miss_1_5mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N2_miss_log_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N512_miss_hazard_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N512_miss_log_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N5_raw_transition_late_hit_5mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N5_raw_transition_miss_5mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N64_miss_log_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N64_miss_ru_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N8_hazard_hit_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N8_hit_log_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N8_miss_hazard_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N8_miss_log_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_N8_miss_ru_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_forms4_complete_triple_miss_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/multi/multi_forms8_complete_triple_near_miss_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/code_hit_brackets_256kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/code_miss_256kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/kelvin_hazard_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/latency_match_end_1kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/latency_match_mid_1kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/latency_match_start_1kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/latency_miss_1kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_hit_sparse_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_miss_1kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_miss_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_miss_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_needle16_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_needle32_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_needle3_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/log_needle8_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/periodic_miss_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/prose_hit_dense_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/prose_miss_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/ru_hit_sparse_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/ru_latency_miss_1kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/ru_miss_1mb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/samechar_miss_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
BenchmarkBar/single/torture_miss_64kb/x_vs_best_win |
1 |
1 |
0 |
0 to 0 |
≥ 0 |
PASSED |
Checks: 11 of 11 passed. Verification: no defect found.