Sapphire Rapids: Stream three-name counts through root candidates

perfloop-oss/casei · INEFFICIENT ALGORITHM

https://perfloop.ai/t/oss/case_qte7kjr85q

Verdict

VERIFIED · settled 2026-10-06 · merged as tsenart/casei#26

What happened: The paired measurements met the required improvement.

Hypothesis

Make the required Sherlock/Holmes/Watson count-spans faster on Sapphire Rapids by selecting the existing root iterator. Results and widths must stay identical; the gain remains unproved.

The594933-byte fixture expects4104 source bytes, not matches. Its bucket already compiles; the four-pattern minimum alone keeps full Each on generic suffix searches. Baseline findFiltered has3.18/3.52 cumulative sampled seconds including1.76 scanning; the change removes restarts, not that whole cost.

Gate-only exploration changed122-124us to88-89us, about28%. The provisional2.063 ratio needs51.5% removal, leaving roughly1.49x: an iterator prerequisite for later confirmation/screen work. case_ttc9g9ze4t/cand_vk3vwyqvzx Verification establishes shipped four/five root Each, not this extension or a fused screen. Its gains are already baseline. No new filter or public API.

Three-pattern anchored replay is selective enough to cost less than repeated generic searches on this fixture.

Pair the complete required operation on model143 against pristine baseline; require repeatable ns/op reduction and independent matches/IDs/order/widths totaling4104. Neutral/slower timing or discrepancy rejects it. Guard existing four/five English/Russian counts, early stop, nested/concurrent use, tails and exercised fallback; preserve existing BenchmarkBar field wins.

The failing eligibility assertion encoded the old exclusion and was revised/rerun; it was not a semantic discrepancy or independent proof. Official pairs and integrated field evidence remain absent. Report setup/retained costs; guest affinity lacks physical exclusivity.

Change to test: On Sapphire Rapids, lower root-bucket Each's minimum from four patterns to three. Retain its upper bound, other guards, source-order iterator and exact trie confirmation.

Where it lives

perfloop-oss/casei · audit/rebar/runner/main.go

Evidence

Rebar imported/sherlock/name-alt5-casei count-spans on a retained Matcher over the 594,933-byte Sherlock fixture (Sherlock|Holmes|Watson) · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 122281 92425 −24.2% (−29653) −31129 to −28778 < −6114 PASSED
MB/s 4865 6437 +32.3% (+1570) +1513 to +1623 ≥ −243.3 PASSED

BenchmarkBar per-row x_vs_best win-status guard across all 38 Sapphire Rapids model-143 rows · 10 sample pairs

metric baseline candidate paired median change confidence range required result
BenchmarkBar/multi/multi_N1_unicode_pair_miss_1_5mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N2_miss_log_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N512_miss_hazard_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N512_miss_log_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N5_raw_transition_late_hit_5mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N5_raw_transition_miss_5mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N64_miss_log_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N64_miss_ru_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N8_hazard_hit_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N8_hit_log_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N8_miss_hazard_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N8_miss_log_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_N8_miss_ru_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_forms4_complete_triple_miss_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/multi/multi_forms8_complete_triple_near_miss_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/code_hit_brackets_256kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/code_miss_256kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/kelvin_hazard_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/latency_match_end_1kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/latency_match_mid_1kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/latency_match_start_1kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/latency_miss_1kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_hit_sparse_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_miss_1kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_miss_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_miss_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_needle16_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_needle32_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_needle3_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/log_needle8_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/periodic_miss_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/prose_hit_dense_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/prose_miss_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/ru_hit_sparse_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/ru_latency_miss_1kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/ru_miss_1mb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/samechar_miss_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED
BenchmarkBar/single/torture_miss_64kb/x_vs_best_win 1 1 0 0 to 0 ≥ 0 PASSED

Checks: 11 of 11 passed. Verification: no defect found.

Timeline