Sapphire Rapids: Add a plan-owned bucket filter for literal alternatives
perfloop/casei · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_109tpvhytb
Verdict
VERIFIED · settled 2026-09-28 · merged as tsenart/casei#22
What happened: The paired measurements met the required improvement.
Hypothesis
The retained three-pass model-143 packet loses the curated English five-literal count row to Hyperscan by about 3.6x. A direct profile of that exact source shape puts 45.15% flat time in `tripleShuftiSkip64` and 93.67% cumulative time in `findFiltered`; the single-literal clean-gap probe is not the owner of this loss.
The current multi-literal path runs a three-byte Shufti filter and then confirms survivors through the shared plan. Hyperscan's multi-literal path uses compiled bucket masks and shifted SIMD state to combine several literal candidates before confirmation. The missing mechanism is a plan-owned multi-literal bucket filter, not a second matcher or benchmark-specific dispatch.
Prove the mechanism on the same curated English five-literal row and its Russian counterpart by counting block survivors and full confirmations, then compare an AVX-512 bucket filter against the current Shufti route. Preserve leftmost and lowest-pattern ordering, exact source widths, Unicode and malformed-byte fallback, and the existing ISA-disabled path. Accept it only if the English loss falls without transferring a material regression to the Russian row or single-literal guards.
Change to test: Add a plan-owned AVX-512 bucket filter for multi-literal ASCII alternatives and retain the existing Shufti, decoded, and portable routes as conservative fallbacks. Make the filter emit only candidate starts and leave exact pattern confirmation and ordering to the shared plan.
Where it lives
perfloop/casei · audit/rebar/runner/main.go
Evidence
Matcher.Each on Rebar curated/02-literal-alternate/sherlock-casei-en: five case-insensitive literals, 725 expected matches over the pinned opensubtitles/en-sampled.txt haystack · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
271923 |
197398 |
−27.2% (−73902) |
−77491 to −72237 |
< −13596 |
PASSED |
MB/s |
3307 |
4556 |
+37.8% (+1250) |
+1202 to +1278 |
≥ −165.3 |
PASSED |
Matcher.Each on Rebar imported/leipzig/tom-sawyer-huckle-fin-insensitive: four case-insensitive literals, 4,152 matches over the pinned 16 MB Leipzig haystack · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
2429196 |
1844860 |
−24.1% (−584454) |
−615908 to −553528 |
< −121460 |
PASSED |
MB/s |
6592 |
8680 |
+31.8% (+2095) |
+1942 to +2176 |
≥ −329.6 |
PASSED |
Matcher.Each on Rebar curated/02-literal-alternate/sherlock-casei-ru: five case-insensitive literals, 971 expected matches over the pinned opensubtitles/ru-sampled.txt haystack · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
495597 |
496322 |
−0.1% (−406) |
−7531 to +6686 |
≤ 24780 |
PASSED |
MB/s |
3169 |
3164 |
+0.1% (+2.59) |
−40.75 to +48.12 |
≥ −158.5 |
PASSED |
Matcher.Each on Rebar imported/sherlock/name-alt5-casei: five case-insensitive literal alternatives over the pinned Sherlock haystack · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
150582 |
122075 |
−19.1% (−28749) |
−30363 to −27083 |
≤ 7529 |
PASSED |
MB/s |
3951 |
4873 |
+23.6% (+931.9) |
+873.3 to +970 |
≥ −197.5 |
PASSED |
Matcher.Each on Rebar curated/01-literal/sherlock-casei-en: one case-insensitive literal over the pinned Sherlock English sample · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
210055 |
210498 |
+0.5% (+996) |
−822 to +3584 |
≤ 10503 |
PASSED |
MB/s |
4281 |
4272 |
−0.5% (−20.22) |
−73.22 to +16.63 |
≥ −214 |
PASSED |
Matcher.Each on Rebar curated/01-literal/sherlock-casei-ru: one case-insensitive literal over the pinned Sherlock Russian sample · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
125454 |
124404 |
−0.7% (−908) |
−3750 to +2080 |
≤ 6273 |
PASSED |
MB/s |
12519 |
12625 |
+0.7% (+91.57) |
−204.3 to +377.2 |
≥ −625.9 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N1_unicode_pair_miss_1_5mb (one Unicode literal, 1.5 MB near-miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.7329 |
0.7347 |
+0.3% (+0.00225) |
−0.0004 to +0.0044 |
≤ 0.03665 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N2_miss_log_1mb (two literal alternatives, 1 MB generated log miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.7982 |
0.8005 |
+0.4% (+0.00285) |
−0.0033 to +0.0158 |
≤ 0.03991 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N512_miss_hazard_64kb (512 fold-hazard literals, 64 KB Cyrillic miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.5133 |
0.5079 |
−1.1% (−0.00575) |
−0.0093 to +0.0009 |
≤ 0.02567 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N512_miss_log_64kb (512 ASCII literals, 64 KB log miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.6625 |
0.6624 |
−0.2% (−0.00115) |
−0.0112 to +0.0035 |
≤ 0.03313 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N5_raw_transition_late_hit_5mb (five Cyrillic alternatives, 5 MB late hit) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.9589 |
0.9612 |
−0.3% (−0.00335) |
−0.0128 to +0.0193 |
≤ 0.04794 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N5_raw_transition_miss_5mb (five Cyrillic alternatives, 5 MB miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.9609 |
0.9619 |
+0.4% (+0.0036) |
−0.0194 to +0.0153 |
≤ 0.04804 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N64_miss_log_64kb (64 ASCII literals, 64 KB log miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.8062 |
0.8073 |
+0.00005 |
−0.0125 to +0.0069 |
≤ 0.04031 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N64_miss_ru_64kb (64 Russian literal alternatives, 64 KB Cyrillic miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.6064 |
0.6064 |
+0.0001 |
−0.0033 to +0.0019 |
≤ 0.03032 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N8_hazard_hit_1mb (eight fold-hazard alternatives, 1 MB planted hit) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.1498 |
0.1507 |
+0.5% (+0.00075) |
−0.0021 to +0.0026 |
≤ 0.00749 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N8_hit_log_1mb (eight ASCII alternatives, 1 MB planted log hits) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.4468 |
0.4457 |
+0.1% (+0.0006) |
−0.0068 to +0.0166 |
≤ 0.02234 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N8_miss_hazard_1mb (eight fold-hazard alternatives, 1 MB prose miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.3214 |
0.3258 |
+1.2% (+0.0039) |
−0.0034 to +0.015 |
≤ 0.01607 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N8_miss_log_1mb (eight ASCII alternatives, 1 MB log miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.4977 |
0.4953 |
−0.00005 |
−0.0092 to +0.003 |
≤ 0.02488 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_N8_miss_ru_1mb (eight Russian literal alternatives, 1 MB Cyrillic miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.6093 |
0.6081 |
−0.1% (−0.0009) |
−0.0341 to +0.0014 |
≤ 0.03046 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_forms4_complete_triple_miss_1mb (complete-triple miss, 1 MB prose) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.3462 |
0.3365 |
−3% (−0.01045) |
−0.0194 to −0.0068 |
≤ 0.01731 |
PASSED |
Matcher.Find on BenchmarkBar/multi/multi_forms8_complete_triple_near_miss_64kb (complete-triple near miss, 64 KB prose) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.6783 |
0.5597 |
−17% (−0.1152) |
−0.1421 to −0.1139 |
≤ 0.03392 |
PASSED |
Matcher.Find on BenchmarkBar/single/ru_latency_miss_1kb (one Russian literal, 1 KB text miss) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
x_vs_best |
0.8717 |
0.8832 |
+1.3% (+0.011) |
+0.0007 to +0.0197 |
≤ 0.04358 |
PASSED |
Checks: 5 of 5 passed. Verification: no defect found.