Win all 18 Rebar literal workloads with one streaming plan
perfloop/casei · DATA PARALLEL GAP
https://perfloop.ai/t/oss/case_7241g226cv
Verdict
VERIFIED · settled 2026-09-27 · merged as tsenart/casei#21
What happened: The paired measurements met the required improvement.
Hypothesis
I inspected `casei.Matcher.Each` in `matcher.go:61-93`. A usable non-empty `rawByteMulti` plan takes `eachRawByteFixedAnchored`; otherwise `Each` loops over `haystack[at:]` and calls `findWithWidth` at line 69 until no match remains. It then adjusts the match start and advances `at` with the existing width, zero-width, and non-overlap rules.
The source establishes a repeated generic enumeration step: every yielded generic match re-enters `findWithWidth` on the remaining suffix. It does not establish that this restart cost dominates any benchmark. The hypothesis is that eligible plan-owned scans can classify contiguous input in wider blocks while preserving the existing exact confirmation and enumeration semantics.
Confirm with differential tests that compare ordered matches, byte offsets, early-yield termination, zero-width advancement, overlap behavior, and Unicode simple folding. Measure end-to-end time or cycles per haystack byte across realistic haystack sizes and match densities on the Rebar enumeration path, including the current generic plans; the proposed path must show a reduced recurring cost without changing results.
Change to test: Replace repeated findWithWidth restarts in Matcher.Each with one plan-owned block scan that carries position, survivor masks, confirmed widths, tags, and non-overlap state across matches while routing non-ASCII candidates through the existing exact plan to preserve Unicode simple folding. Keep one matcher and one plan without per-pattern loops or benchmark-specific dispatch.
Where it lives
perfloop/casei · audit/rebar/runner/main.go
Evidence
Rebar `Matcher.Each` count on the pinned 16,013,977-byte Leipzig text: one `Twain` literal, 965 matches · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
670692 |
579482 |
−13.7% (−91637) |
−98806 to −78688 |
< −33535 |
PASSED |
MB/s |
23877 |
27635 |
+16% (+3821) |
+3177 to +4085 |
> 1194 |
PASSED |
Rebar `Matcher.Each` count on the pinned 899,232-byte English subtitles sample: `Sherlock Holmes`, 522 matches · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
80557 |
70979 |
−12.1% (−9762) |
−10307 to −8832 |
< −4028 |
PASSED |
MB/s |
11163 |
12669 |
+13.6% (+1523) |
+1379 to +1620 |
> 558.1 |
PASSED |
Rebar `Matcher.Each` count-spans on the pinned 594,933-byte Sherlock text: one `the` literal, 23,961 matches · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
309061 |
210861 |
−32% (−98753) |
−103692 to −96143 |
≤ 15453 |
PASSED |
MB/s |
1925 |
2821 |
+47% (+904.9) |
+875.8 to +923 |
≥ −96.25 |
PASSED |
Rebar `Matcher.Each` count on the pinned 1,570,556-byte Russian subtitles sample: `Шерлок Холмс`, 746 matches · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
123719 |
124881 |
+0.7% (+917.5) |
−738 to +3355 |
≤ 6186 |
PASSED |
MB/s |
12695 |
12576 |
−0.7% (−93.9) |
−334.2 to +76.41 |
≥ −634.7 |
PASSED |
Checks: 9 of 9 passed. Verification: no defect found.