Raw-byte multi plans break the lowest-pattern-ID rule on tied starts
perfloop-oss/casei · UNCATALOGUED MECHANISM
https://perfloop.ai/t/oss/case_t755qj6fd4
Verdict
VERIFIED · settled 2026-10-08 · merged as tsenart/casei#30
What happened: The assertion is violated on the comparison and satisfied with this change.
Hypothesis
casei master 4f42896c17 (also e7d8cac) violates its own semantics. NewMatcher([]string{"σοφος","σοφο"}).Find("σοφος") returns {Pattern:1 Start:0} and Each yields Pattern 1 with width 8, on native arm64 and on GOARCH=amd64 with AVX-512 off. Both patterns match at byte 0, so the leftmost-then-lowest-ID rule requires Pattern 0 with width 10. A random differential test over 100k pattern sets found 17 Find and 23 Each mismatches against refFind/refEach, all on rawByteMulti plans. git bisect names a1333ef ("perf: add tagged raw-byte multi-anchor transitions", 2026-08-24) as the first bad commit. The Initiative outcome requires the semantic suite (lowest-pattern-ID, exact widths) to pass at the winning commit, so this blocks completion. The Russian Rebar rows use this route.
The accepted checkout is 4f42896c178aaf032ca0161a1e1225cbb52a27ef. The public contract requires leftmost starts, lowest-ID ties and the selected occurrence's exact source width (matcher.go:10-11,59-63; README.md:177-185). At raw_byte.go:712-730, rawByteMatchAt returns immediately at the first completed anchored terminal, at lines 722-724. A shorter higher-ID prefix can therefore end confirmation before a longer lower-ID pattern completes at that same start. eachRawByteFixedAnchored calls this confirmer at raw_byte.go:808 and compares only its returned match at lines 812-814. Repeating that confirmation at the same start cannot recover the unexamined longer terminal.
Public Each selects this shared route for a non-empty usable rawByteMulti plan (matcher.go:67-73). Public Find selects findRawByteFixedAnchored for usable rawByteMulti plans without the long-input origin gate (plan.go:2519-2523), and that specialization calls the same enumerator (raw_byte.go:736-744). The existing supported operation graph has been extended by one cited call from the raw-byte enumerator to rawByteMatchAt, the behavior owner in the submitted file.
A focused disposable TestPerfloopRawByteTiedStartProbe independently confirmed that the submitted Greek patterns compile a usable rawByteMulti plan. It reproduced the Find and Each discrepancies against refFind/refEach on native AMD64 with AVX-512F/BW/VBMI enabled, with AVX-512 disabled, and with all CPU features disabled. The first two runs also reproduced the discrepancy after a 128-byte prefix. Direct rawByteMatchAt returned Pattern 1 and width 8. The temporary probe was removed without changing tracked source. The native-arm64 observation, e7d8cac result, 100k-set mismatch totals and bisect attribution above remain submitter reports, not independent findings of this run. No Russian-row speed was measured here.
The independent repair proof is the supplied deterministic regression and randomized shared-prefix rawByteMulti differential checks through public Find and Each against refFind/refEach (matcher_test.go:14-29,39-65), including ordered IDs, starts, exact widths and non-overlap under the supported dispatch modes. Any remaining oracle discrepancy rejects the repair. The requested Russian speed guard can be checked with the complete pinned Russian operations in audit/rebar/runner/testdata/rows.tsv:2,4,7-8,18. BenchmarkRebar retains a compiled Matcher outside timing and consumes the full count or source-width total (audit/rebar/runner/bench_test.go:73-115; main.go:108-119). These are regression controls, not evidence of a proposed speedup.
This is an architecture-independent correction to the shared Go confirmation authority, not several target-specific optimizations. It adds no consumer-visible operation, type, configuration or protocol. REBAR.md:28-44 describes continuing Initiative work, and REBAR.md:68-90 documents the accepted shared raw-byte route. That is source-revision maintenance context, not a claim of maintainer approval of a patch.
Change to test: Make the raw-byte multi-pattern route pick the lowest pattern ID among all patterns that match at the leftmost start, as the reference (refFind/refEach) does, on every path (VBMI, AVX2-off and portable). Add the failing case below as a regression test, plus a differential test of Find/Each against refFind/refEach on random rawByteMulti plans that share prefixes. Keep the speed of the Russian Rebar rows.
Where it lives
perfloop-oss/casei · raw_byte.go
Evidence
The assertion is violated on the comparison and satisfied with this change: `For the deterministic Greek lowest-ID tie, Kelvin source-width tie, long-origin tie, EOF-resumption case, and 128 seeded usable shared-prefix plans in TestRawByteMultiSharedPrefixDifferential, together with the variable-width EOF, origin-gate, same-start-tie, earlier-start and alignment cases in TestRawByteMultiVariableWidthEOFDifferential, on usable rawByteMulti plans with valid UTF-8 patterns and possibly opaque-byte haystacks, public Matcher.Find matches refFind on found status, leftmost byte start and lowest pattern ID, and public Matcher.Each matches refEach on the complete ordered non-overlapping results and exact source-byte widths under default, AVX-512-disabled, and AVX-512/AVX2-disabled dispatch.`
Russian Rebar curated Sherlock single-literal count row · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
122260 |
122260 |
+0.1% (+147) |
−1907 to +1237 |
≤ 6113 |
PASSED |
MB/s |
12846 |
12846 |
−0.1% (−15.51) |
−128.8 to +203.3 |
≥ −642.3 |
PASSED |
Russian Rebar curated five-literal alternate count row · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
511019 |
512918 |
+0.2% (+1164) |
−2458 to +6284 |
≤ 25551 |
PASSED |
MB/s |
3073 |
3062 |
−0.2% (−7.005) |
−37.72 to +14.74 |
≥ −153.7 |
PASSED |
Russian Rebar Hyperscan no-SOM single-literal count row · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
24076 |
24044 |
−0.4% (−97) |
−418 to +178 |
≤ 1204 |
PASSED |
MB/s |
25479 |
25513 |
+0.4% (+103.1) |
−188.7 to +440.1 |
≥ −1274 |
PASSED |
Russian Rebar Hyperscan SOM single-literal span row · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
24022 |
24092 |
+0.1% (+29.5) |
−155 to +190 |
≤ 1201 |
PASSED |
MB/s |
25536 |
25462 |
−0.1% (−31.14) |
−202 to +162.3 |
≥ −1277 |
PASSED |
Russian Rebar prefilter single-literal count row · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
24088 |
24130 |
+5.5 |
−442 to +256 |
≤ 1204 |
PASSED |
MB/s |
25467 |
25422 |
−6.525 |
−270.1 to +455 |
≥ −1273 |
PASSED |
BenchmarkBar N5 raw-transition 5MiB Find rows: miss and late-hit · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
bar_miss_x_vs_best |
0.9695 |
0.9647 |
−0.5% (−0.00485) |
−0.0237 to +0.014 |
≤ 0.04848 |
PASSED |
bar_late_hit_x_vs_best |
0.966 |
0.9614 |
−0.2% (−0.00155) |
−0.0069 to +0.0035 |
≤ 0.0483 |
PASSED |
Checks: 2 of 2 passed. Verification: no defect found.