Replace the decoded transition loop with one raw-byte plan

perfloop/casei · DEPENDENT LOAD CHAIN

https://perfloop.ai/t/oss/case_rmg4fdm3me

Verdict

VERIFIED · settled 2026-08-25

What happened: The paired measurements met the required improvement.

Hypothesis

I checked `casei.searchPlan.findFiltered` in `plan.go:2528-2585`. Each non-skipped candidate calls `p.haystackToken(haystack, at)` at line 2569 and passes the returned token directly to `p.advance(state, token)` at line 2570. The model already reaches this body from workload entry points through `casei.searchPlan.find`.

`haystackToken` selects an ASCII or opaque table entry, or decodes a rune and probes `p.runes`. `advance` then indexes the dense transition table with the prior state and that token, or follows node edges and failure links. This is a serial lookup-and-transition chain on each admitted candidate. It suggests a dependent-load-chain cost if those lookups are on the dominant realized iterations and their latency is not hidden by other work.

Confirm this with a reproducible candidate-density sweep that counts non-skipped loop iterations and compares the decoded path with the raw-byte plan. Inspect generated code and hardware counters for the dependent lookup chain, verify that any compiled transition table remains cache-resident and amortizes its build cost, and differentially test Unicode folding, invalid bytes, offsets, and tie behavior.

Change to test: Compile complete simple-fold literal transitions into a compact raw-byte plan that lets filtered search advance without the decoded token boundary, while retaining decoded handling for invalid or unsupported input and feature-gated vector executors. Preserve byte offsets, leftmost and lowest-ID ties, and the portable fallback.

Where it lives

perfloop/casei · matcher.go

Evidence

36-row pinned native BenchmarkBar with N=5 raw root-transition miss and late-hit outcomes · 10 sample pairs

metric baseline candidate paired median change confidence range required result
benchmarkbar_raw_transition_miss_x_vs_best 6.532 0.9977 −84.3% (−5.508) −6.19 to −5.045 < −1 PASSED
benchmarkbar_raw_transition_late_hit_x_vs_best 7.145 0.9251 −86.6% (−6.184) −6.619 to −5.627 < −1 PASSED

Checks: 11 of 11 passed. Verification: no defect found.

Timeline