Sapphire Rapids: Fuse tagged trigram scanning with literal enumeration

perfloop/casei · INEFFICIENT ALGORITHM

https://perfloop.ai/t/oss/case_ttc9g9ze4t

Verdict

VERIFIED · settled 2026-09-29 · merged as tsenart/casei#23

What happened: The paired measurements met the required improvement.

Hypothesis

Speed up retained-Matcher full-text counts for the English five names and Tom/Sawyer/Huckleberry/Finn on Sapphire Rapids without changing which match wins or its byte width. Exploratory three-pass master-5524052 Vectorscan 5.4.12 medians are casei 191.76 versus 73.21 microseconds (needs 61.8% cut) and 1,830 versus 1,300 (needs 29.5%). A fused field win remains unproved.

Rebar supplies 899,232 SHA-verified subtitles bytes with 725 matches and 16,013,977 Leipzig bytes with 4,152 matches. Its timed count compiles Matcher once and enumerates the entire haystack. Current generic Each re-enters findFiltered after matches; that path uses both root filtering and the four-load/four-table bucket. tripleBucketSkipBytes dispatches the block routine when the bucket and VBMI are usable; otherwise it uses the conservative Shufti route. Captured unmodified full-operation five-name CPU samples are 49.1% flat tripleBucketSkip64 and 38.1% flat findFiltered. Even a free bucket would leave about 51% of current cost versus a 38.2% leader-time budget; the combined 87.2% is a loose headroom bound, not removable time. Leipzig bucket samples are 69.8% including untimed preflight.

Pinned Vectorscan's conditional three-mask VBMI Teddy classifies one loaded 64-byte block with register-shifted tags then confirms survivors. A separate untimed matching-configuration debugger entered that engine on these rows, not proof of timed dispatch. A casei-owned fold-stable interior trigram per pattern could screen once per block, replay the same compiled fold-token plan at candidates and defer emission until earliest start and lowest ID are known. S/K can widen under Unicode folding, so recover true source starts and keep decoded and portable paths. Existing two-byte UTF-8 tagged enumeration is an ordering precedent, not an eligible route for these rows. Actual static-key density, confirmation cost and savings require a Case.

All four/five patterns admit selective fold-stable three-unit interior keys. One loaded block and bounded same-plan replay, start recovery and ordered emission cost less than both current screening and repeated suffix traversal on model 143; no separate matcher is needed.

On cpu-sapphire-rapids, first time screening with the ACTUAL compiled tags, setup and tails on both full fixtures. If the screen alone exceeds approximately 73 microseconds or 1.3 milliseconds respectively, this classifier cannot alone beat the leader. If feasible, pair complete BenchmarkRebar Each counts against current master with survivor/replay/CPU receipts; targets must improve and Russian/name-alt5/affected BenchmarkBar rows not regress. Check Unicode, malformed bytes, variable widths, order, ties, early stopping and portable fallback. Screen-only speed does not prove full-operation savings; guarded partial savings may land but do not prove field leadership.

Unpaired guest CPU2 discovery is not Case proof or physical isolation. Retained files include three full model-143 Rebar CSVs, summary, raw 38-row Bar transcript, independent verification, pprof manifest and unmodified profiles; the Tom pprof includes roughly 12.4% untimed preflight samples. Separate untimed debugger entry into VBMI Teddy and captured AVX512VBMI database metadata do not certify the timed loop width. A local unretained static-score census suggested 3,232 five-name and 11,801 Tom three-byte key hits; actual selector and costs are unknown. The 16-file capture limit prevented retaining this census and the native debugger transcript. Oracle design advice is unconfirmed. Prior bucket Case case_109tpvhytb paired 27.2%/24.1% gains at an older revision; ring-only cand_bkfp5q3rbb paired 18%/8.6% but failed a Russian guard. Neither proves this new plan nor closes today's field gap. No model-106 baseline or final-commit certification exists.

Change to test: On Intel Sapphire Rapids model 143, replace eligible generic multi-literal Each's root filtering and repeated suffix searches with a plan-owned, one-load tagged interior-trigram transition that verifies survivors through the same fold-token plan and emits ordered, exact-width results. Retain decoded and feature-off fallbacks.

Where it lives

perfloop/casei · audit/rebar/runner/main.go

Evidence

Rebar curated/02 English five-name retained-Matcher Each count (full 899,232-byte fixture; 725 matches) · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 197158 180822 −8.1% (−16003) −19689 to −13546 < −9858 PASSED
MB/s 4561 4973 +8.9% (+407.1) +340 to +495.3 > 228.1 PASSED

Rebar Leipzig four-name retained-Matcher Each count (full 16,013,977-byte fixture; 4,152 matches) · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 1842441 1713666 −7.3% (−135274) −156845 to −109770 < −92122 PASSED
MB/s 8692 9345 +7.9% (+690.7) +540.6 to +809.4 > 434.6 PASSED

Rebar Russian five-name retained-Matcher Each guard (full pinned haystack) · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 515371 512978 −0.8% (−4070) −5312 to +329 ≤ 25769 PASSED
MB/s 3047 3062 +0.8% (+24.13) −1.96 to +31.78 ≥ −152.4 PASSED

Rebar imported name-alt5 retained-Matcher Each guard (full pinned Sherlock haystack) · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 121958 121812 −0.3% (−389.5) −1180 to +916 ≤ 6098 PASSED
MB/s 4878 4884 +0.3% (+15.65) −36.58 to +47.85 ≥ −243.9 PASSED

Checks: 9 of 9 passed. Verification: no defect found.

Timeline