Parallelize long scanner inputs at full-context pretoken cuts
perfloop/fastokens · DATA PARALLEL GAP
https://perfloop.ai/t/oss/case_cab62jzp3j
Verdict
CLOSED · opened 2026-09-07
What happened: Closed; the assigned hypothesis was retired.
Hypothesis
The current driver has a structural serial gap when no newline supplies a conservative cut, even though the scanner can emit ordered pretokens. The change is deliberately more constrained than re-scanning substrings: its viability rests on proving scanner context and EOF behavior, then on an end-to-end crossover rather than presumed parallel speedup.
First property-test the exact cut selector: compare whole-buffer scanner ranges and IDs with concatenated full-context authority-range results for O200k and Kimi across adversarial whitespace, CRLF, contractions, punctuation, Han, combining marks, emoji, prefix-cache modes, and 1..available worker counts. Any range or ID mismatch rejects it. If parity holds, fresh-process release benchmarks at 63/64/65 KiB, 128 KiB, and 1 MiB no-newline inputs must beat the current serial route without regressing current newline-parallel inputs, allocations/RSS, or p95; otherwise reject.
Architecture-independent scheduling change only. Apply solely after current normalization and added-token segmentation on the existing recognized plain-text scanner branch, retain the one-pass newline path, and do not use naïve isolated substring scans or fixed lookaround. Preserve full right-context/EOF semantics, scanner boundaries, token ordering, exact IDs, prefix-cache reuse-boundary validity, special-token trust, decoding, and public APIs.
Change to test: For eligible long scanner inputs lacking enough newline cuts, add a bounded discovery scan that retains selected emitted-pretoken cut points and use a full-context authority-range scanner/BPE API to process those ranges in parallel before concatenating IDs in order.
Where it lives
perfloop/fastokens · src/lib.rs
Evidence
Timeline
2026-09-07· Case opened2026-09-21· Case closed