Parallelize long scanner inputs at full-context pretoken cuts

perfloop/fastokens · DATA PARALLEL GAP

https://perfloop.ai/t/oss/case_cab62jzp3j

Verdict

CLOSED · opened 2026-09-07

What happened: Closed; the assigned hypothesis was retired.

Hypothesis

The current driver has a structural serial gap when no newline supplies a conservative cut, even though the scanner can emit ordered pretokens. The change is deliberately more constrained than re-scanning substrings: its viability rests on proving scanner context and EOF behavior, then on an end-to-end crossover rather than presumed parallel speedup.

First property-test the exact cut selector: compare whole-buffer scanner ranges and IDs with concatenated full-context authority-range results for O200k and Kimi across adversarial whitespace, CRLF, contractions, punctuation, Han, combining marks, emoji, prefix-cache modes, and 1..available worker counts. Any range or ID mismatch rejects it. If parity holds, fresh-process release benchmarks at 63/64/65 KiB, 128 KiB, and 1 MiB no-newline inputs must beat the current serial route without regressing current newline-parallel inputs, allocations/RSS, or p95; otherwise reject.

Architecture-independent scheduling change only. Apply solely after current normalization and added-token segmentation on the existing recognized plain-text scanner branch, retain the one-pass newline path, and do not use naïve isolated substring scans or fixed lookaround. Preserve full right-context/EOF semantics, scanner boundaries, token ordering, exact IDs, prefix-cache reuse-boundary validity, special-token trust, decoding, and public APIs.

Change to test: For eligible long scanner inputs lacking enough newline cuts, add a bounded discovery scan that retains selected emitted-pretoken cut points and use a full-context authority-range scanner/BPE API to process those ranges in parallel before concatenating IDs in order.

Where it lives

perfloop/fastokens · src/lib.rs

Evidence

Timeline