Share fused pretoken results across worker caches

perfloop/fastokens · CACHE THRASH

https://perfloop.ai/t/oss/case_pmkbqzk9t1

Verdict

PROPOSED · opened 2026-09-07

Hypothesis

Basetenkenizer supplies a publicly visible L1→shared-L2 mechanism absent from the shown target scanner route. It is novel comparison evidence despite declared common ancestry, but shared-cache benefit is intentionally conditional because the target's local cache and fixed worker pool already provide reuse and shared locking can worsen multithreaded work.

Feature-gate the L2 and count eligible spans, L1 misses, L2 probes/hits/inserts/clears, `merge_all_raw_into` calls, per-shard lock wait, RSS, throughput, and p50/p99. Use a rotated-worker repeated-pretoken positive control plus unique code-like and multilingual negative controls in fresh processes; reject if cross-worker L2 hits or avoided merges are negligible, IDs differ, or target-concurrency throughput/p99/RSS regresses.

Architecture-independent cache change only. Scope entries to one immutable BPE instance and exact post-normalization raw spans using the same <=15-byte eligibility and full-key equality as the local cache; release any shard lock before BPE work. Keep memory bounded and keep the raw fused namespace separate from the encoded `shared_cache`; preserve output order, token IDs, added/special-token boundaries, normalization, decoding, and APIs.

Change to test: Add a bounded sharded fused L2 cache keyed by the BPE instance and exact raw pretoken; probe it after a `TL_FUSED_CACHE` miss, backfill the local cache on a hit, and insert only successful computed results after merging.

Where it lives

perfloop/fastokens · src/lib.rs

Evidence

No usable result yet.

Timeline