Share fused pretoken results across worker caches
perfloop/fastokens · CACHE THRASH
https://perfloop.ai/t/oss/case_pmkbqzk9t1
Verdict
PROPOSED · opened 2026-09-07
Hypothesis
Basetenkenizer supplies a publicly visible L1→shared-L2 mechanism absent from the shown target scanner route. It is novel comparison evidence despite declared common ancestry, but shared-cache benefit is intentionally conditional because the target's local cache and fixed worker pool already provide reuse and shared locking can worsen multithreaded work.
Feature-gate the L2 and count eligible spans, L1 misses, L2 probes/hits/inserts/clears, `merge_all_raw_into` calls, per-shard lock wait, RSS, throughput, and p50/p99. Use a rotated-worker repeated-pretoken positive control plus unique code-like and multilingual negative controls in fresh processes; reject if cross-worker L2 hits or avoided merges are negligible, IDs differ, or target-concurrency throughput/p99/RSS regresses.
Architecture-independent cache change only. Scope entries to one immutable BPE instance and exact post-normalization raw spans using the same <=15-byte eligibility and full-key equality as the local cache; release any shard lock before BPE work. Keep memory bounded and keep the raw fused namespace separate from the encoded `shared_cache`; preserve output order, token IDs, added/special-token boundaries, normalization, decoding, and APIs.
Change to test: Add a bounded sharded fused L2 cache keyed by the BPE instance and exact raw pretoken; probe it after a `TL_FUSED_CACHE` miss, backfill the local cache on a hit, and insert only successful computed results after merging.
Where it lives
perfloop/fastokens · src/lib.rs
Evidence
No usable result yet.
Timeline
2026-09-07· Case opened