Direct initial-token lookup for encoded BPE characters
perfloop/basetenkenizer · POOR DATA LOCALITY
https://perfloop.ai/t/oss/case_qe7w2x7hx7
Verdict
OPEN · opened 2026-09-07
Hypothesis
The standard merge setup currently creates temporary UTF-8 text and probes a string hash map once per initial character. A small literal-ASCII table can keep common encoded characters in dense storage while the unchanged map fallback preserves every non-table character and current error behavior; it does not alter rank or tie handling in BPE merging.
On cache-miss standard-BPE encodes across short and long English/code and multilingual inputs, instrument literal-table hits and fallback lookups, compare exact IDs and errors, and profile CPU. Reject if any output or error behavior differs, or if table hits or map-attributed cost are immaterial.
Only the standard encoded-BPE path after its existing whole-token and cache checks. Do not reinterpret UTF-8 bytes through `byte_to_initial_token`, alter ByteLevel semantics, merge ordering, cache policy, or public API.
Change to test: Precompute a `[u32; 128]` literal-ASCII initial-token table in `Bpe::new` and use it in `merge_all_encoded_into`, retaining the current `token_to_id` lookup and error path on every table miss.
Where it lives
perfloop/basetenkenizer · src/lib.rs
Evidence
No usable result yet.
Timeline
2026-09-07· Case opened