Direct initial-token lookup for encoded BPE characters

perfloop/basetenkenizer · POOR DATA LOCALITY

https://perfloop.ai/t/oss/case_qe7w2x7hx7

Verdict

OPEN · opened 2026-09-07

Hypothesis

The standard merge setup currently creates temporary UTF-8 text and probes a string hash map once per initial character. A small literal-ASCII table can keep common encoded characters in dense storage while the unchanged map fallback preserves every non-table character and current error behavior; it does not alter rank or tie handling in BPE merging.

On cache-miss standard-BPE encodes across short and long English/code and multilingual inputs, instrument literal-table hits and fallback lookups, compare exact IDs and errors, and profile CPU. Reject if any output or error behavior differs, or if table hits or map-attributed cost are immaterial.

Only the standard encoded-BPE path after its existing whole-token and cache checks. Do not reinterpret UTF-8 bytes through `byte_to_initial_token`, alter ByteLevel semantics, merge ordering, cache policy, or public API.

Change to test: Precompute a `[u32; 128]` literal-ASCII initial-token table in `Bpe::new` and use it in `merge_all_encoded_into`, retaining the current `token_to_id` lookup and error path on every table miss.

Where it lives

perfloop/basetenkenizer · src/lib.rs

Evidence

No usable result yet.

Timeline