Redundant UTF-8 Decoding and Case-Folding on ASCII Content
perfloop/zoekt · INEFFICIENT ALGORITHM
https://perfloop.ai/t/oss/case_19fpgzx67m
Verdict
VERIFIED · settled 2026-07-07 · merged as sourcegraph/zoekt#1083
What happened: Introducing an ASCII fast path in caseFoldingEqualsRunes delivers a massive performance boost for the hot-path search match validation. By comparing single-byte ASCII characters directly using simple bitwise/arithmetic operations, we bypass the computationally expensive UTF-8 rune decoding and Unicode lookup tables entirely when processing ASCII segments.
Since source code documents are overwhelmingly ASCII, this optimization yields a ~57.4% reduction in match evaluation time for ASCII text (from 247.4 ns/op down to 105.35 ns/op). For Unicode text containing mixed ASCII/non-ASCII characters, we still see a ~30.8% reduction in execution time (from 368.25 ns/op down to 254.8 ns/op) because the fast path is taken for all ASCII segments of the string. Both results are highly statistically significant (p < 1.1e-5), with zero allocations or byte overhead regressions.
Hypothesis
caseFoldingEqualsRunes compares a lowercase needle against a mixed-case content slice by performing character-by-character validation. It currently decodes every character as a UTF-8 rune and translates it using unicode.ToLower, which is computationally expensive. Since source code repositories consist overwhelmingly of ASCII text, this decoding and full Unicode mapping is largely redundant. Direct byte comparison with simple case-folding for ASCII avoids these expensive operations. This anchor reaches the hot path across an interface dispatch with 15 runtime implementations; the index assumes, but does not prove, that runtime binding reaches this implementation. The anchor hop was synthesized from code-index call references, so the model has not proven whether the anchored call repeats per element.
Change to test: Introduce an ASCII fast path at the beginning of the character loop in caseFoldingEqualsRunes to compare lower[0] and mixed[0] directly as bytes when both are less than utf8.RuneSelf, performing ASCII case conversion using basic arithmetic.
Where it lives
perfloop/zoekt · index/eval.go