A/B the complete amd64 backend as Go archsimd

perfloop/casei · INEFFICIENT ALGORITHM

https://perfloop.ai/t/oss/case_37sjyc8f94

Verdict

CLOSED · opened 2026-08-17

What happened: Closed; the settled verdict is recorded on the case.

Hypothesis

Source and model inspection place `casei.rootSkipASCII` on the workload-reachable Case-insensitive substring search path: `casei.searchPlan.find` calls it at `plan.go:2303` and `plan.go:2308`. In `root_amd64.go`, that dispatcher calls `rootSkip64` at line 103, `rootSkip32` at line 111, or `rootSkipScalar` at line 117; the accepted path also reaches the AVX2 `rootSkip32` branch.

The submitted A/B hypothesis is that the complete amd64 vector implementation set behind the unchanged dispatcher could avoid assembly call-boundary or compiler-opacity costs while preserving generated vector-loop quality. This has not been measured. It retains the dispatch condition that a qualifying AVX-512 BW host with at least 64 bytes selects `rootSkip64`, while `rootSkip32` is only the AVX2 32-byte branch; it does not claim `rootSkip32` runs on the 64-byte AVX-512 route.

Confirm or reject this as a complete-backend comparison. Differentially fuzz the assembly and `GOEXPERIMENT=simd-selected` `simd/archsimd` implementations, inspect generated kernels for calls, spills, masks, tails, and intended AVX2/AVX-512 BW/VBMI instructions, then run randomized whole-engine measurements on Ice Lake and Sapphire Rapids. The measurements must include per-kernel throughput, short-input latency, the exact Rebar corpus, and every `BenchmarkBar` row, so no individual kernel is selected independently.

Change to test: Add a `GOEXPERIMENT=simd-selected` pure-Go `simd/archsimd` implementation for every function currently implemented by `root_amd64.s`, with identical feature and length dispatch and no `root_amd64.s` body linked in that build, while retaining the existing assembly backend as the control. Differentially fuzz both complete backends, inspect generated kernels for calls, spills, masks, tails, and expected AVX2/AVX-512 BW/VBMI instructions, and co-measure the whole engines in randomized order on Ice Lake and Sapphire Rapids across per-kernel throughput, short-input latency, the exact Rebar corpus, and every `BenchmarkBar` row; accept or refute at the implementation-boundary level rather than selecting individual kernels.

Where it lives

perfloop/casei · matcher.go

Evidence

Timeline