Restore bounded span-only batches

perfloop/tempo · LOCK CONTENTION

https://perfloop.ai/t/oss/case_wqf5vvek76

Verdict

VERIFIED · settled 2026-08-12 · pull request opened as grafana/tempo#7711

What happened: The paired measurements met the required improvement.

Hypothesis

Static source check: I read pkg/traceql/engine_metrics.go and ran focused rg checks for MetricsEvaluator/iterateBlocks. CompileMetricsQueryRange sets spanOnlyFetch true by default, and metricsEvaluator.Do reaches DoSpansOnly when the query does not need a full trace. DoSpansOnly currently calls Results.Next and then locks e.mtx separately for every fetched span around time filtering, sampling, metricsPipeline.observe, watcher work, exemplar selection, counters, and length checking. Its own TODO says batching was removed to avoid holding the mutex during storage work and to allow SecondPass, and records an approximately 20% performance loss; that comment is source evidence, not a measurement I ran. I also read modules/livestore/instance_search.go: QueryRange creates one rawEval, passes it to each WAL-block callback, iterateBlocks launches bounded goroutines for WAL blocks, and queryRangeWALBlock calls eval.Do. Thus, with at least three eligible WAL blocks and QueryBlockConcurrency at least three, those request-hot block workers can overlap on the same evaluator mutex; actual lock wait and the share of query time remain unmeasured. The removed delta is one lock/unlock and potential handoff per fetched span changed to one per bounded batch, at the real cadence of every span yielded by each concurrent span-only WAL iterator, without moving iterator/storage work under the lock. Proof target: run a reproducible span-only QueryRange over enough WAL blocks and spans to saturate configured block concurrency, compare baseline and batched implementations using end-to-end latency, CPU, and Go mutex/block profiles, and require lower evaluator lock wait plus unchanged series, exemplars, watcher metrics, max-series behavior, and cancellation semantics.

Change to test: Reintroduce an iterator-lifetime-safe bounded span batch in DoSpansOnly so shared evaluator updates acquire e.mtx once per batch rather than once per span, while keeping Next/storage and SecondPass work outside the lock. Preserve per-span filtering, watcher and exemplar behavior, max-series early exit, cancellation, and result equivalence within each batch.

Where it lives

perfloop/tempo · modules/querier/http.go

Evidence

uncapped span-only QueryRange across 10 concurrent WAL blocks · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 10399387 7722757 −26.3% (−2730578) −3604276 to −2269278 < 0 PASSED
B/op 3350283 3395758 +1.3% (+44973) +42289 to +47798 ≤ 131072 PASSED
allocs/op 38503 38599 +0.2% (+96) +91 to +98 ≤ 256 PASSED

uncapped span-only QueryRange CPU and evaluator contention profiles across 10 WAL blocks · 10 sample pairs

metric baseline candidate paired median change confidence range required result
evaluator-mutex-delay-ns/op 25831850 4504590 −79.1% (−20426125) −22470940 to −16503800 < 0 PASSED
evaluator-block-delay-ns/op 28746800 5147445 −79.3% (−22807450) −24443920 to −19388700 < 0 PASSED
span-only-cpu-ns/op 23000000 14500000 −30.4% (−7000000) −10400000 to −5000000 ≤ 1000000 PASSED

Checks: 7 of 7 passed. Verification: no defect found.

Timeline