Batch LSM view setup for cache-miss distances

perfloop/weaviate · UNDER BATCHING

https://perfloop.ai/t/oss/case_152zs6wagt

Verdict

VERIFIED · settled 2026-08-17

What happened: The paired measurements met the required improvement.

Hypothesis

A source-trace check followed the provided gRPC Search path into the HNSW expansion loop and this method. It showed that an empty PrefetchGet result calls distanceToFloatNode once per ID. Standard shard wiring sends that lookup to Shard.vectorByIndexID, which acquires and releases a new objects-bucket consistent view around each secondary-key lookup. GetBySecondaryWithBufferAndView documents that a shared view is reusable across batch lookups and avoids fresh flush-lock acquisition, segment snapshots, and per-segment refcount churn.

The HNSW loop calls DistancesToNodes for every expanded candidate's unvisited neighbor batch. Its normal layer-zero capacity is 2*MaxConnections, or 64 at the default MaxConnections of 32, and filtered ACORN can supply up to 512 IDs at that default. When vectors are absent from cache, including while prefill runs asynchronously or a cache is capped, every missing ID repeats the view setup. The change would remove repeated view setup, not the required per-ID secondary-key reads. Whether that overhead materially affects search latency is a medium-confidence hypothesis because the source does not show cache-miss rate, segment layout, or storage wait time.

Benchmark near-vector gRPC searches immediately after startup with WaitForCachePrefill disabled and with a deliberately capped vector cache. Record PrefetchGet misses, GetConsistentView calls, secondary-key lookups, flush-lock wait, CPU, and end-to-end p50 and p99 latency. Compare the current fallback with one lazy shared view per nonempty miss batch. Confirm that view acquisitions fall from cache-miss count to miss-batch count while lookup count and per-ID error behavior stay unchanged, then require a latency or CPU improvement without a warm-cache regression.

Change to test: Acquire a consistent bucket view lazily on the first cache miss and reuse it for all miss reads in that DistancesToNodes batch through the view-aware vector thunks. Reuse one scratch vector and release the view at batch completion while preserving per-ID errors and multivector semantics.

Where it lives

perfloop/weaviate · adapters/handlers/grpc/v1/service.go

Evidence

64-vector cold HNSW cache-miss distance batch · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 152513 107508 −29.1% (−44410) −49265 to −36626 < −7626 PASSED
cache-misses/op 64 64 0 0 to 0 ≤ 0 PASSED
lookups/op 64 64 0 0 to 0 ≤ 0 PASSED
views/op 64 1 −98.4% (−63) −63 to −63 < −3.2 PASSED
cache-over-capacity/op 0 0 0 0 to 0 ≤ 0 PASSED

512-vector cold HNSW cache-miss distance batch with a 64-vector cache cap · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 1125623 792016 −30.1% (−339281) −359843 to −320314 < −56281 PASSED
cache-misses/op 512 512 0 0 to 0 ≤ 0 PASSED
lookups/op 512 512 0 0 to 0 ≤ 0 PASSED
views/op 512 1 −99.8% (−511) −511 to −511 < −25.6 PASSED
cache-over-capacity/op 448 448 0 0 to 0 ≤ 0 PASSED

post-startup high-EF non-exact gRPC near-vector search with asynchronous prefill and a 64-object vector cache · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 6275572 5629517 −10.6% (−662971) −743709 to −612119 ≤ 250000 PASSED
p50-ns/op 6089972 5474626 −10.5% (−638468) −703097 to −564562 < −304499 PASSED
p99-ns/op 10559519 9183208 −12.8% (−1350787) −1575214 to −1341128 ≤ 500000 PASSED

warm-cache high-EF non-exact gRPC near-vector search · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 1242966 1238668 −0.2% (−2902) −14695 to +3783 ≤ 50000 PASSED

Checks: 5 of 5 passed. Verification: no defect found.

Timeline