Batch-sized scratch allocation exceeds prefetch window

perfloop/weaviate · ALLOCATION HOT LOOP

https://perfloop.ai/t/oss/case_ktn5vy5hsa

Verdict

VERIFIED · settled 2026-08-17

What happened: The paired measurements met the required improvement.

Hypothesis

I ran a narrow overlay benchmark that reset `b.vecs` before each cached 64-ID batch and compared it with a warmed scratch buffer. The cold case reported 1 allocation and 1,815 B/op at 2,878 ns/op; the warmed case reported 0 allocations and 23 B/op at 2,541 ns/op. I also ran `TestFloatBatchDistancerMatchesDistanceToFloatNode` with its test root redirected only to a writable temporary directory; it passed its cached, evicted, missing, short, and regrown batch cases.

`DistancesToNodes` is called for each expanded candidate's `unvisited` neighbor batch in the search loop. A sweeping batch can reach `maximumConnectionsLayerZero`, and an ACORN batch can reach eight times that width. Although the prefetch lead is four, the method grows and retains `b.vecs` to every larger batch width; a newly obtained or GC-cleared pooled distancer therefore creates one escaping, heap-backed pointer buffer proportional to `len(ids)`. At most the current vector and four lookahead vectors need to remain live. The benchmark establishes that allocation delta for a cache-hit calibration, while whether its CPU or GC cost is material in production searches remains a hypothesis.

A case session should compare the current code with the ring under cold pool starts and warmed pools, sweeping production vector dimensions, connection limits, ACORN versus sweeping, and concurrent searches. It must preserve every distance and per-node error, then show the removed allocation and allocated-byte delta persists at realistic batch widths and reduces object-search CPU, GC work, or request p50/p99 latency.

Change to test: Replace the proportional scratch slice with a fixed five-slot ring: the current vector plus four prefetched vectors. Clear and reuse a consumed slot only after its distance is computed, preserving the existing four-item prefetch lead.

Where it lives

perfloop/weaviate · adapters/handlers/grpc/v1/service.go

Evidence

16 concurrent cold cache-hit generated gRPC Searches through Service, DB, Index, Shard, and HNSW: explicit Limit 30 (HNSW k=30 and automatic EF), FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64; one completed GC retains returned pool entries for the live scannable-heap sample · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-float-batch-gc-scan-B/search 120 0 −100% (−120) −150 to −120 < −6 PASSED

cold cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 384 dimensions, M=32 · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-object-search-B/op 60088 58816 −2% (−1192) −1512 to −968 ≤ 3000 PASSED
db-float-batch-B/op 1424 0 −100% (−1424) −1424 to −1424 < −71.2 PASSED
db-float-batch-allocs/op 3 0 −100% (−3) −3 to −3 < −0.15 PASSED

16 concurrent cold cache-hit generated gRPC Searches through the real Service, DB, Index, Shard, and HNSW: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 1024 dimensions, M=64 · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-object-search-B/search 27721 27180 −1.9% (−540.5) −552 to −419 ≤ 1386 PASSED
db-float-batch-B/search 568 0 −100% (−568) −568 to −568 < −28.4 PASSED
db-float-batch-allocs/search 0.5 0 −100% (−0.5) −0.5 to −0.5 < −0.025 PASSED

16 concurrent cold cache-hit generated gRPC Searches through the real Service, DB, Index, Shard, and HNSW: request Limit omitted so the fixture default makes HNSW k=10 with automatic EF, FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64 · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-object-search-B/search 38409 37758 −0.7% (−258) −1085 to −87 ≤ 1920 PASSED
db-float-batch-B/search 120 0 −100% (−120) −150 to −120 < −6 PASSED
db-float-batch-allocs/search 0.25 0 −100% (−0.25) −0.3125 to −0.25 < −0.0125 PASSED

cold cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 30 (HNSW k=30 and automatic EF), FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 768 dimensions, M=32 · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-object-search-B/op 66376 66296 −0.1% (−84) −288 to +392 ≤ 3319 PASSED
db-float-batch-B/op 224 0 −100% (−224) −224 to −224 < −11.2 PASSED
db-float-batch-allocs/op 2 0 −100% (−2) −2 to −2 < −0.1 PASSED

warmed cache-hit default-flat-path generated gRPC Search through the real Service, DB, Index, and Shard: request Limit omitted so the fixture default is 10, supported FlatSearchCutoff 40,000 and automatic EF, 25% filter, 1,024 objects, 1024 dimensions, M=64; the whole-request allocation guard performs 64 additional response-validated warmed searches without explicit collection · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-request-B/search 136099 135695 −0.3% (−392.5) −730 to −75 ≤ 2000 PASSED

warmed cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 1024 dimensions, M=64; the whole-request allocation guard performs 64 additional response-validated warmed searches without explicit collection · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-request-B/search 299170 299456 +0.2% (+597.5) −583 to +1551 ≤ 2000 PASSED

64 warmed cache-hit generated gRPC Searches per timed operation through the real Service, DB, Index, Shard, and HNSW as four 16-request bursts: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 1024 dimensions, M=64; the allocation guard adds one response-validated 64-request burst without explicit collection · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-request-B/search 314410 315427 +0.1% (+458.5) −740 to +1575 ≤ 2000 PASSED

64 warmed cache-hit generated gRPC Searches per timed operation through the real Service, DB, Index, Shard, and HNSW as four 16-request bursts: request Limit omitted so the fixture default makes HNSW k=10 with automatic EF, FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64; the allocation guard adds one response-validated 64-request burst without explicit collection · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-request-B/search 166903 166812 −0.3% (−466) −2049 to +644 ≤ 2000 PASSED

warmed cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 30 (HNSW k=30 and automatic EF), FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64; after a response-validated target warming search, one completed GC samples the returned pool entry's live scannable float-batch storage · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-float-batch-gc-scan-B/op 960 0 −100% (−960) −960 to −960 < −48 PASSED

cold cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW at the supported minimum M=4: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, and 384 dimensions · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-object-search-B/op 53344 53232 −0.1% (−64) −272 to +48 ≤ 2667 PASSED
db-float-batch-B/op 288 0 −100% (−288) −288 to −288 < −14.4 PASSED
db-float-batch-allocs/op 2 0 −100% (−2) −2 to −2 < −0.1 PASSED

warmed cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW at the supported minimum M=4: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, and 384 dimensions; the whole-request allocation guard performs 64 additional response-validated warmed searches without explicit collection · 10 sample pairs

metric baseline candidate paired median change confidence range required result
db-request-B/search 292848 293819 +0.2% (+480.5) +36 to +1451 ≤ 2000 PASSED

Checks: 5 of 5 passed. Verification: no defect found.

Timeline