Batch-sized scratch allocation exceeds prefetch window
perfloop/weaviate · ALLOCATION HOT LOOP
https://perfloop.ai/t/oss/case_ktn5vy5hsa
Verdict
VERIFIED · settled 2026-08-17
What happened: The paired measurements met the required improvement.
Hypothesis
I ran a narrow overlay benchmark that reset `b.vecs` before each cached 64-ID batch and compared it with a warmed scratch buffer. The cold case reported 1 allocation and 1,815 B/op at 2,878 ns/op; the warmed case reported 0 allocations and 23 B/op at 2,541 ns/op. I also ran `TestFloatBatchDistancerMatchesDistanceToFloatNode` with its test root redirected only to a writable temporary directory; it passed its cached, evicted, missing, short, and regrown batch cases.
`DistancesToNodes` is called for each expanded candidate's `unvisited` neighbor batch in the search loop. A sweeping batch can reach `maximumConnectionsLayerZero`, and an ACORN batch can reach eight times that width. Although the prefetch lead is four, the method grows and retains `b.vecs` to every larger batch width; a newly obtained or GC-cleared pooled distancer therefore creates one escaping, heap-backed pointer buffer proportional to `len(ids)`. At most the current vector and four lookahead vectors need to remain live. The benchmark establishes that allocation delta for a cache-hit calibration, while whether its CPU or GC cost is material in production searches remains a hypothesis.
A case session should compare the current code with the ring under cold pool starts and warmed pools, sweeping production vector dimensions, connection limits, ACORN versus sweeping, and concurrent searches. It must preserve every distance and per-node error, then show the removed allocation and allocated-byte delta persists at realistic batch widths and reduces object-search CPU, GC work, or request p50/p99 latency.
Change to test: Replace the proportional scratch slice with a fixed five-slot ring: the current vector plus four prefetched vectors. Clear and reuse a consumed slot only after its distance is computed, preserving the existing four-item prefetch lead.
Where it lives
perfloop/weaviate · adapters/handlers/grpc/v1/service.go
Evidence
16 concurrent cold cache-hit generated gRPC Searches through Service, DB, Index, Shard, and HNSW: explicit Limit 30 (HNSW k=30 and automatic EF), FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64; one completed GC retains returned pool entries for the live scannable-heap sample · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-float-batch-gc-scan-B/search |
120 |
0 |
−100% (−120) |
−150 to −120 |
< −6 |
PASSED |
cold cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 384 dimensions, M=32 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-object-search-B/op |
60088 |
58816 |
−2% (−1192) |
−1512 to −968 |
≤ 3000 |
PASSED |
db-float-batch-B/op |
1424 |
0 |
−100% (−1424) |
−1424 to −1424 |
< −71.2 |
PASSED |
db-float-batch-allocs/op |
3 |
0 |
−100% (−3) |
−3 to −3 |
< −0.15 |
PASSED |
16 concurrent cold cache-hit generated gRPC Searches through the real Service, DB, Index, Shard, and HNSW: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 1024 dimensions, M=64 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-object-search-B/search |
27721 |
27180 |
−1.9% (−540.5) |
−552 to −419 |
≤ 1386 |
PASSED |
db-float-batch-B/search |
568 |
0 |
−100% (−568) |
−568 to −568 |
< −28.4 |
PASSED |
db-float-batch-allocs/search |
0.5 |
0 |
−100% (−0.5) |
−0.5 to −0.5 |
< −0.025 |
PASSED |
16 concurrent cold cache-hit generated gRPC Searches through the real Service, DB, Index, Shard, and HNSW: request Limit omitted so the fixture default makes HNSW k=10 with automatic EF, FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-object-search-B/search |
38409 |
37758 |
−0.7% (−258) |
−1085 to −87 |
≤ 1920 |
PASSED |
db-float-batch-B/search |
120 |
0 |
−100% (−120) |
−150 to −120 |
< −6 |
PASSED |
db-float-batch-allocs/search |
0.25 |
0 |
−100% (−0.25) |
−0.3125 to −0.25 |
< −0.0125 |
PASSED |
cold cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 30 (HNSW k=30 and automatic EF), FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 768 dimensions, M=32 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-object-search-B/op |
66376 |
66296 |
−0.1% (−84) |
−288 to +392 |
≤ 3319 |
PASSED |
db-float-batch-B/op |
224 |
0 |
−100% (−224) |
−224 to −224 |
< −11.2 |
PASSED |
db-float-batch-allocs/op |
2 |
0 |
−100% (−2) |
−2 to −2 |
< −0.1 |
PASSED |
warmed cache-hit default-flat-path generated gRPC Search through the real Service, DB, Index, and Shard: request Limit omitted so the fixture default is 10, supported FlatSearchCutoff 40,000 and automatic EF, 25% filter, 1,024 objects, 1024 dimensions, M=64; the whole-request allocation guard performs 64 additional response-validated warmed searches without explicit collection · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-request-B/search |
136099 |
135695 |
−0.3% (−392.5) |
−730 to −75 |
≤ 2000 |
PASSED |
warmed cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 1024 dimensions, M=64; the whole-request allocation guard performs 64 additional response-validated warmed searches without explicit collection · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-request-B/search |
299170 |
299456 |
+0.2% (+597.5) |
−583 to +1551 |
≤ 2000 |
PASSED |
64 warmed cache-hit generated gRPC Searches per timed operation through the real Service, DB, Index, Shard, and HNSW as four 16-request bursts: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, 1024 dimensions, M=64; the allocation guard adds one response-validated 64-request burst without explicit collection · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-request-B/search |
314410 |
315427 |
+0.1% (+458.5) |
−740 to +1575 |
≤ 2000 |
PASSED |
64 warmed cache-hit generated gRPC Searches per timed operation through the real Service, DB, Index, Shard, and HNSW as four 16-request bursts: request Limit omitted so the fixture default makes HNSW k=10 with automatic EF, FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64; the allocation guard adds one response-validated 64-request burst without explicit collection · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-request-B/search |
166903 |
166812 |
−0.3% (−466) |
−2049 to +644 |
≤ 2000 |
PASSED |
warmed cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW: explicit Limit 30 (HNSW k=30 and automatic EF), FlatSearchCutoff 0, ACORN with a 25% filter and default 0.4 ratio, 1,024 objects, 1536 dimensions, M=64; after a response-validated target warming search, one completed GC samples the returned pool entry's live scannable float-batch storage · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-float-batch-gc-scan-B/op |
960 |
0 |
−100% (−960) |
−960 to −960 |
< −48 |
PASSED |
cold cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW at the supported minimum M=4: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, and 384 dimensions · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-object-search-B/op |
53344 |
53232 |
−0.1% (−64) |
−272 to +48 |
≤ 2667 |
PASSED |
db-float-batch-B/op |
288 |
0 |
−100% (−288) |
−288 to −288 |
< −14.4 |
PASSED |
db-float-batch-allocs/op |
2 |
0 |
−100% (−2) |
−2 to −2 |
< −0.1 |
PASSED |
warmed cache-hit generated gRPC Search through the real Service, DB, Index, Shard, and HNSW at the supported minimum M=4: explicit Limit 64 (HNSW k=64 and automatic EF), FlatSearchCutoff 0, sweeping with no filter, 1,024 objects, and 384 dimensions; the whole-request allocation guard performs 64 additional response-validated warmed searches without explicit collection · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
db-request-B/search |
292848 |
293819 |
+0.2% (+480.5) |
+36 to +1451 |
≤ 2000 |
PASSED |
Checks: 5 of 5 passed. Verification: no defect found.
Timeline
2026-08-12· Case opened