Narrow ACORN candidate-node lock scope
perfloop/weaviate · LOCK CONTENTION
https://perfloop.ai/t/oss/case_ncqv07d22b
Verdict
VERIFIED · settled 2026-08-17
What happened: The paired measurements met the required improvement.
Hypothesis
I ran a source-scope lock audit of `searchLayerByVectorWithDistancerWithStrategy` and `vertex`. The audit showed that `vertex` embeds `sync.Mutex`, and the search takes that mutex once for every candidate popped from the layer queue. In the ACORN branch it copies the candidate adjacency layer, then retains the same mutex through the two-hop expansion loop, including allow-list checks, key lookups, child-node lookup locks, and child adjacency iteration. The candidate node is not accessed again after its layer has been copied until the final unlock.
The relevant cadence is once per expanded candidate, not once per request. Concurrent gRPC Search handler goroutines can execute this per-request path against the same HNSW index, so common entry or hub candidates could make later searches wait on a mutex while the current search expands unrelated nodes. The source audit proves the oversized critical section but does not measure lock waits, so contention and its cost remain a hypothesis.
A case session should run concurrent vector searches with an allow list that selects ACORN, collect a Go mutex profile, and compare the narrowed-lock implementation with the baseline. The proof signal is reduced mutex wait attributed to this search path, accompanied by better p95 search latency or throughput at unchanged recall; the removed interval is the work from the copied candidate layer through completion of its two-hop expansion.
Change to test: In the ACORN branch, release the candidate node mutex immediately after checking its level and copying its adjacency layer into the scratch slice. Run the two-hop expansion from that snapshot while preserving panic-safe cleanup around the short critical section.
Where it lives
perfloop/weaviate · adapters/handlers/grpc/v1/service.go
Evidence
Steady-state GOMAXPROCS=8 v1 Service.Search with 16 concurrent benchmark workers through the DB, index, shard, inverted-filter allow list, and HNSW path: a deterministic 512-node, 64-D cosine graph is queried by near vector plus equality filter for 8 IDs, selecting ACORN with FlatSearchCutoff=8 and the default 0.4 ACORN ratio; every matching UUID must be returned exactly once. · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
77347 |
68424 |
−11.8% (−9103) |
−10002 to −7419 |
< −3867 |
PASSED |
Steady-state GOMAXPROCS=8 v1 Service.Search with eight concurrent benchmark workers through the production default flat-search cutoff: the deterministic 512-node, 64-D cosine fixture uses a near vector plus equality filter for 64 IDs and requires every matching UUID exactly once; the query profile verifies the default flat-search branch. · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
163215 |
164531 |
+0.4% (+678) |
−3529 to +5593 |
≤ 8161 |
PASSED |
Steady-state GOMAXPROCS=8 v1 Service.Search with eight concurrent benchmark workers through DB, index, shard, inverted-filter allow list, and the supported sweeping HNSW strategy: the deterministic 512-node, 64-D cosine fixture queries 64 equality-filtered IDs, requires every matching UUID exactly once, and uses FlatSearchCutoff=8. · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
282154 |
276205 |
−1.2% (−3478) |
−11936 to +7230 |
≤ 14108 |
PASSED |
Checks: 7 of 7 passed. Verification: no defect found.
Timeline
2026-08-12· Case opened