Bulk tenant resolution removes per-object leader queries

perfloop/weaviate · N+1

https://perfloop.ai/t/oss/case_x82s8m33e0

Verdict

VERIFIED · settled 2026-08-13

What happened: The paired measurements met the required improvement.

Hypothesis

Source inspection found that putObjectBatch iterates over every object and synchronously calls shardResolver.ResolveShard before it groups by shard. In the multi-tenant resolver, each ResolveShard calls ValidateTenants for one tenant; ValidateTenants calls schemaReader.TenantsShards. The production schema Manager sorts/compacts its input and directly calls schemaManager.QueryTenantsShards once per invocation, and the Raft implementation builds a query documented as directed to the leader. Therefore an all-valid N-object, one-class multi-tenant batch structurally issues N serial leader-status queries, including repeated queries for the same tenant. The already-present ResolveShards path instead extracts unique tenants, performs one bulk ValidateTenants call, and returns targets for all objects on success. This is not already coalesced at the Manager layer: each current one-tenant invocation directly creates its own QueryTenantsShards call.

The workload cadence is one processRequest dequeued by worker.Loop at a time; its adaptive BatchStream sizing starts at 200 objects and clamps to 100–1000, so a single-class object-heavy work item can expose N in that range. I ran `go test ./adapters/repos/db/sharding -run 'Test_ShardResolution_MultiTenant|Test_ResolveShard_MultiTenant' -count=1`; it passed, confirming the existing resolver's multi-tenant batch target behavior, but it did not measure request counts or latency. The leader-query cost share remains a hypothesis.

Proof target: trace one valid multi-tenant BatchStream batch at N=100, 200, and 1000 with U unique tenants, counting Manager/Raft QueryTenantsShards spans and their serial waterfall plus end-to-end BatchStream p50/p95. The baseline should show N queries per one-class batch; the bulk path should show one query with U tenant names while preserving per-object results for mixed-validity batches.

Change to test: For the successful multi-tenant batch path, use the resolver's bulk target-resolution path before grouping objects so one-class batches replace N per-object tenant-status queries with one query containing the U unique tenants. Preserve putObjectBatch's partial per-object error behavior: either have the bulk resolver return per-object outcomes or use an error-path fallback, rather than turning one invalid tenant into a whole-batch failure.

Where it lives

perfloop/weaviate · adapters/handlers/grpc/v1/batch/worker.go

Evidence

BatchStream multi-tenant write: 100 objects across 10 hot tenants · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 5378513 4451728 −22.2% (−1194285) −1585572 to −527361 < 0 PASSED
p50-ns/op 3862405 3145162 −20.3% (−785922) −963522 to −648177 ≤ 0 PASSED
p95-ns/op 13235476 11634104 −11.4% (−1514026) −4337341 to −290602 ≤ 1000000 PASSED
leader-tenant-shard-queries/op 100 1 −99% (−99) −99 to −99 < 0 PASSED
go-user-cpu-ns/op 14313879 12976638 −15.4% (−2197210) −2994870 to −392136 ≤ 1000000 PASSED
B/op 4460880 3727117 −18.3% (−814438) −990201 to −595170 ≤ 0 PASSED
allocs/op 17383 13180 −24.2% (−4207) −4213 to −4195 ≤ 0 PASSED

BatchStream multi-tenant write: 200 objects across 10 hot tenants · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 5754439 4164855 −28.2% (−1622691) −7257911 to −1099743 < 0 PASSED
p50-ns/op 4313643 3152235 −29.4% (−1266253) −1421178 to −986838 ≤ 0 PASSED
p95-ns/op 12966177 9541469 −29% (−3754074) −56188541 to −1694094 ≤ 1000000 PASSED
leader-tenant-shard-queries/op 200 1 −99.5% (−199) −199 to −199 < 0 PASSED
go-user-cpu-ns/op 14482619 12326007 −18.3% (−2647225) −6856960 to −1515258 ≤ 1000000 PASSED
B/op 6597586 4959571 −24.6% (−1624200) −1810234 to −1415361 ≤ 0 PASSED
allocs/op 33211 24674 −25.7% (−8537) −8543 to −8532 ≤ 0 PASSED

BatchStream multi-tenant write: 1000 objects across 10 hot tenants · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 19712572 12614253 −35.7% (−7043137) −9132596 to −6479781 < 0 PASSED
p50-ns/op 17374783 11497831 −33% (−5737164) −7281063 to −5328701 ≤ 0 PASSED
p95-ns/op 35380966 23774689 −33.7% (−11923135) −16856328 to −8421869 ≤ 1000000 PASSED
leader-tenant-shard-queries/op 1000 1 −99.9% (−999) −999 to −999 < 0 PASSED
go-user-cpu-ns/op 49714229 38827437 −21.8% (−10857270) −15633335 to −9363337 ≤ 1000000 PASSED
B/op 25452124 17380862 −30% (−7626018) −8994288 to −7130510 ≤ 0 PASSED
allocs/op 160284 116722 −27.2% (−43531) −43665 to −43449 ≤ 0 PASSED

Checks: 8 of 8 passed. Verification: no defect found.

Timeline