Replace periodic commit-log Stat checks with byte-threshold signaling

perfloop/weaviate · POLLING LOOP

https://perfloop.ai/t/oss/case_2k1tjm5ncx

Verdict

VERIFIED · settled 2026-08-16

What happened: The paired measurements met the required improvement.

Hypothesis

startSwitchLogs always enters switchCommitLogs(false). That function takes l.Mutex and calls currentFile.Stat before it can return for a sub-threshold file, while every WAL mutation uses the same mutex. The HNSW commit-log scheduler runs these checks on a 500 ms to 10 s cadence, so inactive and still-growing logs repeatedly take a metadata-poll path without rotating.

The active WALWriter is the sole append path and already runs while that mutex is held. A per-file byte counter can therefore identify the threshold crossing without a filesystem query. The no-pending cycle would stop at an atomic read; a real rotation still pays its required Flush, Sync, open, pointer handoff, and close. Seeding once after an append-open preserves the existing timestamp-collision case, while retaining the pending bit on an error preserves retry behavior.

A many-shard, below-threshold import and idle profile should record fstat syscall counts, logger-mutex blocked time, and end-to-end write latency; this hypothesis is falsified if those no-op polls are not material on the traced workload. Forced-backup and injected Flush, Sync, and Open failure tests must also show that a sealed file is never skipped and a failed rotation remains eligible for retry.

Change to test: Place a counting io.Writer between compact.WALWriter and each active bufio.Writer. Maintain accepted WAL bytes under the existing hnswCommitLogger mutex and set an atomic rotationPending bit only when the active file crosses maxSizeIndividual. Let startSwitchLogs return before taking the mutex when that bit is clear; recheck it under the mutex when set, use the tracked size and filepath.Base(currentFileName) instead of the per-cycle Stat, and seed the counter with one post-open Stat for an O_APPEND collision. Forced backup switches bypass the gate, and Flush/Sync/Open failures leave the bit set until a successful handoff resets it.

Where it lives

perfloop/weaviate · adapters/repos/db/vector/hnsw/commit_logger.go

Evidence

64-shard below-threshold HNSW commit-log maintenance checks at 4 CPUs · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 971.6 0.7068 −99.9% (−970.9) −991.2 to −847.3 < −48.58 PASSED

single-shard buffered HNSW tombstone WAL writes with rotations suppressed by a 1<<62-byte threshold · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 58.33 32.13 −44.6% (−26.03) −26.66 to −25.05 ≤ 2.916 PASSED
B/op 16 0 −100% (−16) −16 to −16 < −0.8 PASSED
allocs/op 1 0 −100% (−1) −1 to −1 < −0.05 PASSED

single-shard, one-CPU-pinned sparse-file HNSW tombstone writes beginning 8 bytes below the default 100 MiB threshold, then flushed and rotated after every 9-byte write · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 1673443 1641222 −1.6% (−26655) −53166 to −11838 ≤ 10000 PASSED
B/op 37738 4859 −87.1% (−32881) −32890 to −32859 < −1887 PASSED
allocs/op 49 45 −8.2% (−4) −4 to −4 < −2.45 PASSED

single-shard, one-CPU-pinned HNSW tombstone writes from a normal new file through one full default 100 MiB threshold crossing, flush, and rotation · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 1029872743 761998531 −25.8% (−265417753) −346771709 to −223817440 < −51493637 PASSED
B/op 186452760 5952 −100% (−186446808) −186446944 to −186441232 < −9322638 PASSED
allocs/op 11650923 69 −100% (−11650854) −11650856 to −11650847 < −582546 PASSED

Checks: 4 of 4 passed. Verification: no defect found.

Timeline