Batch conflict-deletion durability flushes

perfloop/weaviate · UNDER BATCHING

https://perfloop.ai/t/oss/case_gnbzdw8xn2

Verdict

CLOSED · opened 2026-08-15

What happened: Closed; the assigned hypothesis was retired.

Hypothesis

I traced `AsyncReplicationScheduler.runBatch` through `runEntry`, `runHashbeatCycle`, and `Shard.hashBeat`. When propagation returns a `RepairResponse` with `Deleted` set and the deletion strategy permits resolution, `hashBeat` iterates the responses and invokes `resolveObjectConflict` once per response. Local selection is capped by `propagationLimit`, which defaults to 1,000 objects, and a cycle that sent objects is rescheduled at the default five-second propagating cadence; the source does not establish the live deletion-conflict width.

For every eligible response, `DeleteObject` calls `Store.WriteWALs` and flushes every configured vector and geo queue. `Store.WriteWALs` visits every bucket, and `Bucket.WriteWAL` documents that an explicit call writes buffered WAL data and that one call is sufficient for a larger batch. The existing `DeleteObjectBatch` path performs its per-ID deletes and then runs this WAL and queue-flush stage once, so the source shows a K-to-1 reduction in those boundary calls for K batched repair conflicts. The source does not measure whether the forced writes or flushes dominate cycle time.

A case should inject 1, 10, 100, and up to the configured limit of deletion-conflict responses into one hashbeat. It should compare hashbeat duration, `Store.WriteWALs` and queue-flush call counts, storage I/O, and CPU before and after batching. The differential test must retain per-object errors, durability before successful completion, and the race where a newer local write arrives before a time-based conflict delete.

Change to test: Collect eligible deletion-conflict responses into a repair-specific batch that retains each response's deletion timestamp and the time-based local-newer recheck. Perform WAL and vector/geo queue flushes once after that batch, while preserving per-ID errors, counters, and durability before the hashbeat reports success.

Where it lives

perfloop/weaviate · adapters/repos/db/async_replication_scheduler.go

Evidence

Timeline