Avoid high-frequency heap allocation and buffer copying of immutable bytes in VAD

perfloop/pipecat · ALLOCATION HOT LOOP

https://perfloop.ai/t/oss/case_ak17b7sxar

Verdict

VERIFIED · settled 2026-09-17 · merged as pipecat-ai/pipecat#5811

What happened: The paired measurements met the required improvement.

Hypothesis

In src/pipecat/audio/vad.VADAnalyzer._run_analyzer, which processes real-time audio frames at high frequency (typically every 10-30ms), self._vad_buffer is managed as an immutable bytes object. On every frame, self._vad_buffer += buffer creates a new bytes object and copies the old buffer contents. Within the processing loop, self._vad_buffer = self._vad_buffer[num_required_bytes:] similarly allocates a new bytes object and copies all remaining bytes. This creates a high rate of heap allocations and memory copying on the audio hot path. Transitioning to a mutable bytearray with .extend() and del (or a read index) eliminates these redundant allocations.

Change to test: Initialize _vad_buffer as a mutable bytearray instead of immutable bytes, append incoming chunks using .extend(), and extract or delete analyzed slices in-place using del or a sliding read index to avoid creating new heap allocations and copying the entire remaining buffer on every frame.

Where it lives

perfloop/pipecat · src/pipecat/audio/vad/vad_controller.py

Evidence

PipelineWorker with VADProcessor and the real SileroVADAnalyzer replaying scripts/provider-watch/assets/speech-16k.wav (mono 16 kHz) as 20 ms, 640-byte InputAudioRawFrame chunks in one continuous 50-trace precision sample (ten identical five-trace batches), with the ordinary PipelineWorker RTVI integration and user-audio-level observer enabled on every frame at a 0.15-second reporting period, using the default VAD settings; CPU is reported per five-trace batch · 10 sample pairs

metric baseline candidate paired median change confidence range required result
AudioVolumeTracker rolling-window payload bytes copied per update 26388 13588 −48.5% (−12800) −12800 to −12800 < −1319 PASSED
AudioVolumeTracker rolling-buffer allocation events per update 2 0 −100% (−2) −2 to −2 < −0.1 PASSED
VAD and RTVI process CPU nanoseconds per five-trace replay batch 816000000 812500000 −0.6% (−5000000) −14000000 to +2000000 ≤ 13400000 PASSED

Checks: 6 of 6 passed. Verification: no defect found.

Timeline