Avoid high-frequency heap allocation and buffer copying of immutable bytes in VAD
perfloop/pipecat · ALLOCATION HOT LOOP
https://perfloop.ai/t/oss/case_ak17b7sxar
Verdict
VERIFIED · settled 2026-09-17 · merged as pipecat-ai/pipecat#5811
What happened: The paired measurements met the required improvement.
Hypothesis
In src/pipecat/audio/vad.VADAnalyzer._run_analyzer, which processes real-time audio frames at high frequency (typically every 10-30ms), self._vad_buffer is managed as an immutable bytes object. On every frame, self._vad_buffer += buffer creates a new bytes object and copies the old buffer contents. Within the processing loop, self._vad_buffer = self._vad_buffer[num_required_bytes:] similarly allocates a new bytes object and copies all remaining bytes. This creates a high rate of heap allocations and memory copying on the audio hot path. Transitioning to a mutable bytearray with .extend() and del (or a read index) eliminates these redundant allocations.
Change to test: Initialize _vad_buffer as a mutable bytearray instead of immutable bytes, append incoming chunks using .extend(), and extract or delete analyzed slices in-place using del or a sliding read index to avoid creating new heap allocations and copying the entire remaining buffer on every frame.
Where it lives
perfloop/pipecat · src/pipecat/audio/vad/vad_controller.py
Evidence
PipelineWorker with VADProcessor and the real SileroVADAnalyzer replaying scripts/provider-watch/assets/speech-16k.wav (mono 16 kHz) as 20 ms, 640-byte InputAudioRawFrame chunks in one continuous 50-trace precision sample (ten identical five-trace batches), with the ordinary PipelineWorker RTVI integration and user-audio-level observer enabled on every frame at a 0.15-second reporting period, using the default VAD settings; CPU is reported per five-trace batch · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
AudioVolumeTracker rolling-window payload bytes copied per update |
26388 |
13588 |
−48.5% (−12800) |
−12800 to −12800 |
< −1319 |
PASSED |
AudioVolumeTracker rolling-buffer allocation events per update |
2 |
0 |
−100% (−2) |
−2 to −2 |
< −0.1 |
PASSED |
VAD and RTVI process CPU nanoseconds per five-trace replay batch |
816000000 |
812500000 |
−0.6% (−5000000) |
−14000000 to +2000000 |
≤ 13400000 |
PASSED |
Checks: 6 of 6 passed. Verification: no defect found.