Bypassed VAD speech activity throttling causes excessive frame processing overhead
perfloop/pipecat · CHATTY IO
https://perfloop.ai/t/oss/case_asvh8wp7fx
Verdict
VERIFIED · settled 2026-09-16
What happened: The paired measurements met the required improvement.
Hypothesis
In VADController._handle_audio, when the user is speaking (self._vad_state == VADState.SPEAKING), the event handler on_speech_activity is called directly on every single raw audio frame (e.g., every 10-20ms) instead of calling self._maybe_speech_activity(), which implements the configured speech_activity_period throttling (default 200ms). This causes the VADProcessor to broadcast UserSpeakingFrame down the entire pipeline of processors and transports up to 50 times/sec instead of 5 times/sec, generating massive, redundant frame-processing overhead. This completely bypasses the intended throttling/coalescing logic.
Change to test: Call self._maybe_speech_activity() instead of calling self._call_event_handler("on_speech_activity") directly in VADController._handle_audio.
Where it lives
perfloop/pipecat · src/pipecat/audio/vad/vad_controller.py
Evidence
PipelineWorker VADProcessor with SileroVADAnalyzer replaying the repository's bundled 16 kHz speech trace five times as 20 ms frames with the default 200 ms speech-activity period · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
VADProcessor process CPU nanoseconds for five speech-trace replays |
770000000 |
650000000 |
−16.9% (−130000000) |
−170000000 to −110000000 |
< −38500000 |
PASSED |
VADProcessor elapsed nanoseconds for five speech-trace replays |
813448193 |
682118949 |
−17.5% (−142736115) |
−159842874 to −95137728 |
< −40672410 |
PASSED |
Checks: 6 of 6 passed. Verification: no defect found.