Bypassed VAD speech activity throttling causes excessive frame processing overhead

perfloop/pipecat · CHATTY IO

https://perfloop.ai/t/oss/case_asvh8wp7fx

Verdict

VERIFIED · settled 2026-09-16

What happened: The paired measurements met the required improvement.

Hypothesis

In VADController._handle_audio, when the user is speaking (self._vad_state == VADState.SPEAKING), the event handler on_speech_activity is called directly on every single raw audio frame (e.g., every 10-20ms) instead of calling self._maybe_speech_activity(), which implements the configured speech_activity_period throttling (default 200ms). This causes the VADProcessor to broadcast UserSpeakingFrame down the entire pipeline of processors and transports up to 50 times/sec instead of 5 times/sec, generating massive, redundant frame-processing overhead. This completely bypasses the intended throttling/coalescing logic.

Change to test: Call self._maybe_speech_activity() instead of calling self._call_event_handler("on_speech_activity") directly in VADController._handle_audio.

Where it lives

perfloop/pipecat · src/pipecat/audio/vad/vad_controller.py

Evidence

PipelineWorker VADProcessor with SileroVADAnalyzer replaying the repository's bundled 16 kHz speech trace five times as 20 ms frames with the default 200 ms speech-activity period · 10 sample pairs

metric baseline candidate paired median change confidence range required result
VADProcessor process CPU nanoseconds for five speech-trace replays 770000000 650000000 −16.9% (−130000000) −170000000 to −110000000 < −38500000 PASSED
VADProcessor elapsed nanoseconds for five speech-trace replays 813448193 682118949 −17.5% (−142736115) −159842874 to −95137728 < −40672410 PASSED

Checks: 6 of 6 passed. Verification: no defect found.

Timeline