RTVIObserver keeps the id of every frame no branch handles, audio included, for the whole session
perfloop/pipecat · RESOURCE LEAK
https://perfloop.ai/t/oss/case_n9kqmgc5rj
Verdict
VERIFIED · settled 2026-09-23
What happened: The assertion is violated on the comparison and satisfied with this change.
Hypothesis
Source-derived and deterministic; its size in a real session is unmeasured and is what the case measures. `RTVIObserver.on_push_frame` sets `mark_as_seen = True` before its branch chain and, at the end (observer.py:618-619), adds `frame.id` to `self._frames_seen`. Nothing in the class removes ids from that set, so it only grows until the observer is discarded. With the default `RTVIObserverParams` (`user_audio_level_enabled=False` at observer.py:184, `bot_audio_level_enabled=False` at observer.py:178), an `InputAudioRawFrame` or `TTSAudioRawFrame` matches none of the branches, yet `mark_as_seen` stays True and its id is added. The dedup check at observer.py:443 only matters for frames a branch acts on; for a frame no branch handles, a second sighting does nothing either way. `PipelineWorker` installs an `RTVIObserver` by default, so a voice session adds one id per pushed audio frame (tens per second at the usual 10 to 20 ms frame sizes) for its whole length. Proof target: with default params, drive a long stream of `InputAudioRawFrame`s and `TTSAudioRawFrame`s through `on_push_frame`; an assertion that `_frames_seen` stays bounded (for example at most 64 ids) fails before the change and passes after. The repository's RTVI and worker tests pass. Every RTVI message the observer emits for a scripted conversation (speaking events, transcriptions, LLM and TTS text, metrics) is identical before and after. Retained memory after gc over a long default-worker audio replay is measured before and after.
Change to test: Record a frame id in `_frames_seen` only when a branch of `on_push_frame` actually handled the frame (sent an RTVI message or updated observer state). A frame that no branch handles needs no dedup entry: seeing it again does nothing either way. Keep the existing dedup, the `mark_as_seen = False` retry paths and every message the observer sends unchanged. This change touches the RTVI observer only; the IdleFrameObserver's retention is a separate change and out of scope.
Where it lives
perfloop/pipecat · src/pipecat/processors/frame_processor.py
Evidence
The assertion is violated on the comparison and satisfied with this change: `With a default PipelineWorker's automatically installed RTVIObserver using default parameters, replaying 4096 InputAudioRawFrame instances and 4096 TTSAudioRawFrame instances through on_push_frame leaves at most 64 frame IDs in _frames_seen.`
default PipelineWorker RTVI audio replay (4096 input + 4096 TTS frames) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
retained_bytes |
753964 |
216 |
−100% (−753748) |
−753748 to −753748 |
< −37698 |
PASSED |
Checks: 6 of 6 passed. Verification: no defect found.