Logprob streaming recopies cumulative token IDs before slicing each response delta

perfloop/dynamo · AVOIDABLE COPY

https://perfloop.ai/t/oss/case_yqda6tatak

Verdict

VERIFIED · settled 2026-09-07 · pull request opened as ai-dynamo/dynamo#14402

What happened: The paired measurements met the required improvement.

Hypothesis

TensorRT-LLM streams cumulative token IDs. In `HandlerBase._generate_locally_impl`, the response loop first creates `out["token_ids"] = output.token_ids[tokens_so_far:]`, then calls `_extract_logprobs` with the same cursor. The shared helper copies the entire cumulative list with `list(output.token_ids)` and creates another suffix only to match IDs with new logprobs. Its test describes that list as every token emitted so far.

On logprob-enabled aggregate streams, each update copies prior history again. One-token updates through N generated tokens therefore create the 1+...+N copied-history shape per output choice even though transport needs only the delta. Threading the first delta into logprob extraction removes the historical copy and duplicate suffix. The remaining floor is delta-sized logprob transformation and response construction.

A case should falsify this with long, one-token-per-update aggregate requests that request logprobs. Heap-allocation and handler-CPU profiling should show whether cumulative token-list copying disappears, while a wire-level chunk comparison checks token/logprob alignment, finish data, and error bailouts remain unchanged.

Change to test: Pass the already-built `out["token_ids"]` delta from `HandlerBase._generate_locally_impl` into `_extract_logprobs`/`extract_from_completion_output`. Use that delta to align the cursor-sliced logprob window, removing `list(output.token_ids)` and the second token-ID suffix slice while retaining the empty-logprobs fast path and full-chunk alignment bailout.

Where it lives

perfloop/dynamo · components/src/dynamo/trtllm/request_handlers/aggregated_handler.py

Evidence

256-token cumulative native TensorRT-LLM aggregate logprob updates with four interleaved choices and one-token native increments through source-bound AggregatedHandler.generate replay (handler CPU wall time) · 10 sample pairs

metric baseline candidate paired median change confidence range required result
ns/op 17710858 14201442 −19.6% (−3466246) −4391164 to −2941180 < −885543 PASSED
ops/s 56.46 70.42 +25.3% (+14.28) +11.67 to +17.13 ≥ 0 PASSED

Checks: 1 of 1 passed. Verification: no defect found.

Timeline