Logprob streaming recopies cumulative token IDs before slicing each response delta
perfloop/dynamo · AVOIDABLE COPY
https://perfloop.ai/t/oss/case_yqda6tatak
Verdict
VERIFIED · settled 2026-09-07 · pull request opened as ai-dynamo/dynamo#14402
What happened: The paired measurements met the required improvement.
Hypothesis
TensorRT-LLM streams cumulative token IDs. In `HandlerBase._generate_locally_impl`, the response loop first creates `out["token_ids"] = output.token_ids[tokens_so_far:]`, then calls `_extract_logprobs` with the same cursor. The shared helper copies the entire cumulative list with `list(output.token_ids)` and creates another suffix only to match IDs with new logprobs. Its test describes that list as every token emitted so far.
On logprob-enabled aggregate streams, each update copies prior history again. One-token updates through N generated tokens therefore create the 1+...+N copied-history shape per output choice even though transport needs only the delta. Threading the first delta into logprob extraction removes the historical copy and duplicate suffix. The remaining floor is delta-sized logprob transformation and response construction.
A case should falsify this with long, one-token-per-update aggregate requests that request logprobs. Heap-allocation and handler-CPU profiling should show whether cumulative token-list copying disappears, while a wire-level chunk comparison checks token/logprob alignment, finish data, and error bailouts remain unchanged.
Change to test: Pass the already-built `out["token_ids"]` delta from `HandlerBase._generate_locally_impl` into `_extract_logprobs`/`extract_from_completion_output`. Use that delta to align the cursor-sliced logprob window, removing `list(output.token_ids)` and the second token-ID suffix slice while retaining the empty-logprobs fast path and full-chunk alignment bailout.
Where it lives
perfloop/dynamo · components/src/dynamo/trtllm/request_handlers/aggregated_handler.py
Evidence
256-token cumulative native TensorRT-LLM aggregate logprob updates with four interleaved choices and one-token native increments through source-bound AggregatedHandler.generate replay (handler CPU wall time) · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
ns/op |
17710858 |
14201442 |
−19.6% (−3466246) |
−4391164 to −2941180 |
< −885543 |
PASSED |
ops/s |
56.46 |
70.42 |
+25.3% (+14.28) |
+11.67 to +17.13 |
≥ 0 |
PASSED |
Checks: 1 of 1 passed. Verification: no defect found.
Timeline
2026-09-04· Case opened2026-09-06· PR opened