Resolve `previous_response_id` to a prompt prefix instead of replaying every Responses item
perfloop/ds4 · REDUNDANT SERIALIZATION
https://perfloop.ai/t/oss/case_avefhzk2jy
Verdict
VERIFIED · settled 2026-08-30
What happened: The paired measurements met the required improvement.
Hypothesis
`client_main` currently passes the entire body to `parse_responses_request`, whose array-input path makes `parse_responses_input` walk every replayed item, decode its strings through `json_string` and growable buffers, construct transient `chat_msgs`, then render and tokenize the whole history. Only after that conversion does `generate_job` get a chance to recognize a `responses-visible` live prefix and prefill just a suffix. The source explicitly rejects non-null `previous_response_id` and `conversation`, emits `resp_` ids only while serializing a response, and retains only a per-slot live frontier, while the Responses note documents clients that resend full visible replay; consequently the same semantic history is repeatedly JSON-encoded, decoded, rendered, and tokenized before the existing prefix machinery can reuse it. A durable response-id prefix index would remove that historical conversion class on stateful continuations while preserving hidden reasoning and tool-call bindings through the saved exact checkpoint; its hit-path floor is state lookup, validation, delta parsing, tail rendering/tokenization, and any checkpoint load or suffix prefill, while first turns, misses, and edited branches keep the current full replay behavior. A later controlled growing-conversation `/v1/responses` trace should separately record request bytes, client-thread parse/render/tokenize allocation and CPU, state-hit rate, and cached versus prefetched tokens; it falsifies this ceiling if stateful continuations are not materially exercised or those history-sized ingress terms do not become bounded by the new tail.
Change to test: Generate one stable `resp_` id before either streaming or non-streaming emission and maintain a bounded, tenant-scoped append-only response-state index keyed by it. Each record should retain the exact token/KV checkpoint reference, prompt syntax and configuration fingerprint, continuation rendering boundary, visible/model prefix metadata, and generated call ids. For a request carrying `previous_response_id`, resolve and validate that record, parse only the supplied delta input, render and tokenize only a generic continuation tail, and synchronize that tail onto the saved prefix; retain the existing full-input replay path for stateless requests, expired or mismatched records, and explicit edits.
Where it lives
perfloop/ds4 · ds4_server.c
Evidence
in-memory steady-state 8k-token Responses response-id host ingress with explicit max_output_tokens=1 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
host_ingress_ns/op |
3344668 |
406000 |
−87.8% (−2935367) |
−3045565 to −2857818 |
< −167233 |
PASSED |
disk-backed 64-turn stateless full-visible-replay Responses host ingress with explicit max_output_tokens=1 · 10 sample pairs
| metric | baseline | candidate | paired median change | confidence range | required | result |
|---|---|---|---|---|---|---|
host_ingress_ns/op |
3570595 |
3233399 |
−9.8% (−349333) |
−433463 to −232656 |
≤ 0 |
PASSED |
Checks: 5 of 5 passed. Verification: no defect found.
Timeline
2026-08-08· Case opened2026-08-09· PR opened