Stop the fast client spinning forever on chunked responses larger than its buffer
perfloop/fortio · NON PROGRESS LOOP
https://perfloop.ai/t/oss/case_pj53ak1dd3
Verdict
VERIFIED · settled 2026-09-27 · pull request opened as fortio/fortio#1168
What happened: The assertion is violated on the comparison and satisfied with this change.
Hypothesis
A chunked HTTP response whose first chunk exceeds Fortio's fast-client buffer can leave a load-test request cycling without consuming any more bytes. The request cannot finish while the loop stays in that state. How often load-test targets produce such responses, and the CPU cost in a real run, are not independently measured here.
The benefiting operation is `fortio load` against an HTTP/1.1 endpoint sending a size-prefixed first chunk larger than the configured response buffer, at revision 5c19725ff61c9f7ad944b91ec32d96a399341d87. The CLI defaults to the fast client and keep-alive (bincommon/commonflags.go:43-48,186-195); its `-httpbufferkb` flag controls the buffer, which defaults to 128 KiB (bincommon/commonflags.go:106-107; fhttp/http_client.go:69-77,853). `cli.fortioLoad` selects `fhttp.RunHTTPTest` for an HTTP URL (cli/fortio_main.go:433-475). `RunHTTPTest` constructs a client per thread and calls `StreamFetch` on warmup (fhttp/httprunner.go:146-197); every scheduled `HTTPRunnerResults.Run` calls `StreamFetch` (fhttp/httprunner.go:62-74), which invokes `readResponse` (fhttp/http_client.go:991-1055). A raw server returning `Transfer-Encoding: chunked` with an 8 KiB first chunk and `-httpbufferkb 4` is the required reproducible input; the default 128 KiB configuration encounters the same boundary when the first chunk plus headers exceeds 128 KiB. The prevalence of either input in real load tests is unknown.
`readResponse` fills `c.buffer[c.size:]` (fhttp/http_client.go:1110-1115). Once it parses the first chunk size, it logs a warning and caps `maxV` to the buffer length if that chunk will not fit (fhttp/http_client.go:1207-1241). When the buffer fills, it asks `ParseChunkSize` to parse the next size from `c.buffer[maxV:c.size]`, which is empty; the parser returns -1 and the loop continues (fhttp/http_client.go:1252-1274; fhttp/http_utils.go:188-202). On the next iteration `DelayedErrorReader` passes the empty `c.buffer[c.size:]` to the socket (fhttp/http_client.go:1078-1088,1113-1115). The checked Go/Linux TCP implementation returns 0,nil immediately for a zero-length read. Neither `c.size`, the remaining chunk bytes, nor `maxV` changes; the read deadline set in `StreamFetch` (fhttp/http_client.go:1015-1016) cannot end that immediate-return cycle. This claim concerns the first oversized chunk on the default plain-HTTP fast-client path, not an observed production incident.
The source-inferred cost is an unbounded number of parse/read/continue cycles for each eligible response, starting on the first full buffer. With one client (`-c 1`), one affected warmup or scheduled request can prevent the load run from finishing; a stuck worker remains runnable rather than waiting for more network data on this TCP implementation. The submitter reports that with a 4 KiB buffer, an 8 KiB first chunk and a 300 ms timeout, `Fetch` failed to return in three of three local attempts after more than two seconds. That report is not a profile or an independently reproduced measurement. Assuming the immediate empty-read behavior, rejecting a full buffer before `Read` should eliminate the busy cycles and permit the run to finish after a bounded response; the CPU-seconds saved would grow with the time a stuck worker would otherwise remain active. CPU share, affected-call rate and the actual workload-level saving remain unmeasured.
The falsifying check is a watchdog-bounded `fhttp` test using the raw chunked server, a 4 KiB buffer, an 8 KiB first chunk, and two `Fetch` calls on one fast client: both calls must return without retaining the incomplete socket. Also run the ordinary `fortio load -c 1 -n 2 -timeout 300ms -httpbufferkb 4 <URL>` against that endpoint before and after the change, recording completion, response counts, wall time and process CPU time. The claim is supported if the baseline stalls with repeated no-progress reads and the changed load run terminates without sustained busy CPU; it is rejected if the baseline makes progress, or the changed client still stalls or consumes the same busy-loop resources.
Change to test: In `FastClient.readResponse`, before `conn.Read`, if `c.size == len(c.buffer)`, set `keepAlive = false` and break so the incomplete socket is closed instead of reading into an empty slice. Add a watchdog-bounded first-oversized-chunk test with two fetches on the same client.
Where it lives
perfloop/fortio · cli/fortio_main.go
Evidence
The assertion is violated on the comparison and satisfied with this change: `For Fortio fhttp FastClient.Fetch with a 4 KiB buffer and 1 s request timeout, both sequential calls on one client return HTTP 200 for each complete local HTTP/1.1 chunked response layout from TestFastClientOversizedChunks: one 8 KiB first chunk, and a flushed 1-byte first chunk followed by an 8 KiB second chunk; the focused native test completes within its 4 s timeout.`
Checks: 3 of 3 passed. Verification: no defect found.
Timeline
2026-09-26· Case opened2026-09-27· PR opened