Skip to content

GH-29: bound response buffering with pass-through overflow - #36

Merged
blainemotsinger merged 15 commits into
mainfrom
GH-29
Sep 25, 2026
Merged

blainemotsinger merged 15 commits into
mainfrom
GH-29

Conversation

@blainemotsinger

@blainemotsinger blainemotsinger commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Summary

Replaces unbounded bufferedRecorder response buffering with a bounded buffer plus pass-through overflow, so a concurrent burst of large completions can no longer grow gateway memory without limit.

  • New config key max_buffered_response_bytes in [gateway] (default 8388608 = 8 MiB; -1 = unlimited; 0/omitted = default).
  • attemptRecorder buffers up to the limit. Over the limit, or when upstream Content-Length exceeds it, the buffered prefix is committed to the client and the remainder streams through.
  • Once attemptResult.committed is true the handler never retries and never writes a 502 — an upstream failure yields a truncated body (same contract as the existing streaming commitGate).
  • New metric gateway_response_passthrough_total{model,reason} with reasons size_limit / content_length.
  • stream: true and request-side maxBodyBytes are unchanged.

Design

Spec: tmp/docs/sdd/2026-09-24-gh29-response-buffer-design.md
Plan: tmp/docs/sdd/2026-09-24-gh29-response-buffer-plan.md

Key invariant: committed means "bytes may have reached the client", not "commit fully succeeded". attemptRecorder.commit() sets the flag immediately after WriteHeader, matching commitGate in internal/api/stream.go, so a failed body write can never fall through to a second status/headers write.

Commits

  • dfdb73e add max_buffered_response_bytes config key
  • 841bb20 bound response buffering with pass-through overflow
  • f066d6c refuse retry and 502 after committed response
  • 7c0342f document bounded response buffering
  • 8b5bfb3 document passthrough metric and commit-once wording
  • c62f789 set committed before body write and gate passthrough metric

Testing

  • go test ./... — all packages pass
  • go test -race ./internal/api/ ./internal/config/ ./internal/metrics/
  • go test -race -count=5 ./internal/api/ — stability on go1.26.3 and go1.22.12
  • GOTOOLCHAIN=go1.22.12 go test ./... && GOTOOLCHAIN=go1.22.12 go vet ./...
  • gofmt -l . clean
  • Mutation gates verified: reverting the committed guard, the commit() ordering, or the metric gate each fail their tests

Follow-ups (not in this branch)

  • Request-side maxBodyBytes configurability
  • classifyProxyError maps response-header timeout to outcomeClientAborted, contradicting the README retry wording — pre-existing, needs its own issue
  • docs/api.md Prometheus table is stale (missing this metric and two pre-existing ones)
  • Test-gap cluster: size_limit-then-truncate untested at handler level; writeThrough tail write errors not captured to err

closes #29

…errors

classifyProxyError treated net/http's response-header timeoutError as a
client abort because that type declares
Is(err) { return err == context.DeadlineExceeded } without exposing an
Unwrap chain. The client got an empty 200 after a 60s hang and the attempt
was never retried, contradicting the documented contract. Walk the unwrap
chain for a literal context.DeadlineExceeded instead of trusting errors.Is.

writeThrough dropped client write errors that happened after commit, so a
truncated pass-through was counted as a success with no log. Capture them
in attemptRecorder.err alongside commit()'s buffered flush.
TestProxyToVLLMOverflowPassesThrough now asserts the passthrough reason is
size_limit, so dropping the Flush() that forces chunked framing fails the
test instead of silently falling through to the Content-Length fast-path.

Add the missing size_limit-then-truncate combination at the handler level,
using chunk-framed output that is never terminated so the read ends in
io.ErrUnexpectedEOF rather than the clean EOF of a close-delimited body.
The Prometheus table omitted gateway_proxy_attempts_total,
gateway_response_passthrough_total and health_check_panics_total. Document
all eight metrics and the label sets for the labeled counters.
@blainemotsinger blainemotsinger self-assigned this Sep 25, 2026
commitGate.Write returned the client write error but never recorded it, so a
stream whose client stopped reading was reported as a clean completion. Capture
it in writeErr rather than err, which already carries upstream transport and
body-read failures and drives the retry path — a failing client must not be
classified as an upstream problem or leave the node in the offline set.

Brings the streaming path in line with attemptRecorder.writeThrough.
…Recorder

attemptRecorder.err was filled by three different sources: ReverseProxy's
ErrorHandler and errCaptureBody, which report upstream failures, and a.w.Write,
which reports a client that stopped reading. The handler could not tell them
apart, so a dead client was logged as an upstream truncation and left the node
in the offline set.

Route client write failures to writeErr, matching commitGate.writeErr on the
streaming path, and clear the node instead of reporting it. An upstream failure
still takes precedence when both are set.
The streaming and pass-through paths logged the same root cause under two
different strings, which makes it harder to grep or alert on a client that
stopped reading. Collapse both onto "client write failed after commit".

The two upstream messages are untouched so existing alerts keep matching:
"stream terminated by upstream failure" and "response truncated after commit".
commitResponse discarded the error from writing a fully buffered response to
the client, so outcomeOK logged nothing at all. A client that stopped reading
was therefore visible only when the response happened to exceed the buffer
limit and take the pass-through path, or when it was streaming.

Return the write error and log it. The node answered correctly, so the offline
set and the metrics are unchanged: this is a missing log line, not a behavior
change.
The three paths that log "client write failed after commit" were the only
signal that a client stopped reading, and logs do not alert well.
gateway_response_client_write_failures_total{model} makes the rate visible.

gateway_requests_total keeps counting these as 200: it records what the
gateway did, and the gateway did commit a response. Delivery to the client is
a separate concern, so it gets a separate counter rather than overloading the
status label and breaking existing series.
The request body cap was a hardcoded package variable in the handler, so
operators could not raise it for large prompts or lower it to protect the
gateway. Expose it as max_request_body_bytes alongside
max_buffered_response_bytes.

A request body cap is a resource control, so unlike the response buffer there
is no unlimited setting: negative values are rejected and 0 keeps the previous
32 MiB default.
@blainemotsinger
blainemotsinger merged commit 79aee05 into main Sep 25, 2026
4 checks passed
@blainemotsinger
blainemotsinger deleted the GH-29 branch September 25, 2026 17:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bound proxy response buffering (memory per attempt)

1 participant