Skip to content

fix(netty): stop inflating a suspended HTTP/1.1 response - #2362

Open
mkurz wants to merge 1 commit into
AsyncHttpClient:mainfrom
mkurz:fix/demand-bounded-decompression
Open

mkurz wants to merge 1 commit into
AsyncHttpClient:mainfrom
mkurz:fix/demand-bounded-decompression

Conversation

@mkurz

@mkurz mkurz commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Summary

  • While an HTTP/1.1 response is suspended through ResponseBodyControl, stop decompressing its body. Decompressed body parts are at most 64 KiB, so a handler that suspends after each part receives one part per resume().
  • Http1ContentDecompressor becomes a ChannelDuplexHandler on Netty's pull-based Decompressor, instead of extending Netty's HttpContentDecompressor. Codings, multi-member gzip, header rewrites and the maxDecompressedResponseSize limit are kept.
  • HTTP/2 is unchanged; see below.

Problem

Suspending a response stops further socket reads, but Netty's HttpContentDecompressor inflates every chunk completely as soon as it is read. So one socket read of a highly compressible body still became many parts that nobody had asked for.

For example, one 64 KiB read of gzipped zeros (about 1032:1) became some 67 MB of body parts for a handler that had suspended the response. Only maxDecompressedResponseSize (256 MiB by default) bounded it.

With 64 MiB of gzipped zeros and a handler that suspends inside each body callback:

  • main delivered 1,955,380 bytes in 30 more parts past the suspending part;
  • after that, it delivered up to 33,713,376 bytes per resume, so the whole body arrived in 2–3 resumes.

AHC 2 behaved the same way, so this is hardening, not a regression fix. It matters to consumers that bound their buffering through ResponseBodyControl, such as Play WS's streaming API.

Change

Http1ContentDecompressor uses Netty's pull-based Decompressor (Netty 4.2.17+, AHC pins 4.2.18):

  • Codings: gzip and x-gzip, including bodies made of several gzip members (RFC 1952, section 2.2) as before; deflate and x-deflate; br and zstd when available; and snappy.
  • Kept behavior: the header rewrites, the 100 Continue pass-through, and the maxDecompressedResponseSize accounting and message. The limit counts every gzip member.
  • Output in bounded parts: it hands the body on in parts of at most 64 KiB. Decompressor output larger than that, such as zstd's when Netty's io.netty.compression.defaultMaxForwardBytes is raised, is split into retained slices. A slice shares the larger buffer until the slice and the rest have both been released, so the 64 KiB bound applies to delivered parts, not to allocations. Before each part, the handler checks whether the response is suspended.
  • While suspended: the compressed input stays in the handler, together with every message that arrived after it. To fill a part, the handler may take one more decompressor output buffer than it hands on, so decompression can run ahead of delivery by one such buffer (at most 64 KiB with Netty's defaults).
  • On resume: the handler intercepts the read request that resume() issues. It continues decompression first, and only then lets the read reach the socket. This works the same with auto-read on and off.
  • Without suspension: each part goes straight through, as before, and a single resume can deliver many parts.
  • Connection close: if the connection closes while input is held back, the rest is decompressed and delivered before channelInactive, even to a suspended response, because Netty tears the pipeline down with the channel.
  • End of the exchange: once it has ended (cancel, ABORT, timeout), held input is dropped instead of decompressed.
  • Small outputs: they are copied together into parts of up to 64 KiB, so part counts match main.

A suspension made from another thread takes effect when its queued task runs. Until then, the event loop can keep reading and decompressing, as before.

Why HTTP/2 is unchanged

After END_STREAM, Netty's HTTP/2 stream channel hands every queued frame over on the next read, regardless of autoRead, and then closes. So a handler in the stream pipeline cannot hold compressed input back, and would have to inflate up to a whole stream window at once (16 MiB by default). That is exactly the decompression-bomb case, so a partial fix would be misleading. It needs either a Netty change (honor demand after END_STREAM and close only when drained) or AHC delivering the body independently of the stream channel.

Behavior changes

  • Suspended HTTP/1.1 responses receive no further part until they resume. A handler that suspends after each part receives one part of at most 64 KiB per resume.
  • Parts are now at most 64 KiB; before, they were usually just over 64 KiB.
  • Framing errors on the final chunk, such as a bad chunk size, now fail the request. Netty's decoder used to turn such a failure into a successful, truncated response.
  • The gzip header check is stricter: the pull decoder checks both gzip magic bytes.
  • -Dio.netty.noJdkZlibDecoder is no longer honored on HTTP/1.1; the JDK's zlib is always used.
  • A corrupt gzip body may deliver fewer bytes before it fails. It fails in the same cases as before, but output decompressed in the same step that detects the damage (for example a checksum mismatch, or trailing data that is not another member) is dropped instead of handed on first.
  • Unchanged, compared with main on the same inputs:
    • several gzip members are decompressed in full, and a corrupt checksum or length in a later member fails the request;
    • bytes after a gzip member must form another member, so other trailing data fails the request, unless it is shorter than a gzip header;
    • bytes after the end of a deflate stream are dropped;
    • truncated or empty compressed bodies are accepted. That includes corrupt compressed data that makes the stream look truncated.
  • Performance without suspension: median elapsed time of 10 × 64 MiB over loopback on JDK 11, from two alternating runs. Measured on the first revision of this change; the later multi-member and part-splitting fixes do not touch the inflate path.
    • Gzipped zeros: unchanged (117–120 ms on main, 120 ms here).
    • Gzipped text: about 3–7% slower (135–143 ms against 145–147 ms), which comes from Netty's pull decompressor making more, smaller inflate calls.
    • Plain bodies: unchanged.

Compatibility

  • API: Http1ContentDecompressor is a public class, so this is a binary-incompatible change to it: Revapi reports its changed superclass, plus the members inherited from Netty's decoder and its hooks. AHC only uses it internally (ChannelManager instantiates it, and its public constructor is unchanged), so nothing changes for code that uses AHC's client API. Each difference is listed as an exact Revapi entry with a justification in pom.xml, so any other change to the class is still flagged.
  • Netty: Netty still marks its Decompressor API as unstable (@UnstableApi) while it migrates its codecs to it (netty/netty#16743). A future Netty release may require changes to this handler.

AI disclosure

Claude Code on behalf of Matthias Kurz. The commit includes Co-Authored-By per AGENTS.md.

Test plan

New tests:

  • Http1DecompressionSuspensionTest (7 tests, live server):

    • the bound with auto-read on and off: full length, all-zeros content;
    • connection reuse, a chunked body with trailers, and a body ended by connection close;
    • cancel, ABORT, and HEAD and empty gzip responses.

    On main, its two bound tests fail deterministically.

  • Http1ContentDecompressorTest (19 tests, EmbeddedChannel):

    • every encoding, the header rewrites, and unsupported encodings passed through;
    • bodiless responses, trailers and 100 Continue;
    • one part per read, and suspension before the body;
    • close while input is held, removal releasing held input, and the size limit on the resume path;
    • gzip bodies of two members, in one buffer, in 1-byte fragments and split at the member boundary;
    • suspension across a member boundary, a corrupt checksum or length in the second member, trailing data after gzip and deflate streams, and the size limit across members;
    • every coding with fragmented input and a resume per part, with auto-read on and off.
  • Http1ContentDecompressorLargeOutputTest (1 test):

    • starts a fresh JVM with -Dio.netty.compression.defaultMaxForwardBytes=262144, because Netty reads that property once per JVM. The child writes to a file, the test fails if the child has not finished after 60 seconds, and the child is killed in any case. With a deliberately hanging child, the test failed after 64 seconds and left no process behind;
    • first checks that zstd output in that JVM really exceeds 64 KiB, then that parts stay within 64 KiB, one per resume;
    • checks that a tracking allocator sees every buffer released, also when the handler is removed while it holds the rest of a split output.

Broken variants, each run against these 27 tests:

  • Without the suspension check before each part: 11 failures and 1 error, including both bound tests.
  • Read requests passed straight through, without continuing decompression first: 9 failures and 2 errors, including both bound tests and the chunked-body test.
  • Held input not decompressed when the connection closes: 2 failures, the close-delimited body test and the channelInactive ordering test.
  • Without decoding concatenated gzip members: 6 failures, all in the new gzip member tests.
  • Without splitting output larger than a part: the fresh-JVM test fails.

Runs:

  • Targeted tests (the 27 above, plus Http1DecompressionLimitTest, AutomaticDecompressionTest, Http2ContentDecompressorTest, ResponseBodyControlTest, Expect100ContinueTest and LargeResponseTest) on JDK 11: 60 tests passed.
  • ./mvnw -B -ntp -Dgpg.skip=true clean verify on JDK 11: 1,818 tests, 0 failures, 0 errors, 22 skipped; Revapi passed. No test-skipping flags were used.
  • The same 60 targeted tests, compiled on JDK 11 and run on JDK 17, 21 and 25 (Maven Surefire's -Djvm): 0 failures on each. The fresh-JVM test's child runs on the same JDK.
  • Compared with Netty's HttpContentDecompressor (what main uses) on 15 kinds of input, each in one buffer and in 1-byte fragments (concatenated, truncated and corrupt gzip members, trailing data after gzip, zlib and raw deflate streams): every input succeeds or fails as on main, with the same body when it succeeds. In 5 of the 30 runs both fail, and this branch hands on fewer bytes before the failure.
  • All four open AHC changes merged together (raw Cookie headers on retries, ResponseBodyControl.execute, suspend/resume order, demand-bounded decompression): they merge without conflicts, together and in every pair, and ./mvnw -B -ntp -Dgpg.skip=true clean install on JDK 11 passes with 1,853 tests, 0 failures, 0 errors, 22 skipped; Revapi passed.
  • Play WS against that merged build, Scala 2.13 and 3.3, JDK 17: all tests pass.
    • Its gzip streaming test demands one part early from another thread. It finds 0 bytes delivered beyond demand, against 1,875,916 bytes with AHC 3.0.14. It measures delivered parts, not memory inside the decompressor.
    • Play WS's shaded jar includes Netty's pull decompressors.

Motivation:
HTTP/1.1 automatic decompression ran Netty's HttpContentDecompressor,
which inflates every chunk completely as soon as it is read. Suspending
a response through ResponseBodyControl stops further socket reads, but
not that. One 64 KiB read of a highly compressible gzip body (about
1032:1) still became some 67 MB of body parts for a handler that had
asked for none; only maxDecompressedResponseSize (256 MiB) bounded it.
AHC 2 behaved the same way, so this is hardening, not a regression fix.

Modification:
Http1ContentDecompressor is now a ChannelDuplexHandler built on
Netty's pull-based Decompressor (Netty 4.2.17+; Netty still marks this
API unstable). It keeps the codings (gzip, including bodies of several
gzip members, deflate, br, snappy, zstd), the header rewrites and the
maxDecompressedResponseSize accounting. It hands the body on in parts
of at most 64 KiB, splitting larger decompressor output, and checks
before each part whether the response is suspended. If it is, the
compressed input stays in the handler, and so does every message
behind it. The read request that resume() issues continues
decompression before it reaches the socket. Without suspension, each
part goes straight through as before.

The pipeline is torn down with its channel, so if the connection
closes while input is held back, the rest is inflated and delivered
before channelInactive. Once the exchange has ended (cancel, ABORT,
timeout), held input is dropped instead of inflated. A framing failure
on the final chunk now reaches the handler; Netty's decoder used to
replace it with success.

Revapi reports the changed superclass of this public handler class and
the members inherited from Netty's decoder; they are justified in the
pom.

HTTP/2 is unchanged. After END_STREAM, Netty's stream channel hands
every queued frame over on the next read and then closes. A handler in
the stream pipeline therefore cannot hold compressed input back.

Result:
A response whose handler suspends it after each body part receives one
part of at most 64 KiB per resume, and decompression runs ahead of the
delivered parts by at most one decompressor output buffer. For 64 MiB
of gzipped zeros, main delivered 1.96 MB in 30 more parts past the
suspending part, then up to 33.7 MB per resume; now nothing past it,
and 65,536 bytes per resume.

Http1DecompressionSuspensionTest covers auto-read on and off, chunked
bodies with trailers, close-delimited bodies, cancel, ABORT, bodiless
responses and connection reuse; its two bound tests fail on main.
Http1ContentDecompressorTest covers the handler on its own, including
gzip bodies of several members, corrupt and trailing data, and every
coding across repeated suspension.
Http1ContentDecompressorLargeOutputTest covers decompressor output
larger than a part, in a JVM with Netty's
io.netty.compression.defaultMaxForwardBytes raised. Loopback throughput
without suspension: unchanged for gzipped zeros, about 5% more elapsed
time for gzipped text.

Claude Code on behalf of Matthias Kurz

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mkurz

mkurz commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Heads-up: #2357, opened before this PR, which I missed, rewrites the same Http1ContentDecompressor for a different goal: no per-response EmbeddedChannel and Inflater, and no copy of direct input. The two conflict, but the goals are complementary. #2357 still inflates a suspended response in full, and this PR doesn't pool inflaters or avoid the copy.

If you'd like #2357 first, I'd rebuild the demand bounding on top of its inflate loop instead of Netty's pull Decompressor. That drops this PR's dependency on Netty's @UnstableApi for gzip and deflate, and keeps the pull decompressors only for br, zstd and snappy. Happy to coordinate with @pavel-ptashyts either way.

@mkurz

mkurz commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Related to "Why HTTP/2 is unchanged" above: netty/netty#17774 makes Http2StreamChannel respect auto-read and the read limits for frames queued at END_STREAM. Once that is in a Netty release AHC uses, an HTTP/2 counterpart of this change becomes possible; output from the final DATA frame would still need handling in AHC.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant