Repository navigation
fix(vault): read test responses in full - #452
Open
LKSNDRTMLKV wants to merge 1 commit into
Open
LKSNDRTMLKV wants to merge 1 commit into
LKSNDRTMLKV wants to merge 1 commit into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #336.
What was wrong
publish_serve_cycle,textileandsuspensionhung on the request sent straight afterPOST …/publish— about one run in a thousand on Linux — until the client timeout. It was never in the vault. It was the test client.Server::encode status=200, body=Known(10037)).hyperreads a socket 8 KiB at a time, so when the headers arrive ~2 KiB are still unread, and thereqwest::Responseowns its connection until someone reads them or drops it.resp.status()and then writelet resp = …for the next call. A shadowed binding is not dropped, so the publish response lives to the end of the test.hyper-util's pool hands the same still-busy connection to the next request, which queues behind a body nobody is going to read while the test waits for it. Nothing is running anywhere.This matches every sighting: all were the request right after a publish, and in the first one the machine was idle for the last ~114 s (380 of 381 tests had finished at 02:42:26; the hung one was killed at 02:44:20).
Evidence
Reproduced on Linux (
rust:1.96.0-bookworm, 4 CPUs, sharedpostgres:17), runningtest_suspension_flowas fresh processes four at a time: 3 hits in ~3,700 runs. Withhyper's own tracing on, the failing run shows the mechanism directly:A diagnostic added for this run reported, at the timeout: router has no request in flight; last answered
POST /dpp → 201,POST …/publish → 200; server-side socket empty; client-side socket holding 2,022 unread bytes (10,214 sent − 8,192 read). Thesuspendrequest never reached the router.The change
TestClientreads every response to the end before returning it, so a test cannot hold a response that still owns a connection. It returns an ordinaryreqwest::Responseover the buffered bytes; no test changed (the suite only usesstatus,json,text,headers).tests/test_client.rspins the invariant: a server sends the head and the first 8 KiB, pauses, then sends the rest;get()must not return before the tail arrives. It fails (in 3.8 ms) with the buffering removed and passes with it.Verification — and what is not done
just checkrc=0 (1,487 unit tests),just lint-integrationrc=0, vault integration tier 476 passed / 12 skipped.hyper-utilre-offering a mid-body connection is intended behaviour or a defect. The harness no longer depends on it.Not covered
dpp-node/tests/smoke.rsanddpp-resolver/tests/resolver_e2e.rsbuild their own sharedreqwest::Clients and shadowlet respthe same way. They may carry the same latent hang; unconfirmed, and out of scope for #336 (vault tier).Picking this up
connectingvsno response in 45s; <router state>), which is what the old message could not.suspensionandtextilein a 4-CPUrust:1.96.0-bookwormcontainer, pointODAL_TEST_PG_ADMIN_URLat a Postgres with--shm-size=512m, and run--exact test_suspension_flowas fresh processes 4 at a time. Recreate that Postgres every few thousand runs — it does not drop per-test databases. Unfixed, a hit takes ~5–10 minutes.