Repository navigation
Three timing tests take their two minima interleaved, not in blocks - #106
Merged
Merged
Conversation
TestAThreeComponentJPEGIsCachedOnItsSampleTriple failed once during other
work, reporting 8.9x against a threshold of 3x, and passed three times in
a row alone a minute later. The load average was 59.
Each of the three tests took three draws of the witness, then three of the
subject, and compared the minima. That is enough to fail on a busy
machine: the witness can take all three of its turns in a quiet moment and
the subject all three of its own under a spike, and the ratio then reports
the machine rather than the code.
MEASURED rather than argued, because the failure could not be reproduced
on demand. A throwaway test computed the ratio BOTH WAYS on the same work,
40 rounds per scheme per run, under load:
max of 40 max of 40 max of 40 over 3x
blocked 1.61 2.27 8.80 1 of 120
interleaved 1.53 1.52 1.61 0 of 120
The 8.80 is the failure, reproduced -- and close to the 8.9 that started
this. The blocked scheme also returned a ratio of 0.50 in the first run,
which says the calibrated subject ran twice as fast as the device witness:
impossible for the quantity being measured, and a direct demonstration
that its two minima come from two different moments.
So pairedBest draws the two alternately and keeps each one's minimum. Its
comment carries the numbers above, because a test whose threshold is a
ratio needs a reader to know what the ratio's own spread is.
The fourth ratio assertion in the suite is left alone:
TestAType3GlyphThatShowsItselfIsRefusedRatherThanBounded compares
OPERATION COUNTS, and says in its own comment why that is not a duration.
It cannot drift with load and does not need this.
100.0% of statements.
This was referenced Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Independent of #103 and #105 — branched from
main, touches only two test files.The flake
TestAThreeComponentJPEGIsCachedOnItsSampleTriplefailed once during otherwork, reporting 8.9× against a threshold of 3×, and passed three times in a
row alone a minute later. The load average was 59.
Each of three tests in
jpegspace_test.gotook three draws of the witness, thenthree of the subject, and compared the minima. That is enough to fail on a busy
machine: the witness can take all three of its turns in a quiet moment and the
subject all three of its own under a spike, and the ratio then reports the
machine rather than the code.
Measured, not argued
The failure could not be reproduced on demand, so instead of guessing, a
throwaway test computed the ratio both ways on the same work — 40 rounds per
scheme per run, under load:
The 8.80 is the failure, reproduced — and close to the 8.9 that started
this.
The blocked scheme also returned a ratio of 0.50 in the first run, which
says the calibrated subject ran twice as fast as the device witness. That is
impossible for the quantity being measured, and it is the plainest
demonstration that its two minima come from two different moments.
The change
pairedBest(t, witness, subject, rounds)inbuild_test.godraws the twoalternately and keeps each one's minimum; the three call sites hoist their two
documents and call it. The helper's comment carries the table above, because a
test whose threshold is a ratio needs its reader to know what that ratio's own
spread is.
The fourth ratio assertion in the suite is deliberately left alone.
TestAType3GlyphThatShowsItselfIsRefusedRatherThanBoundedcompares operationcounts and says in its own comment why that is neither a count of today's
parser nor a duration on this machine. It cannot drift with load.
gofmt,go vetclean; 100.0% of statements; no non-test file changed.🤖 Generated with Claude Code