Charge a decode what its shape costs, not four bytes a pixel - #98
Merged
Merged
Conversation
A codec was bounded by a PIXEL count -- maxImageBytes divided by four -- as though every codestream decoded to four bytes a pixel. None does. Measured on whole pages of the corpus at 72 dpi, a JPEG 2000 decode costs 6.8 to 9.7 bytes a pixel for one component and 20.2 to 21.1 for three. Charging every shape alike was wrong in both directions, and visibly: sim_unitarian-...-1825-06-25_4_25.pdf needs 516 MB and was REFUSED, while cabepcc_000084.pdf needs 667 MB and is drawn. The bound refused pages cheaper than pages it admitted. maxDecodeBytes -- what a CODEC may spend, as against maxImageBytes which is what the picture it produces may cost to hold -- is 24 times the old pixel bound, which is exactly 1.5 GB. So a codestream charged 24 bytes a pixel has the bound it had before TO THE PIXEL: a JPEG's does not move, and neither does a three-component JPEG 2000's. What changes is that a one-component picture stops paying for colour it does not have. Three pages of the measured corpus come back drawn where they came back blank, and they are EXACTLY the three a census of blank-for-us-drawn-by-them named: 9 449 by 13 701, 8 680 by 8 680 and 8 424 by 8 400, all one component, all past the old bound. 3212 of 3215 documents are byte-identical and those three are the difference -- no page changed that was not meant to. Measured peaks for the newly admitted pages: 881 MB and 516 MB, against the 667 MB this package already tolerates for a colour page it admits. Against the judge, per document: both of the three that can be compared pixel to pixel move CLOSER to poppler, by 43.5 and 19.7 levels of 255. The third cannot be compared by that instrument because poppler renders it one column wider -- 1356 against 1355 -- which is a page-size rounding difference that predates this and is not its business. A four-component JPEG 2000 is charged as three. Nothing in the measured corpus is that shape, so that figure is a guess where the other two are readings, and it is the one place this bound is not measured.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The bound refused pages cheaper than pages it admitted
A codec was bounded by a pixel count —
maxImageBytes / 4— as though everycodestream decoded to four bytes a pixel. None does. Measured on whole pages of
the corpus at 72 dpi, minimum of three runs:
Charging every shape alike was wrong in both directions, and visibly so:
cabepcc_000084.pdf— 31.6 Mpx, coloursim_unitarian-…-1825-06-25_4_25.pdf— 75.3 Mpx, greybulletinno38tasm.pdf— 129.5 Mpx, greyNothing that was admitted is refused
maxDecodeBytes— what a codec may spend, as againstmaxImageByteswhichis what the picture it produces may cost to hold — is 24 times the old
pixel bound, which is exactly 1.5 GB. A codestream charged 24 bytes a pixel
therefore has the bound it had before to the pixel. A JPEG's does not move.
A three-component JPEG 2000's does not move. What changes is that a
one-component picture stops paying for colour it does not have.
What it draws, and whether that is better
Three pages of 3215 change, and they are exactly the three a census named
— a sweep over both corpora counting pages blank for us and drawn by the judge,
and the reverse, since a census that looks only for the failure it expects
finds the number it went looking for. The other 3212 are byte-identical.
Judged per document against poppler, before and after:
The third cannot be compared pixel to pixel by that instrument: poppler renders
it one column wider, 1356 against 1355. That is a page-size rounding
difference which predates this change and is not its business.
They are not drawn well: the conformance run puts 14.7% and 19.2% of their
pixels apart from poppler, against a corpus median near 0.2%. A page these two
disagree about is still a page, and a blank one disagrees about all of its ink —
which is 24% to 79% of these three. Closer is the claim; close is not.
Tests
Nine of nine mutations killed, coverage gate at exact 100%.
The last one took two attempts, and the reason is worth reading. A test that
charged a JPEG the one-component cost was supposed to draw a page it should
refuse — but it used stand-in bytes, so a refusal and a decode that fails
leave the same blank page, and the test could not tell them apart. It needed
real JPEG data and a control proving that data draws at its own size.
The pairs are the point throughout: the same picture at each cost, admitted and
refused; a JPEG at the old bound and one pixel past it; and
jpxSizeasked on areal one-component codestream, because the test above it substitutes
jpxSizeand a version that never reports a single component — which puts every scan back
to blank — passes without it.
Not claimed
Speed, though these pages are drawn in a fraction of poppler's time. The
machine has carried
glpi-agentat 94% of a core throughout.A four-component JPEG 2000 is charged as three. Nothing in the measured
corpus is that shape — 65 034 one-component streams and 54 130 three-component,
none other — so that figure is a guess where the other two are readings, and it
is the one place this bound is not measured.
🤖 Generated with Claude Code