Skip to content

Charge a decode what its shape costs, not four bytes a pixel - #98

Merged
tannevaled merged 1 commit into
mainfrom
decode-cost-by-shape
Sep 29, 2026
Merged

tannevaled merged 1 commit into
mainfrom
decode-cost-by-shape

Conversation

@tannevaled

Copy link
Copy Markdown
Contributor

The bound refused pages cheaper than pages it admitted

A codec was bounded by a pixel count — maxImageBytes / 4 — as though every
codestream decoded to four bytes a pixel. None does. Measured on whole pages of
the corpus at 72 dpi, minimum of three runs:

decode bytes a pixel
JPEG 2000, one component 6.8, 6.9, 7.7, 9.7
JPEG 2000, three components 20.2, 21.1

Charging every shape alike was wrong in both directions, and visibly so:

page peak today
cabepcc_000084.pdf — 31.6 Mpx, colour 667 MB drawn
sim_unitarian-…-1825-06-25_4_25.pdf — 75.3 Mpx, grey 516 MB refused
bulletinno38tasm.pdf — 129.5 Mpx, grey 881 MB refused

Nothing that was admitted is refused

maxDecodeBytes — what a codec may spend, as against maxImageBytes which
is what the picture it produces may cost to hold — is 24 times the old
pixel bound, which is exactly 1.5 GB. A codestream charged 24 bytes a pixel
therefore has the bound it had before to the pixel. A JPEG's does not move.
A three-component JPEG 2000's does not move. What changes is that a
one-component picture stops paying for colour it does not have.

What it draws, and whether that is better

Three pages of 3215 change, and they are exactly the three a census named
— a sweep over both corpora counting pages blank for us and drawn by the judge,
and the reverse, since a census that looks only for the failure it expects
finds the number it went looking for. The other 3212 are byte-identical.

Judged per document against poppler, before and after:

moved closer to the judge 2 of 2 comparable
moved further 0
by 43.5 and 19.7 levels of 255

The third cannot be compared pixel to pixel by that instrument: poppler renders
it one column wider, 1356 against 1355. That is a page-size rounding
difference which predates this change and is not its business.

They are not drawn well: the conformance run puts 14.7% and 19.2% of their
pixels apart from poppler, against a corpus median near 0.2%. A page these two
disagree about is still a page, and a blank one disagrees about all of its ink —
which is 24% to 79% of these three. Closer is the claim; close is not.

Tests

Nine of nine mutations killed, coverage gate at exact 100%.

The last one took two attempts, and the reason is worth reading. A test that
charged a JPEG the one-component cost was supposed to draw a page it should
refuse — but it used stand-in bytes, so a refusal and a decode that fails
leave the same blank page
, and the test could not tell them apart. It needed
real JPEG data and a control proving that data draws at its own size.

The pairs are the point throughout: the same picture at each cost, admitted and
refused; a JPEG at the old bound and one pixel past it; and jpxSize asked on a
real one-component codestream, because the test above it substitutes jpxSize
and a version that never reports a single component — which puts every scan back
to blank — passes without it.

Not claimed

Speed, though these pages are drawn in a fraction of poppler's time. The
machine has carried glpi-agent at 94% of a core throughout.

A four-component JPEG 2000 is charged as three. Nothing in the measured
corpus is that shape — 65 034 one-component streams and 54 130 three-component,
none other — so that figure is a guess where the other two are readings, and it
is the one place this bound is not measured.

🤖 Generated with Claude Code

A codec was bounded by a PIXEL count -- maxImageBytes divided by four -- as
though every codestream decoded to four bytes a pixel. None does. Measured on
whole pages of the corpus at 72 dpi, a JPEG 2000 decode costs 6.8 to 9.7 bytes
a pixel for one component and 20.2 to 21.1 for three.

Charging every shape alike was wrong in both directions, and visibly:
sim_unitarian-...-1825-06-25_4_25.pdf needs 516 MB and was REFUSED, while
cabepcc_000084.pdf needs 667 MB and is drawn. The bound refused pages cheaper
than pages it admitted.

maxDecodeBytes -- what a CODEC may spend, as against maxImageBytes which is
what the picture it produces may cost to hold -- is 24 times the old pixel
bound, which is exactly 1.5 GB. So a codestream charged 24 bytes a pixel has
the bound it had before TO THE PIXEL: a JPEG's does not move, and neither does
a three-component JPEG 2000's. What changes is that a one-component picture
stops paying for colour it does not have.

Three pages of the measured corpus come back drawn where they came back blank,
and they are EXACTLY the three a census of blank-for-us-drawn-by-them named:
9 449 by 13 701, 8 680 by 8 680 and 8 424 by 8 400, all one component, all past
the old bound. 3212 of 3215 documents are byte-identical and those three are
the difference -- no page changed that was not meant to.

Measured peaks for the newly admitted pages: 881 MB and 516 MB, against the
667 MB this package already tolerates for a colour page it admits.

Against the judge, per document: both of the three that can be compared pixel
to pixel move CLOSER to poppler, by 43.5 and 19.7 levels of 255. The third
cannot be compared by that instrument because poppler renders it one column
wider -- 1356 against 1355 -- which is a page-size rounding difference that
predates this and is not its business.

A four-component JPEG 2000 is charged as three. Nothing in the measured corpus
is that shape, so that figure is a guess where the other two are readings, and
it is the one place this bound is not measured.
@tannevaled
tannevaled merged commit 0889b42 into main Sep 29, 2026
9 checks passed
@tannevaled
tannevaled deleted the decode-cost-by-shape branch September 29, 2026 09:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant