Repository navigation
Images returns the inline ones, so 4 147 pictures of one population can be compared at all - #105
Merged
Merged
Conversation
…g ones
qpdf's own fixture says it in its name. form-xobjects-some-resources2.pdf
draws six 15x15 images on page 1 through forms, some of which carry only
part of what they use, and this renderer drew TWO of them.
images pdfimages -list takes out 6 (objects 10 12 14 16 18 20)
images render.Images returned 2 (objects 10 12)
It was not only the extraction API. Through go-pdfkit/conformance's
compare against pdftoppm at 72 dpi, on these fixtures alone:
before after
pixels differing 8.54% 0.03%
byte for byte 94.06% 99.56%
mean |diff| 10.936 0.519 levels
worst square mean 130.27 8.80 levels
worst pixel 254 149
A worst square mean of 130 is not resampling; four of six stamps were
missing from the page.
Both paths had HALF the rule. drawForm and imagesDrawn each handled an
ABSENT /Resources by falling back to the parent, and a present one
replaced it wholesale -- so a name the form did not list resolved to
nothing and drawXObject returned early. A form XObject's resource
dictionary is not required to be complete, and a name it does not provide
is resolved in the resources in force where the form was painted, which is
what poppler does.
So: renderer.resChain, the dictionaries in force outermost first, pushed by
every stream that brings its own -- run() covers a form, a Type 3 glyph
procedure, a pattern cell, an annotation's appearance and a soft mask; the
extraction path pushes its own because it does not go through run(). The
seven lookups that each read one category out of ONE dictionary --
XObject, Font, ColorSpace, Pattern, Shading, ExtGState, Properties -- go
through named(), which tries the dictionary the caller holds and then
walks outward.
named() takes that dictionary rather than reading it off the chain, and
that is not decoration: the first attempt dropped the parameter, and
TestABDCThatNamesNoLayerHidesNothing failed at once because the tests call
these functions directly. A lookup whose first dictionary is implicit is a
lookup nobody can call.
Checked in both directions. On the whole forms corpus, 2 268 documents in
eighteen populations through conformance's images harness, exactly ONE
population moves: gh-qpdf, pictures 63 -> 73, and the ten recovered land
in (samples) direct as 51 of 51 exact and 51 of 51 byte-identical with
poppler. They are not merely present; they are right. Nothing else in any
population changes a cell.
Mutation-checked: removing the walk fails the three inheritance tests,
and dropping the caller's dictionary fails the one that calls directly.
100.0% of statements.
Fixes #102.
…ation can be compared at all
Images returned nothing for a picture written into the content stream with
BI, and said why: "they are named by no resource and are objects of
nothing, so there is nothing for a tool that extracts a file's objects to
hand back beside them." The first half is true and the conclusion was
never measured. pdfimages EXTRACTS inline images and lists them with an
object of 0.
Measured over the conformance corpus's fr-cerfa, 450 forms, page 1 each:
4 147 pictures the reference got out and nothing of ours was ever compared
with -- 2 614 of them stencils, 2 592 of them 16x16, one document holding
2 592. A defect in inline-image decoding was invisible to every figure in
that baseline, because a picture nobody returns is a picture nobody
compares.
ONE ENTRY PER BI, which is the only unit that can be right: an inline
image IS its draw, there is no object to collapse repeats onto, and
pdfimages lists one row per draw as well. Object 0, which is what the file
gives it and what pdfimages prints. The name is an ordinal, BI#1 upward
across the whole call including across forms, because there is no resource
name and a reader of a difference has to be able to say which picture it
was. No mask is read beside it: an inline dictionary has no abbreviation
for /SMask.
Over the WHOLE forms corpus, 2 268 documents in eighteen populations,
through conformance's images harness:
pictures 5 917 -> 11 296
unseen 4 243 -> 80
paired by size 0 -> 5 368
The third line is the honest caveat and it is not a side effect: the
pairing falls back to size exactly when an object number is missing, and
until now the one class whose number is missing never reached it. That
fallback has been in the harness for weeks with a share of zero. These
5 368 pictures are matched by size and order, which is weaker than by
object, and conformance#92 made that share visible for this reason.
In fr-cerfa 2 729 of the new pictures come out exact and byte-identical
and none differs. The differing ones are three rows of ONE safedocs
fixture, Inline_Image_Abbreviations_InlineAbbreviations.pdf, whose eight
dictionaries each carry the abbreviation AND the long name with
CONFLICTING values -- /W 20 /Width 10 /H 10 /Height 40 /BPC 8
/BitsPerComponent 4, with a comment reading "COMMENT OUT THIS LINE TO SEE
EFFECT". poppler resolves those in favour of the long name, hence its
10x40 at 4 bits against our 20x10 at 8. reader's InlineImage.Expanded
already documents that choice, cites Table 93 for it, and notes that
pdf.js and MuPDF disagree with each other about it. Nothing here is a
defect; it is an ambiguity that was decided on purpose and could not be
seen until the pictures reached a comparison.
Mutation-checked: dropping the ordinal gives "BI#0", advancing it before
the decode succeeds leaves a gap a reader cannot account for, and skipping
afford() lets a page past the budget come back without an error.
100.0% of statements.
Based on #103, which the same counter found. Fixes #101.
This was referenced Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #101. Based on #103 — that branch is its first commit, so review this one's second commit alone, or merge #103 first.
The hole
Imagesreturned nothing for a picture written into the content stream withBI, and said why:The first half is true and the conclusion was never measured.
pdfimagesextracts inline images and lists them, with an object of
0. Over theconformance corpus's
fr-cerfa— 450 French government forms, page 1 each —that is 4 147 pictures the reference gets out and nothing of ours was ever
compared with: 2 614 of them stencils, 2 592 of them 16×16, one document
holding 2 592. A defect in inline-image decoding was invisible to every figure
in that baseline, because a picture nobody returns is a picture nobody compares.
The choices, which are what made this an issue first
BI. An inline image is its draw; there is no object tocollapse repeats onto; and
pdfimageslists one row per draw too.0— what the file gives it and whatpdfimagesprints.BI#1upward across the whole call includingacross forms, because there is no resource name and a reader of a difference
has to be able to say which picture it was.
/SMask.Measured over the whole forms corpus
2 268 documents, eighteen populations, through conformance's
imagesharness:unseen(the reference got it, we did not)The third row is the honest caveat and it is not a side effect. The pairing
falls back to matching on decoded size exactly when an object number is missing,
and until now the one class whose number is missing never reached it — that
fallback has sat in the harness with a share of zero. These 5 368 pictures are
matched by size and order, which is weaker than by object, and
conformance#92 made the share
visible for this reason.
In
fr-cerfa, 2 729 of the new pictures come out exact and byte-identicaland none differs.
The differing ones are one fixture, and they are not a defect
Three rows, all in
gh-safedocs'sInline_Image_Abbreviations_InlineAbbreviations.pdf, whose eight dictionarieseach carry the abbreviation and the long name with conflicting values:
poppler resolves those in favour of the long name — hence its
10x40at 4 bitsagainst our
20x10at 8, the same 600 bytes read two ways — and the misalignedsize pairing that follows accounts for the rest.
reader'sInlineImage.Expandedalready documents that choice, citesTable 93 for preferring the abbreviation, and notes that pdf.js and MuPDF
disagree with each other about it. So nothing here is a defect: it is an
ambiguity someone decided on purpose, and it could not be seen until the
pictures reached a comparison at all.
Checks
Mutations: dropping the ordinal gives
BI#0; advancing it before the decodesucceeds leaves a gap in the names a reader cannot account for; skipping
afford()lets a page past the budget come back without an error. Each fails atest.
gofmt,go vetclean; 100.0% of statements.One note on the suite:
TestAThreeComponentJPEGIsCachedOnItsSampleTriplefailed once during this work at a load average of 59 and passed three times
in a row alone. It compares two wall-clock minima taken one after the other
rather than interleaved, which is enough to fail under a spike. Not touched
here — flagged because the next person to see it should not go looking for a
regression.
🤖 Generated with Claude Code