Skip to content

Images returns the inline ones, so 4 147 pictures of one population can be compared at all - #105

Merged
tannevaled merged 3 commits into
mainfrom
inline-images
Oct 6, 2026
Merged

tannevaled merged 3 commits into
mainfrom
inline-images

Conversation

@tannevaled

Copy link
Copy Markdown
Contributor

Fixes #101. Based on #103 — that branch is its first commit, so review this one's second commit alone, or merge #103 first.

The hole

Images returned nothing for a picture written into the content stream with
BI, and said why:

Inline images — the ones written into the content stream with BI — are not
returned: they are named by no resource and are objects of nothing, so there
is nothing for a tool that extracts a file's objects to hand back beside them.

The first half is true and the conclusion was never measured. pdfimages
extracts inline images and lists them, with an object of 0.
Over the
conformance corpus's fr-cerfa — 450 French government forms, page 1 each —
that is 4 147 pictures the reference gets out and nothing of ours was ever
compared with
: 2 614 of them stencils, 2 592 of them 16×16, one document
holding 2 592. A defect in inline-image decoding was invisible to every figure
in that baseline, because a picture nobody returns is a picture nobody compares.

The choices, which are what made this an issue first

  • One entry per BI. An inline image is its draw; there is no object to
    collapse repeats onto; and pdfimages lists one row per draw too.
  • Object 0 — what the file gives it and what pdfimages prints.
  • The name is an ordinal, BI#1 upward across the whole call including
    across forms, because there is no resource name and a reader of a difference
    has to be able to say which picture it was.
  • No mask beside it: an inline dictionary has no abbreviation for /SMask.
  • The budget is charged before the decode, exactly as for an XObject.

Measured over the whole forms corpus

2 268 documents, eighteen populations, through conformance's images harness:

before after
pictures 5 917 11 296
unseen (the reference got it, we did not) 4 243 80
paired by size 0 5 368

The third row is the honest caveat and it is not a side effect. The pairing
falls back to matching on decoded size exactly when an object number is missing,
and until now the one class whose number is missing never reached it — that
fallback has sat in the harness with a share of zero. These 5 368 pictures are
matched by size and order, which is weaker than by object, and
conformance#92 made the share
visible for this reason.

In fr-cerfa, 2 729 of the new pictures come out exact and byte-identical
and none differs.

The differing ones are one fixture, and they are not a defect

Three rows, all in gh-safedocs's
Inline_Image_Abbreviations_InlineAbbreviations.pdf, whose eight dictionaries
each carry the abbreviation and the long name with conflicting values:

/W 20  %% COMMENT OUT THIS LINE TO SEE EFFECT!  /Width 10
/H 10  %% COMMENT OUT THIS LINE TO SEE EFFECT!  /Height 40
/BPC 8 /BitsPerComponent 4

poppler resolves those in favour of the long name — hence its 10x40 at 4 bits
against our 20x10 at 8, the same 600 bytes read two ways — and the misaligned
size pairing that follows accounts for the rest.

reader's InlineImage.Expanded already documents that choice, cites
Table 93 for preferring the abbreviation, and notes that pdf.js and MuPDF
disagree with each other about it. So nothing here is a defect: it is an
ambiguity someone decided on purpose, and it could not be seen until the
pictures reached a comparison at all.

Checks

Mutations: dropping the ordinal gives BI#0; advancing it before the decode
succeeds leaves a gap in the names a reader cannot account for; skipping
afford() lets a page past the budget come back without an error. Each fails a
test.

gofmt, go vet clean; 100.0% of statements.

One note on the suite: TestAThreeComponentJPEGIsCachedOnItsSampleTriple
failed once during this work at a load average of 59 and passed three times
in a row alone. It compares two wall-clock minima taken one after the other
rather than interleaved, which is enough to fail under a spike. Not touched
here — flagged because the next person to see it should not go looking for a
regression.

🤖 Generated with Claude Code

…g ones

qpdf's own fixture says it in its name. form-xobjects-some-resources2.pdf
draws six 15x15 images on page 1 through forms, some of which carry only
part of what they use, and this renderer drew TWO of them.

    images pdfimages -list takes out       6   (objects 10 12 14 16 18 20)
    images render.Images returned          2   (objects 10 12)

It was not only the extraction API. Through go-pdfkit/conformance's
compare against pdftoppm at 72 dpi, on these fixtures alone:

                           before    after
    pixels differing        8.54%    0.03%
    byte for byte          94.06%   99.56%
    mean |diff|            10.936    0.519  levels
    worst square mean      130.27     8.80  levels
    worst pixel               254      149

A worst square mean of 130 is not resampling; four of six stamps were
missing from the page.

Both paths had HALF the rule. drawForm and imagesDrawn each handled an
ABSENT /Resources by falling back to the parent, and a present one
replaced it wholesale -- so a name the form did not list resolved to
nothing and drawXObject returned early. A form XObject's resource
dictionary is not required to be complete, and a name it does not provide
is resolved in the resources in force where the form was painted, which is
what poppler does.

So: renderer.resChain, the dictionaries in force outermost first, pushed by
every stream that brings its own -- run() covers a form, a Type 3 glyph
procedure, a pattern cell, an annotation's appearance and a soft mask; the
extraction path pushes its own because it does not go through run(). The
seven lookups that each read one category out of ONE dictionary --
XObject, Font, ColorSpace, Pattern, Shading, ExtGState, Properties -- go
through named(), which tries the dictionary the caller holds and then
walks outward.

named() takes that dictionary rather than reading it off the chain, and
that is not decoration: the first attempt dropped the parameter, and
TestABDCThatNamesNoLayerHidesNothing failed at once because the tests call
these functions directly. A lookup whose first dictionary is implicit is a
lookup nobody can call.

Checked in both directions. On the whole forms corpus, 2 268 documents in
eighteen populations through conformance's images harness, exactly ONE
population moves: gh-qpdf, pictures 63 -> 73, and the ten recovered land
in (samples) direct as 51 of 51 exact and 51 of 51 byte-identical with
poppler. They are not merely present; they are right. Nothing else in any
population changes a cell.

Mutation-checked: removing the walk fails the three inheritance tests,
and dropping the caller's dictionary fails the one that calls directly.
100.0% of statements.

Fixes #102.
…ation can be compared at all

Images returned nothing for a picture written into the content stream with
BI, and said why: "they are named by no resource and are objects of
nothing, so there is nothing for a tool that extracts a file's objects to
hand back beside them." The first half is true and the conclusion was
never measured. pdfimages EXTRACTS inline images and lists them with an
object of 0.

Measured over the conformance corpus's fr-cerfa, 450 forms, page 1 each:
4 147 pictures the reference got out and nothing of ours was ever compared
with -- 2 614 of them stencils, 2 592 of them 16x16, one document holding
2 592. A defect in inline-image decoding was invisible to every figure in
that baseline, because a picture nobody returns is a picture nobody
compares.

ONE ENTRY PER BI, which is the only unit that can be right: an inline
image IS its draw, there is no object to collapse repeats onto, and
pdfimages lists one row per draw as well. Object 0, which is what the file
gives it and what pdfimages prints. The name is an ordinal, BI#1 upward
across the whole call including across forms, because there is no resource
name and a reader of a difference has to be able to say which picture it
was. No mask is read beside it: an inline dictionary has no abbreviation
for /SMask.

Over the WHOLE forms corpus, 2 268 documents in eighteen populations,
through conformance's images harness:

    pictures         5 917 -> 11 296
    unseen           4 243 -> 80
    paired by size       0 -> 5 368

The third line is the honest caveat and it is not a side effect: the
pairing falls back to size exactly when an object number is missing, and
until now the one class whose number is missing never reached it. That
fallback has been in the harness for weeks with a share of zero. These
5 368 pictures are matched by size and order, which is weaker than by
object, and conformance#92 made that share visible for this reason.

In fr-cerfa 2 729 of the new pictures come out exact and byte-identical
and none differs. The differing ones are three rows of ONE safedocs
fixture, Inline_Image_Abbreviations_InlineAbbreviations.pdf, whose eight
dictionaries each carry the abbreviation AND the long name with
CONFLICTING values -- /W 20 /Width 10 /H 10 /Height 40 /BPC 8
/BitsPerComponent 4, with a comment reading "COMMENT OUT THIS LINE TO SEE
EFFECT". poppler resolves those in favour of the long name, hence its
10x40 at 4 bits against our 20x10 at 8. reader's InlineImage.Expanded
already documents that choice, cites Table 93 for it, and notes that
pdf.js and MuPDF disagree with each other about it. Nothing here is a
defect; it is an ambiguity that was decided on purpose and could not be
seen until the pictures reached a comparison.

Mutation-checked: dropping the ordinal gives "BI#0", advancing it before
the decode succeeds leaves a gap a reader cannot account for, and skipping
afford() lets a page past the budget come back without an error.
100.0% of statements.

Based on #103, which the same counter found. Fixes #101.
@tannevaled
tannevaled merged commit 4e54b0b into main Oct 6, 2026
13 checks passed
@tannevaled
tannevaled deleted the inline-images branch October 6, 2026 14:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Images returns no inline image, so 4 147 pictures of one corpus population are compared against nothing

1 participant