Skip to content

A form XObject's incomplete /Resources no longer shadows the enclosing ones - #103

Merged
tannevaled merged 1 commit into
mainfrom
form-resources-inherit
Oct 6, 2026
Merged

tannevaled merged 1 commit into
mainfrom
form-resources-inherit

Conversation

@tannevaled

Copy link
Copy Markdown
Contributor

Fixes #102.

The defect

qpdf's own fixture says it in its name. form-xobjects-some-resources2.pdf
draws six 15×15 images on page 1 through form XObjects, some of which carry
only part of what they use, and this renderer drew two.

page 1
images pdfimages -list takes out 6 (objects 10, 12, 14, 16, 18, 20)
images render.Images returned 2 (objects 10, 12)

It was not only the extraction API. Through go-pdfkit/conformance's compare
against pdftoppm at 72 dpi, on these fixtures alone:

before after
pixels differing 8.54% 0.03%
byte for byte 94.06% 99.56%
mean abs difference 10.936 levels 0.519
worst square mean 130.27 levels 8.80
worst pixel 254 149

A worst square mean of 130 is not resampling. Four of the six stamps were
missing from the page we drew.

The cause: both paths had half the rule

drawForm and imagesDrawn each handled an absent /Resources by falling
back to the parent:

resources, ok := r.doc.GetDict(stream.Dict, "Resources")
if !ok {
	resources = parent
}

and a present one replaced the parent wholesale — so a name the form did not
list resolved to nothing and drawXObject returned early. The comment beside
one of them, "A form with no resources of its own draws against the ones in
force where it was drawn"
, is exactly right and exactly half.

A form XObject's resource dictionary is not required to be complete, and a name
it does not provide is resolved in the resources in force where the form was
painted. poppler does that — hence six against two.

The change

renderer.resChain holds the dictionaries in force, outermost first. Every
stream that brings its own resources pushes them: run() covers a form, a
Type 3 glyph procedure, a pattern cell, an annotation's appearance and a soft
mask, and the extraction path pushes its own because it does not go through
run().

The seven lookups that each read one category out of one dictionary —
XObject, Font, ColorSpace, Pattern, Shading, ExtGState,
Properties — now go through named(), which tries the dictionary the caller
holds and then walks outward. maxFormDepth already bounds the depth.

named() takes that dictionary rather than reading it off the chain, and
that is not decoration. The first attempt dropped the parameter, and
TestABDCThatNamesNoLayerHidesNothing failed immediately, because the tests
call these functions directly. A lookup whose first dictionary is implicit is a
lookup nobody can call.

Checked in both directions

Nothing else moves. The whole forms corpus — 2 268 documents in eighteen
populations, through conformance's images harness — has exactly one
population change: gh-qpdf, pictures 63 → 73. The ten recovered land in
(samples) direct as 51 of 51 exact and 51 of 51 byte-identical with
poppler. They are not merely present; they are right. No other population
changes a cell.

Mutations. Removing the outward walk fails the three inheritance tests;
dropping the caller's own dictionary fails the one that calls directly.

gofmt, go vet clean; 100.0% of statements, which is what the gate asks.

How it was found

By a counter added in
conformance#92 for the
pictures the reference extracts and we do not. Of gh-qpdf's thirteen such
rows, three are inline images (which Images does not return by design,
#101) and ten were this.

🤖 Generated with Claude Code

…g ones

qpdf's own fixture says it in its name. form-xobjects-some-resources2.pdf
draws six 15x15 images on page 1 through forms, some of which carry only
part of what they use, and this renderer drew TWO of them.

    images pdfimages -list takes out       6   (objects 10 12 14 16 18 20)
    images render.Images returned          2   (objects 10 12)

It was not only the extraction API. Through go-pdfkit/conformance's
compare against pdftoppm at 72 dpi, on these fixtures alone:

                           before    after
    pixels differing        8.54%    0.03%
    byte for byte          94.06%   99.56%
    mean |diff|            10.936    0.519  levels
    worst square mean      130.27     8.80  levels
    worst pixel               254      149

A worst square mean of 130 is not resampling; four of six stamps were
missing from the page.

Both paths had HALF the rule. drawForm and imagesDrawn each handled an
ABSENT /Resources by falling back to the parent, and a present one
replaced it wholesale -- so a name the form did not list resolved to
nothing and drawXObject returned early. A form XObject's resource
dictionary is not required to be complete, and a name it does not provide
is resolved in the resources in force where the form was painted, which is
what poppler does.

So: renderer.resChain, the dictionaries in force outermost first, pushed by
every stream that brings its own -- run() covers a form, a Type 3 glyph
procedure, a pattern cell, an annotation's appearance and a soft mask; the
extraction path pushes its own because it does not go through run(). The
seven lookups that each read one category out of ONE dictionary --
XObject, Font, ColorSpace, Pattern, Shading, ExtGState, Properties -- go
through named(), which tries the dictionary the caller holds and then
walks outward.

named() takes that dictionary rather than reading it off the chain, and
that is not decoration: the first attempt dropped the parameter, and
TestABDCThatNamesNoLayerHidesNothing failed at once because the tests call
these functions directly. A lookup whose first dictionary is implicit is a
lookup nobody can call.

Checked in both directions. On the whole forms corpus, 2 268 documents in
eighteen populations through conformance's images harness, exactly ONE
population moves: gh-qpdf, pictures 63 -> 73, and the ten recovered land
in (samples) direct as 51 of 51 exact and 51 of 51 byte-identical with
poppler. They are not merely present; they are right. Nothing else in any
population changes a cell.

Mutation-checked: removing the walk fails the three inheritance tests,
and dropping the caller's dictionary fails the one that calls directly.
100.0% of statements.

Fixes #102.
tannevaled added a commit to go-pdfkit/conformance that referenced this pull request Oct 4, 2026
… fixed

gh-qpdf's 13 unseen rows are 3 inline images, 0 annotation pictures, and
TEN that were a defect. The fixture names itself:
form-xobjects-some-resources2.pdf.

Both of render's paths had half a rule. drawForm and imagesDrawn each fell
back to the parent when a form XObject's /Resources was ABSENT, and a
PRESENT one replaced it wholesale -- so a name the form did not list
resolved to nothing and the picture was silently not drawn. A form's
resource dictionary is not required to be complete.

Six images against our two on that page, and it reached the page and not
only the extraction: through this repository's own compare against
pdftoppm, 8.54% of pixels away at a worst square mean of 130 levels.
After the fix, 0.03% and 8.80.

go-pdfkit/render#103 fixes it, filed as #102, checked in both directions
over the whole forms corpus: exactly one population moves, gh-qpdf's
pictures 63 -> 73, and the ten recovered land in (samples) direct as 51 of
51 exact and 51 of 51 byte-identical with poppler.

Which is what the counter was for. Unseen was added as a field documenting
a requirement, its first number was wrong twice, and what survived the two
corrections found a rendering defect every fidelity figure in this file
had been blind to -- because a picture nobody returns is a picture nobody
compares.
tannevaled added a commit to go-pdfkit/conformance that referenced this pull request Oct 6, 2026
* images: count the pictures paired by size, and say so

PairedBy's own comment states the requirement:

	A run whose size share is large is a run whose numbers are worth
	less, and that has to be visible.

It was not visible anywhere. The pairing already records how each picture
was matched -- by object number, or by falling back to size -- but Tally
threw that away, so neither the report nor the JSON record carried it. A
reader of either had no way to tell a run where every picture was matched
by its object number from one where the matching was guesswork.

Counts gains SizePaired, Tally increments it on PairedBySize only, the
filter line ends with "N paired by size", and FilterCounts carries it into
the record (omitempty, so existing records that have no such number stay
as they are).

The test asserts both directions. A counter that increments on every
picture reads exactly like one that works, until the run that is all
object-paired reports every picture as suspect -- so the JPXDecode row,
which has no size-paired picture in it, must come out zero. Checked by
mutation: counting unconditionally, dropping the report column, and
dropping the record field each make it fail.

* baseline: answer the pairing-audit entry, and date the records

Two entries in "What is not measured, and why" are addressed, one of them
by closing it and one by naming what blocks it.

The pairing entry asked whether a per-picture pairing audit had been run
since conformance#13 closed. It is now counted in every run rather than
audited once, and the first reading is ZERO over 1 606 pictures -- 330
across the five pdfscans populations, 1 276 across all eighteen of
pdfforms, 53 filter rows, none paired by size.

A zero needs a control, because a counter that never fires reads exactly
like a fallback that never fires. Ablating the object test -- object
forced to 0 inside match, the only input its first loop reads -- pairs
EVERY picture of the same sample by size. The counter is live and the
zero belongs to the corpus.

§30 is the other half, and it is the uncomfortable one. §29 says this
file must be read backwards because a section is true of the run that
wrote it; that applies to the records too, and there a strikethrough will
not do. All twenty-three are dated 2026-09-24 at render v0.35.0 and
jpeg2000 v0.1.0, while go.mod holds v0.60.0/v0.9.1 and §29 measures
v0.67.0/v0.13.2 -- four subjects in one document, with the
machine-readable half the oldest of them. Nineteen jpeg2000 releases have
landed since the version the records name, including the decoder of §25
and the three blank pages of §26.

That is what blocks the aggregate bound. The twenty-one buckets carrying
terms do invite one -- mse.worst 0.654, |mean.worst| 0.579, peak.worst 4
everywhere but gh-openpdf's converted bucket at 20 -- but a bound fitted
to them is fitted to v0.35.0, and a gate borrowed from code that no
longer exists reddens on the first honest change and is then deleted. The
next step is the records, not the bound.

No gate is added. A test asserting the records name go.mod's version
would redden every dependency bump; the move here is §28's and
conformance#92's -- make the number visible and let the reader see that a
record is six days and thirty-two releases behind what it is cited for.

* fix(deps): update dependencies (#55)

* baseline: say how the twenty-three records are re-taken

§30 names the records as the next step, and nothing in the repository said
how to take them. The top-level README shows one invocation of
"images -json"; neither file said how the SET is taken, what it is named,
or what it costs.

So: one invocation per population, which is where the filenames in
baseline/ come from, and omitting -only gives one record carrying every
population of that corpus -- not the shape baseline/ holds.

With the cost, measured one at a time on 2026-10-01: fr-cerfa (450 forms)
about two minutes, ia-biodiversity (250 scans) about fifty. And with the
reason to take them one at a time, which is §22's: the judge is held to a
wall-clock bound, that bound is a property of the machine rather than of
the corpus, and two runs beside each other can turn a document into a
hung entry neither would produce alone.

* images: ask the judge before calling a page refusal a defect, and §31

§28 is titled "The one cell in this document that said «defect», and
closing it", and it is not closed. Re-taken 2026-10-03 at the versions §28
names -- render v0.67.0, go-images/jpeg2000 v0.13.2, one population at a
time, with a v0.60.0 run of the same population as control -- four of its
five cells reproduce and the fifth does not:

                   §28 before  §28 after  re-measured
    pictures              754        755          755  ok
    direct                705        706          706  ok
    compared              650        651          651  ok
    exact                 650        651          651  ok
    refused                 1          0            1  NOT

bulletinno38tasm.pdf is still refused, byte-identically at v0.60.0 and
v0.67.0, so the seven releases §28 credits changed nothing about it. It is
not §26's ceiling either. Page 1 names THREE pictures of 9449 by 13701 --
two JPEG 2000 greys and a JBIG2 soft mask, all three of which pdfimages
takes out -- each 129 460 749 pixels. Two spend 258 921 498 of the
268 435 456 budget and leave 9 513 958, which is the number in the
message; the third cannot be afforded. The charge is honest for the path:
decodeBase says r.bounded is true exactly on the Images path, which hands
pictures out through raster.Image and so spends four bytes a pixel. Three
pictures of 129.5 Mpx is 1.447 GiB at four bytes against a 1.000 GiB page
budget, and 0.362 GiB as the single-component greys they are. Tracked as
go-pdfkit/render#100.

Looking for the document turned up a defect in the instrument instead.
Ours is documented as "ours would not draw it AND THE JUDGE WOULD", and it
is set at two places: reader.Open failing, which goes through blame(), and
render.Images failing on a page, which asked nobody. A page refusal went
into the defect column whatever poppler did with the same page. Here it
happens to be right, but by hand, not by instrument.

judgePage now asks, of the same tool on the same page, the way the branch
below it already does: pictures back is Ours, a tool that does not finish
is Hung, and a refusal -- or a run that takes nothing out, which
judgeShots reports the same way -- is Neither. The vocabulary now says
that the last two are not told apart instead of implying a distinction the
code cannot make.

It changed an answer at once: TestAPageThatIsNotThereSaysSo asks for page
nine of a one-page document and asserted Ours, so the instrument called a
page that does not exist a defect of ours. Asked, poppler cannot give page
nine either. Checked by mutation -- never asking the judge, and swapping
Hung for Neither, each make the tests fail. 100.0% of statements in every
package.

And §30 is why this sat for nine days: §28's figures were taken in a tree
with the modules upgraded by hand and never written to baseline/, so the
only copy was prose, which cannot be re-run. The committed record said
refused: 1 the whole time. The repository contained its own refutation.

* baseline: check the refusal fix against the corpus, not only its tests

ia-biodiversity re-taken a third time with the asking in place: refused is
still 1, unopenable and declined still 0, every filter row identical to the
run before it. A classification that quietly moved a real defect into
Neither would be indistinguishable from a fix when read off the test suite
alone; asked of the document it was written for, it leaves it where it
belongs, because pdfimages does hand three pictures back.

Says explicitly that none of the three runs is committed to baseline/:
they were taken at render v0.67.0 while go.mod holds v0.60.0, which is
what §30 is about.

* images: the page draws, only the extraction refuses -- and the census is one

A correction to §31, found the same day, by asking the renderer the
question the column's own words name. afford() and r.budget are reached
from images.go ALONE, and bounded:true is set at exactly one place in the
whole module -- inside Images. So:

    Images  REFUSED: ... with 9513958 of the 268435456 pixels left
    Page    ok: 411x596, 23659 dark pixels, 216735 light, 4562 between

Not a blank page with a clean exit either. Through this repository's own
compare, against pdftoppm at 72 dpi: 7.30% of pixels differ, mean 14.875
levels, and we take 795ms against poppler's 2.992s. A 9449x13701 scan
reduced to 411x596 is a 23x downscale, which is the resampling
disagreement §1 describes and the reason this baseline measures codecs
through images rather than through a rasteriser.

So Ours, whose comment said "would not open the document or draw the
page", is here NEITHER: ours draws the page. The comment now says what the
column measures -- no picture came out of the extraction -- and names the
instance, because a vocabulary that overstates is how §28 came to be
written. render#100 drops from a gap against poppler to a limit of one
API, and says so.

Then the census, so that the limit is sized rather than guessed. Every
document of both corpora, page 1, our side only, never poppler:

    3281 documents
      67 our side produced nothing for
      66 of them refused at reader.Open
       1 of them refused by the page budget

One document in 3 281. And the sixty-six are §27's sixty-six, recovered by
an instrument that asks neither build nor poppler: 39 unsupported security
handler, 23 wrong password, 2 no indirect objects, 1 /Brotli, 1 no
startxref, in the same populations (ia-americana 30, gh-openpdf 14,
gh-pdfbox 8, ia-texts 7, gh-safedocs 5, gh-pypdf 1, ia-medical 1). Two
instruments that share no dependency agreeing on a census is a control
this file has been short of -- three agreed once before and all three read
the same layer.

* images: count the pictures paired by size, and say so

PairedBy's own comment states the requirement:

	A run whose size share is large is a run whose numbers are worth
	less, and that has to be visible.

It was not visible anywhere. The pairing already records how each picture
was matched -- by object number, or by falling back to size -- but Tally
threw that away, so neither the report nor the JSON record carried it. A
reader of either had no way to tell a run where every picture was matched
by its object number from one where the matching was guesswork.

Counts gains SizePaired, Tally increments it on PairedBySize only, the
filter line ends with "N paired by size", and FilterCounts carries it into
the record (omitempty, so existing records that have no such number stay
as they are).

The test asserts both directions. A counter that increments on every
picture reads exactly like one that works, until the run that is all
object-paired reports every picture as suspect -- so the JPXDecode row,
which has no size-paired picture in it, must come out zero. Checked by
mutation: counting unconditionally, dropping the report column, and
dropping the record field each make it fail.

* baseline: answer the pairing-audit entry, and date the records

Two entries in "What is not measured, and why" are addressed, one of them
by closing it and one by naming what blocks it.

The pairing entry asked whether a per-picture pairing audit had been run
since conformance#13 closed. It is now counted in every run rather than
audited once, and the first reading is ZERO over 1 606 pictures -- 330
across the five pdfscans populations, 1 276 across all eighteen of
pdfforms, 53 filter rows, none paired by size.

A zero needs a control, because a counter that never fires reads exactly
like a fallback that never fires. Ablating the object test -- object
forced to 0 inside match, the only input its first loop reads -- pairs
EVERY picture of the same sample by size. The counter is live and the
zero belongs to the corpus.

§30 is the other half, and it is the uncomfortable one. §29 says this
file must be read backwards because a section is true of the run that
wrote it; that applies to the records too, and there a strikethrough will
not do. All twenty-three are dated 2026-09-24 at render v0.35.0 and
jpeg2000 v0.1.0, while go.mod holds v0.60.0/v0.9.1 and §29 measures
v0.67.0/v0.13.2 -- four subjects in one document, with the
machine-readable half the oldest of them. Nineteen jpeg2000 releases have
landed since the version the records name, including the decoder of §25
and the three blank pages of §26.

That is what blocks the aggregate bound. The twenty-one buckets carrying
terms do invite one -- mse.worst 0.654, |mean.worst| 0.579, peak.worst 4
everywhere but gh-openpdf's converted bucket at 20 -- but a bound fitted
to them is fitted to v0.35.0, and a gate borrowed from code that no
longer exists reddens on the first honest change and is then deleted. The
next step is the records, not the bound.

No gate is added. A test asserting the records name go.mod's version
would redden every dependency bump; the move here is §28's and
conformance#92's -- make the number visible and let the reader see that a
record is six days and thirty-two releases behind what it is cited for.

* baseline: say how the twenty-three records are re-taken

§30 names the records as the next step, and nothing in the repository said
how to take them. The top-level README shows one invocation of
"images -json"; neither file said how the SET is taken, what it is named,
or what it costs.

So: one invocation per population, which is where the filenames in
baseline/ come from, and omitting -only gives one record carrying every
population of that corpus -- not the shape baseline/ holds.

With the cost, measured one at a time on 2026-10-01: fr-cerfa (450 forms)
about two minutes, ia-biodiversity (250 scans) about fifty. And with the
reason to take them one at a time, which is §22's: the judge is held to a
wall-clock bound, that bound is a property of the machine rather than of
the corpus, and two runs beside each other can turn a document into a
hung entry neither would produce alone.

* images: ask the judge before calling a page refusal a defect, and §31

§28 is titled "The one cell in this document that said «defect», and
closing it", and it is not closed. Re-taken 2026-10-03 at the versions §28
names -- render v0.67.0, go-images/jpeg2000 v0.13.2, one population at a
time, with a v0.60.0 run of the same population as control -- four of its
five cells reproduce and the fifth does not:

                   §28 before  §28 after  re-measured
    pictures              754        755          755  ok
    direct                705        706          706  ok
    compared              650        651          651  ok
    exact                 650        651          651  ok
    refused                 1          0            1  NOT

bulletinno38tasm.pdf is still refused, byte-identically at v0.60.0 and
v0.67.0, so the seven releases §28 credits changed nothing about it. It is
not §26's ceiling either. Page 1 names THREE pictures of 9449 by 13701 --
two JPEG 2000 greys and a JBIG2 soft mask, all three of which pdfimages
takes out -- each 129 460 749 pixels. Two spend 258 921 498 of the
268 435 456 budget and leave 9 513 958, which is the number in the
message; the third cannot be afforded. The charge is honest for the path:
decodeBase says r.bounded is true exactly on the Images path, which hands
pictures out through raster.Image and so spends four bytes a pixel. Three
pictures of 129.5 Mpx is 1.447 GiB at four bytes against a 1.000 GiB page
budget, and 0.362 GiB as the single-component greys they are. Tracked as
go-pdfkit/render#100.

Looking for the document turned up a defect in the instrument instead.
Ours is documented as "ours would not draw it AND THE JUDGE WOULD", and it
is set at two places: reader.Open failing, which goes through blame(), and
render.Images failing on a page, which asked nobody. A page refusal went
into the defect column whatever poppler did with the same page. Here it
happens to be right, but by hand, not by instrument.

judgePage now asks, of the same tool on the same page, the way the branch
below it already does: pictures back is Ours, a tool that does not finish
is Hung, and a refusal -- or a run that takes nothing out, which
judgeShots reports the same way -- is Neither. The vocabulary now says
that the last two are not told apart instead of implying a distinction the
code cannot make.

It changed an answer at once: TestAPageThatIsNotThereSaysSo asks for page
nine of a one-page document and asserted Ours, so the instrument called a
page that does not exist a defect of ours. Asked, poppler cannot give page
nine either. Checked by mutation -- never asking the judge, and swapping
Hung for Neither, each make the tests fail. 100.0% of statements in every
package.

And §30 is why this sat for nine days: §28's figures were taken in a tree
with the modules upgraded by hand and never written to baseline/, so the
only copy was prose, which cannot be re-run. The committed record said
refused: 1 the whole time. The repository contained its own refutation.

* baseline: check the refusal fix against the corpus, not only its tests

ia-biodiversity re-taken a third time with the asking in place: refused is
still 1, unopenable and declined still 0, every filter row identical to the
run before it. A classification that quietly moved a real defect into
Neither would be indistinguishable from a fix when read off the test suite
alone; asked of the document it was written for, it leaves it where it
belongs, because pdfimages does hand three pictures back.

Says explicitly that none of the three runs is committed to baseline/:
they were taken at render v0.67.0 while go.mod holds v0.60.0, which is
what §30 is about.

* images: the page draws, only the extraction refuses -- and the census is one

A correction to §31, found the same day, by asking the renderer the
question the column's own words name. afford() and r.budget are reached
from images.go ALONE, and bounded:true is set at exactly one place in the
whole module -- inside Images. So:

    Images  REFUSED: ... with 9513958 of the 268435456 pixels left
    Page    ok: 411x596, 23659 dark pixels, 216735 light, 4562 between

Not a blank page with a clean exit either. Through this repository's own
compare, against pdftoppm at 72 dpi: 7.30% of pixels differ, mean 14.875
levels, and we take 795ms against poppler's 2.992s. A 9449x13701 scan
reduced to 411x596 is a 23x downscale, which is the resampling
disagreement §1 describes and the reason this baseline measures codecs
through images rather than through a rasteriser.

So Ours, whose comment said "would not open the document or draw the
page", is here NEITHER: ours draws the page. The comment now says what the
column measures -- no picture came out of the extraction -- and names the
instance, because a vocabulary that overstates is how §28 came to be
written. render#100 drops from a gap against poppler to a limit of one
API, and says so.

Then the census, so that the limit is sized rather than guessed. Every
document of both corpora, page 1, our side only, never poppler:

    3281 documents
      67 our side produced nothing for
      66 of them refused at reader.Open
       1 of them refused by the page budget

One document in 3 281. And the sixty-six are §27's sixty-six, recovered by
an instrument that asks neither build nor poppler: 39 unsupported security
handler, 23 wrong password, 2 no indirect objects, 1 /Brotli, 1 no
startxref, in the same populations (ia-americana 30, gh-openpdf 14,
gh-pdfbox 8, ia-texts 7, gh-safedocs 5, gh-pypdf 1, ia-medical 1). Two
instruments that share no dependency agreeing on a census is a control
this file has been short of -- three agreed once before and all three read
the same layer.

* images: count the pictures the judge got out and we did not

The pairing walks OUR pictures and appends one result for each, so a row
of the judge's that nothing claimed left no trace in any count. The one
direction this harness could not see was the direction that matters most:
a picture the field gets out of a file and we do not.

Its own test suite had already set the case up and recorded the blind spot
as the expected answer. TestAPictureTheOtherSideDoesNotHave stands in a
judge with a 9x9 against a page holding a 2x1: ours is unmatched AND
theirs is, the first has been reported since the day it was written, the
second was dropped, and the test asserted len == 1.

So: Missing.Unseen, one result per unclaimed row with Name left empty so
Tally skips it -- there is no filter of OURS to group it under, because
there is no picture of ours -- Summary.Unseen (omitempty), and a report
line, for the reason the hangs have one: a report that leaves it out
cannot be told from one that has none.

It is known to be non-empty by construction for at least one class.
render.Images does not return inline images, and says why: they are the
objects of nothing. pdfimages lists them with an object of 0. Whether
this corpus holds any is what the count is for.

Two mutations, and the SECOND needed a test written for it: dropping the
loop fails TestAPictureTheOtherSideDoesNotHave, but counting the CLAIMED
rows as well passed everything, because no test had a judge with one row
paired and one not. TestOnlyTheJudgePicturesNothingClaimedAreCounted is
that case, and it is the same hole SizePaired was guarded against -- a
counter that fires on everything reads exactly like one that works.

100.0% of statements in every package. §31's closing note is amended:
#55 landed on 2026-10-03, so a record taken at v0.67.0 now agrees with
the tree that reads it.

* images: tell a repeated draw apart from a picture we failed to get

The counter added in the commit before this one said 7 614 over the
eighteen forms populations, 7 480 of them in fr-cerfa, and that number was
wrong in the way a number is wrong when one document carries 304 of the
316 in a sixty-document sample.

The document is one this file already names, cerfa_10074.pdf, the page of
211 uniform 2x2 swatches under /SMasks. Asked directly:

    rows pdfimages -list prints              729
    distinct objects among them              214
    rows for object 147 alone                 62
    pictures render.Images returns           425
    729 - 425                                304   = the unseen count

pdfimages lists one row PER DRAW. The judgePage comment said a repeated
draw came back "against pdfimages's one"; that was never measured and is
wrong, and it is corrected where it stands.

Ours is one entry per object on purpose -- decoding the same stream twice
says nothing about a codec -- so the asymmetry is expected and harmless to
every fidelity figure here. Counting it as pictures we failed to get is
not. Missing.Repeated is now its own count, in the record and on a report
line printed beside Unseen so the small number is read as small.

What is left in Unseen is a scope difference and it is real: six rows in
those sixty documents, one of which is cerfa_10011.pdf's 106x56 DCTDecode
reached as /Annots -> /Widget /Btn -> /MK /I -> Form /FRM -> /Im0, a push
button's ICON. Images walks the content stream and cannot see it;
pdfimages renders the page with its annotations. A tool written to count
exactly those agreed DOCUMENT BY DOCUMENT on 56 of 60, and the four it
missed are the repeat-draw case. render.Page draws annotations, so this is
not a decoder gap -- it is the extraction API's scope, which is §31's
finding in a different place.

Both classifications are mutation-checked in both directions: never
repeated gives "repeated 0, want 1" and "unseen 2, want 1"; always
repeated breaks the two tests that need an unseen. 100.0% of statements.

§32 has the measurement. Whether Images should walk /AP and /MK is named
there and not decided: it changes every count in the file.

* images: name the real cause of Unseen -- inline images, not annotations

§32 offered annotations as the cause and that was fitted to a sample where
the thing being counted was almost always nought. Fifty-six of sixty
documents "agreeing" was fifty-four pairs of noughts.

Over the whole fr-cerfa population instead -- 450 forms, page 1 each --
there are 4 153 unseen rows after the repeat split, and a program that
counts the pictures only an annotation reaches finds 25 of them in all
450. Grouped by what the judge's own listing says:

    4 147  object 0
    2 614  stencil type
    1 538  image type
    2 592  16x16
       13  documents of 450 carrying any
    2 592  cerfa_12481.pdf alone

An object of nought is what pdfimages prints for an INLINE image, and
render.Images says plainly that it does not return those. Read at the
byte level, the first three of that document's eight content streams hold
751 BI...ID images, the first being

    /IM true /W 16 /H 16 /BPC 1 /D[1 0] /F /CCF /DP<</K -1 /Columns 16>>

a 16x16 CCITT-fax image mask, which is exactly the stencil 16x16 row
poppler lists.

So a class of 4 147 pictures in one population is drawn by render.Page
and compared against the reference by nothing: a defect in inline-image
decoding would be invisible to every figure in this baseline. Filed as
go-pdfkit/render#101.

It also explains a nought from two sections earlier. SizePaired is zero in
every population because the pairing falls back to size exactly when an
object number is missing, and the one class whose object number is missing
never reaches the pairing. The fallback was written for a case the API
cannot deliver.

The annotation case stays, measured and small, with cerfa_10011.pdf's
106x56 DCTDecode reached through /MK /I. Neither case is a decoder gap --
render.Page draws both -- and §32 keeps the method note, because this is
the third total in one day dominated by a single document.

* baseline: the third cause of Unseen was a rendering defect, and it is fixed

gh-qpdf's 13 unseen rows are 3 inline images, 0 annotation pictures, and
TEN that were a defect. The fixture names itself:
form-xobjects-some-resources2.pdf.

Both of render's paths had half a rule. drawForm and imagesDrawn each fell
back to the parent when a form XObject's /Resources was ABSENT, and a
PRESENT one replaced it wholesale -- so a name the form did not list
resolved to nothing and the picture was silently not drawn. A form's
resource dictionary is not required to be complete.

Six images against our two on that page, and it reached the page and not
only the extraction: through this repository's own compare against
pdftoppm, 8.54% of pixels away at a worst square mean of 130 levels.
After the fix, 0.03% and 8.80.

go-pdfkit/render#103 fixes it, filed as #102, checked in both directions
over the whole forms corpus: exactly one population moves, gh-qpdf's
pictures 63 -> 73, and the ten recovered land in (samples) direct as 51 of
51 exact and 51 of 51 byte-identical with poppler.

Which is what the counter was for. Unseen was added as a field documenting
a requirement, its first number was wrong twice, and what survived the two
corrections found a rendering defect every fidelity figure in this file
had been blind to -- because a picture nobody returns is a picture nobody
compares.

* baseline: the annotation case dominates another population, so neither ranks

§32 demoted annotations to "real and small: 25 in 450", which is true of
fr-cerfa and false as a general claim -- the same mistake as the one the
section is about, in miniature.

    population   documents  unseen  object 0  annotation-reached
    fr-cerfa           450    4153      4147                  25
    us-uscis            88      73         0                  74

In us-uscis it is ONE picture per document in 73 of 88, and not one is
inline: g-1041.pdf's is a 1125x75 one-bit strip, which is a barcode. On a
population of government forms the barcode or logo that identifies the
form is the one picture per page nothing ever compares.

Filed as go-pdfkit/render#104. Neither cause ranks above the other -- each
dominates a different population -- and a single ordering of them would
have been wrong whichever way round it was written.

* baseline: re-take the twenty-three records, and §33 on what moved

§30 said the records were thirty-two render releases and nineteen
jpeg2000 releases behind the tree that reads them, and that the next step
was the records and not a bound. Re-taken 2026-10-04 at render v0.67.0 and
go-images/jpeg2000 v0.13.2, one population at a time for §22's reason,
against the same pdfimages 26.04.0.

TWO CELLS MOVED in thirty-two releases:

    ia-biodiversity  JPXDecode direct  pictures  496 -> 497
                                       exact     496 -> 497
    gh-pdfbox        JPXDecode conv.   mean    -0.0020404 -> -0.0020666

and nothing else, in twenty-three populations: every other bucket is
identical cell for cell, to every digit recorded. The first is §26's page,
and the picture it gained comes out exact -- the one claim of §28's table
that §31 did not refute, now standing on a record instead of on prose. The
second is a signed mean moving by 0.000026 of a level on one picture.

So §30 was right about the fact and wrong about the risk it implied: a
record can be thirty-two releases stale in its metadata and current in its
numbers, and re-taking it is the only way to know which. What it buys is
that the prose and the records now describe the same code, which is what
§31 found had failed.

The fresh records also carry what the old ones could not: sizePaired 0,
unseen 4 277, repeated 3 371 over 3 280 documents and 8 067 pictures.

And §33 reconciles 66 with 63. §27 counts 66 documents neither draws; the
records say unopenable 63. Four scans files are on disk and in no manifest
row, and three of them are refusals -- two in ia-americana, one in
ia-medical. images and compare walk the MANIFEST; the sweeps behind §27
and §31 walk the directory. Neither is wrong and no count is transferable
between them without saying which it walked, which is why §24 reads 3 208,
§27 reads 3 281 and §33 reads 3 280.

One row has a fix waiting: render#103 takes gh-qpdf from 63 pictures to
73. These records are at the go.mod as it stands, so that row is expected
to move when it lands, and §33 says so before it does.

* baseline: the Conditions are the fresh run's, and §30 was wrong about the gate

The re-take needed the Conditions table to follow it, and the repository
said so itself: TestTheBaselineReadmeDescribesTheRecordsBesideIt failed on
the first run with three rows stale -- render v0.35.0, go-gfx v0.31.0 and
jpeg2000 v0.1.0 against records taken at v0.67.0, v0.34.0 and v0.13.2.

One of the three is worth naming. The jpeg2000 mention the guard reads is
a PROSE sentence, and my first correction put the new version on the next
line of the rewrapped paragraph -- so the version was in the file and not
on the row, which is exactly the case the guard's own comment says it was
narrowed to catch. Reworded so the version sits on the matched line.

And §30 is corrected. It argued against adding a gate, and a gate was
already there -- a better one. It asserts the README's rows match the
RECORDS, not go.mod, so a dependency bump alone cannot redden it; what
reddens it is a re-take whose README was not amended, which is what
happened here. A guard on agreement between two things this repository
holds costs nothing to keep green; a guard on agreement with the outside
world is a tax on every upgrade.

The "these figures still hold twelve releases later" paragraph is struck:
they held for thirty-two, and §33 says which two cells moved.

* README: correct what it says about the record, and document three counts

Two paragraphs had outlived their subject, which is the hazard §29 is
about, in the file a reader opens first.

"refused is 4 across 3280 documents, all in ia-biodiversity, and the same
four documents as before" -- it is 1, in the records re-taken 2026-10-04,
and the one document is bulletinno38tasm.pdf, whose page render.Page
DRAWS: it is the extraction API that refuses it under a per-page budget
the Images path spends four bytes a pixel of. §31 has the measurement and
render#100 holds the question.

"The instrument folds cannot-read and declined-to-decode into one count
and should not" -- it no longer asserts the other half either. Ours is
documented as "ours produced nothing and the judge did", and the page path
set it WITHOUT ASKING poppler. It asks now, and the first thing that
changed was a test demanding Ours for page nine of a one-page document.

"The bound is unexercised by images on this corpus" -- §22 struck that
months ago: the bound fired three times and it is a property of the
machine, not of the corpus. The narrower true thing is that none of these
23 records has a hung entry.

And the three counts the records now carry are described where the record
is described: sizePaired, unseen and repeated, with what each is FOR and
what it reads today -- 0, 4 277 and 3 371. All omitempty.

* baseline: §34, the page compared against poppler over the whole corpus

Everything above §34 measures extracted pictures -- images against
pdfimages, 23 records, every codec drilled to one picture -- and §24
measures time. The page itself was never recorded: baseline/pages/ held
one population, as a control of one binary against itself.

compare against pdftoppm at 72 dpi with -cropbox, one population at a
time, 2026-10-04, at the go.mod of the records beside it. All 23 raw
outputs are in baseline/pages/, unedited.

    3214 pages compared, 66 not
    3001 of 3214 differ on under 1% of their pixels  (93.4%)
    the MEDIAN is at or under 0.35% in every population, 0.000% in five

The 66 not compared are 63 refused, 2 of different pixel sizes and 1 hung,
and the 63 are the records' unopenable 63 by an instrument that renders
instead of extracting. That this run's total is ALSO 66 is a coincidence
of three causes rather than §27's 66, and §34 says so, because the number
invites the mistake.

The one hung makes concrete the paragraph the top-level README carries:
images reports hung 0 in all 23 records because that document draws no
picture on its first page and the extraction harness never asks poppler
about it. compare meets it.

And §34 says not to read a low byte-identical share as a defect.
ia-medical is 15.69% byte-identical against gh-verapdf's 99.60% and is not
twenty times worse -- its median difference is 0.000% of pixels. A
9000-pixel scan reduced to 600 is a 15x to 23x downscale, where any
resampling difference puts a level into most pixels and byte-identity
collapses while the picture does not.

The worst page of the corpus is ia-americana's 19.23% and is not diagnosed
here. Sixth worst is gh-qpdf's 8.54%, which render#103 takes to 0.03%:
these records are at the go.mod as it stands, and §33 says the same of the
row it moves.

Two things these files are not, both named: the timing columns were
measured while the load moved between 7 and 59 and are not to be read --
only the fidelity columns are deterministic, which is the sole reason the
run is usable -- and they are text, because compare has no -json. That is
a real gap against this repository's own "a number that is not written
down cannot be regressed against", and §34 records why it was not closed
here: the record type and its helpers live in cmd/images, package main, so
compare cannot import them, and moving them touches files an unmerged pull
request is changing.

* corpus: check that a corpus is what its manifest says, and say so

The package comment of corpus has said since it was written that the
manifest "is what makes the corpus a measurement rather than a pile: it
says where each file came from, when, how big it is and what its bytes
hash to, so a figure quoted from it can be reproduced and a file that
changed underneath can be noticed".

Nothing noticed. The hash was written and never read back, and a file in
no row was invisible to every tool here. A documented capability with no
implementation.

corpus.Check reports a row with no file, a file that will not open, a size
or digest that moved, and a file on disk in no row. harvest -check prints
them and exits NON-ZERO, so it can be the first line of a measuring script
rather than something to remember to run; harvest because harvest is what
writes the manifest.

Run over the two corpora:

    pdfforms   2268 rows, clean
    pdfscans   1012 rows, clean, and FOUR files on disk in no row

    ia-americana/calcflh_000254.pdf
    ia-americana/epn11-1968countclip2restricted.pdf
    ia-biodiversity/bidragtillknne38suom.pdf
    ia-medical/b22346703.pdf

Three of those four are documents our reader refuses, which is the whole
of why images and compare count 63 refusals and a directory-walking sweep
counts 66 (§33). And every digest that IS recorded verifies, so no
document has changed underneath -- 1 012 full hashes and 2 268 prefixes.

The digest is compared over AS MANY CHARACTERS AS THE MANIFEST RECORDED.
The forms corpus keeps sixteen under the header "sha256-8"; the scans
corpus keeps sixty-four. Sixteen hex characters is sixty-four bits, which
answers "did this file change" perfectly well, and a check demanding all
sixty-four would call every one of 2 268 rows changed and be deleted by
the first person to run it. That is the mutation this is proved against.

An unreadable file is reported as unreadable and NOT as missing: missing
would send a reader to re-fetch a document that is already on the disk.

One branch was removed rather than covered: filepath.Rel cannot fail on a
path WalkDir has just produced under its own root, and an error branch
nothing can reach is a branch nobody can test. The relative path is
trimmed instead.

Three mutations, each caught: comparing the whole digest instead of the
recorded prefix, not looking for files in no row, and calling an
unreadable file missing. 100.0% of statements.

* corpus: compare the manifest's slashes against the manifest's slashes

The windows-latest lane caught a defect in the check itself, not in its
tests: Entry.Path is slash-separated -- deliberately, so a corpus can be
moved -- and WalkDir hands back paths separated by the host's separator.
On Windows the comparison was "a\one.pdf" against the manifest's
"a/one.pdf", so EVERY file of a corpus came back as being in no manifest
row. A check that reports 3 280 disagreements on a clean corpus is worse
than no check.

filepath.ToSlash on the walked path, which is a no-op on the two platforms
I can run and the whole of the fix on the one I cannot.

And two tests are skipped there rather than failing: Chmod(0) denies
neither a file read nor a directory listing on Windows, so the unreadable
cases cannot be built at all. Skipped with the reason, beside the existing
skip for root, which cannot build them either.

This fix is verified by CI and not locally -- on macOS ToSlash is the
identity and the local run cannot tell the two versions apart. Said here
because it is the kind of claim worth marking.

* images: a page we drew no picture for is still judged

judgePage returned nil as soon as render.Images handed back an empty
slice, so a page we produce nothing for produced no row -- not a refusal,
not an unseen picture, nothing. It has done that since it was written.

The corpus's WORST page is one of those. Page 1 of
sim_unitarian-...-1825-06-25_4_25.pdf:

    render.Images  ->  0 images, err == nil
    render.Page    ->  1269x1775, 154 055 dark pixels
    pdfimages      ->  3 pictures of 7048x9856

Each picture is 69 465 088 pixels, 3.5% past the per-picture ceiling, and
decodeBase drops it by returning nil -- which is not an error. compare
measures that page at 19.23% of pixels differing from poppler, the worst
of 3 214, and this instrument measured it not at all. Filed as
go-pdfkit/render#108.

Which also makes the count added in §32 an UNDERCOUNT, and systematically
so for the heaviest pages -- the ones a reader would most want it for.

So the judge is asked, and what it took out is counted as Unseen, which is
exactly what that word means. A page neither side produced a picture for
returns no row still: that is what most pages of most documents look like,
and a row each would bury the ones above. A judge that will not finish is
Hung, named.

Two mutations, caught: returning nil again, and counting them as Neither.
Both run after a go build, because a mutation that does not compile has
not been tested -- which is how a mutation "survived" earlier in this
branch. 100.0% of statements.

* Treat the manifest and a fetched name as data something else wrote

A security audit of the harness, from one question: what in here is data
that something else wrote? Three answers, and two had no guard.

1. A MANIFEST PATH WAS READ VERBATIM AND JOINED TO THE CORPUS.

   Entry.Path comes out of MANIFEST.tsv -- a data file beside the
   documents, in a shared directory, which nothing here necessarily wrote
   -- and was joined to the corpus directory by os.Stat, os.Open, a hash
   and the argv of a poppler invocation. filepath.Join neutralises an
   ABSOLUTE path by construction and does NOT neutralise "..":

       Join("/corpus", "../../etc/passwd") == "/etc/passwd"

   corpus.Read now refuses such a row. Refuses rather than cleans: a row
   trying to leave the corpus is not a row with a typo, and rewriting it
   would put a file in a measurement under a name the manifest did not
   give.

2. A REMOTE SERVER'S IDENTIFIER BECAME A LOCAL FILE NAME.

   harvest names what it fetches ReplaceAll(id, "/", "_") + ".pdf", where
   id comes from the archive's search results. The slash was replaced,
   which blocks a directory escape; a leading dash was not. Every poppler
   tool parses its arguments with getopt and the document path is
   positional, so a corpus file called "-v.pdf" makes pdfimages print its
   version and take no picture -- and the harness records the page as one
   the judge drew nothing for. A wrong measurement, from a file name.

   Guarded twice on purpose: internal/poppler.Document prefixes "./" at
   every call site (pdfinfo, pdfimages, pdfimages -list, pdftoppm), and
   corpus.safeName strips a leading "-" or "." when the file is created so
   the awkward name never enters the corpus. The "." goes for §33's
   reason: a file nothing lists is measured by every tool that walks the
   manifest and by none that walks the directory.

3. What was already right is stated in §36 rather than assumed: no shell
   anywhere (exec takes an argv, so command injection was never available
   -- argument injection was), every poppler call bounded and a bound that
   fires named, a decode charged what it will cost before a byte is spent,
   and temporary files under os.MkdirTemp.

Both fixes mutation-checked, each after a go build: accepting a ".." row
fails two cases, leaving a leading dash alone fails the poppler test.
100.0% of statements.

* baseline: §35, the undercount measured, and a hung column that was empty because nobody asked

judgePage returned nil the moment render.Images handed back an empty slice,
so a page we produce no picture for produced no row at all. Re-taken over
all 23 populations with the fix:

    fr-cerfa        4153 -> 5351   +1198
    gh-safedocs        0 ->   11     +11
    ia-americana      24 ->   32      +8
    gh-pdfbox          4 ->   10      +6
    gh-qpdf           13 ->   19      +6
    four others                   +1 each
    ----------------------------------------
    23 populations  4277 -> 5510  +1233   = a 29% undercount

Nine of twenty-three moved. fr-cerfa's 1 198 come from THIRTEEN documents
whose page 1 draws only inline images -- 1 197 of the 1 198 rows carry
object 0 -- so those pages did not lose a count, they vanished whole. And
§34 had already found the consequence without looking for it: the page
compare ranks worst of 3 214, at 19.23%, is one of these, and so is the
second worst at 14.70%.

ONE CELL NOBODY EXPECTED. Everything else in all 23 records is identical
cell for cell, and the exception is not an unseen:

    gh-qpdf  hung: qpdf_qtest_qpdf_shared-unnamed-field.pdf page 1, pdfimages

That document draws no picture on page 1, so images NEVER ASKED poppler
about it -- which is precisely what this file and the top-level README
have both said for weeks to explain hung: 0. The sentence was right and it
described a hole, not a corpus. Asking the judge on a page we drew nothing
for asks about that document for the first time, and pdfimages does not
come back. §34's independent page comparison reached the same file, page
and tool. Two instruments, neither looking for it, agreeing on the one
document in 3 280 that stops a poppler tool.

The top-level README's hung paragraph and its unseen row are corrected to
1 and 5 510.

* images: bound what the judge's answer costs to read back

The audit's fourth finding, from carrying the same question past the
inputs a FILE provides to the inputs a TOOL provides. render refuses to
decode a page past a ceiling; three png.Decode calls read poppler's answer
about the same page with no bound at all.

Measured on the page of §35, where our own decoder refuses to spend
277 MB:

    PNGs pdfimages writes for page 1     3 x 7048 x 9856
    on disk                              44 MB
    png.Decode -> three *image.Gray      69.5 MB each
    raster.FromImage -> RGBA             278 MB each
    retained                             834 MB

An instrument that refuses to spend what it asks the other side to spend
is measuring two different things, and nothing here limits what poppler is
asked to write.

maxJudgePixels is the same number render bounds a page by, in the same
unit, read from png.DecodeConfig BEFORE a byte is decoded -- for the reason
afford() gives: a limit noticed after the allocation has not helped.
Today's corpus is inside it (that page is 3 x 69 465 088 = 208 395 264
against 268 435 456), so no figure in baseline/ moves; ia-texts re-taken
to confirm, cell for cell.

And the silence beside it is gone. A PNG that could not be read was
SKIPPED with continue, so the pairing saw fewer of the judge's pictures
and said nothing -- the same mistake as §35, three functions away. It is
reported now, the way the hung listing beside it already was.

readPNG reads the file once and decodes from memory twice rather than
seeking one open file back to the start: the Seek error is a branch
nothing can reach on a regular file, and a branch nothing can reach is a
branch nobody can test.

Mutation-checked, each after a go build: dropping the bound lets a picture
past the budget be read, restoring the continue turns a page that should
be reported into a page measured on nothing. 100.0% of statements.
@tannevaled
tannevaled merged commit a80d887 into main Oct 6, 2026
9 checks passed
@tannevaled
tannevaled deleted the form-resources-inherit branch October 6, 2026 13:33
tannevaled added a commit that referenced this pull request Oct 6, 2026
…an be compared at all (#105)

* A form XObject's incomplete /Resources no longer shadows the enclosing ones

qpdf's own fixture says it in its name. form-xobjects-some-resources2.pdf
draws six 15x15 images on page 1 through forms, some of which carry only
part of what they use, and this renderer drew TWO of them.

    images pdfimages -list takes out       6   (objects 10 12 14 16 18 20)
    images render.Images returned          2   (objects 10 12)

It was not only the extraction API. Through go-pdfkit/conformance's
compare against pdftoppm at 72 dpi, on these fixtures alone:

                           before    after
    pixels differing        8.54%    0.03%
    byte for byte          94.06%   99.56%
    mean |diff|            10.936    0.519  levels
    worst square mean      130.27     8.80  levels
    worst pixel               254      149

A worst square mean of 130 is not resampling; four of six stamps were
missing from the page.

Both paths had HALF the rule. drawForm and imagesDrawn each handled an
ABSENT /Resources by falling back to the parent, and a present one
replaced it wholesale -- so a name the form did not list resolved to
nothing and drawXObject returned early. A form XObject's resource
dictionary is not required to be complete, and a name it does not provide
is resolved in the resources in force where the form was painted, which is
what poppler does.

So: renderer.resChain, the dictionaries in force outermost first, pushed by
every stream that brings its own -- run() covers a form, a Type 3 glyph
procedure, a pattern cell, an annotation's appearance and a soft mask; the
extraction path pushes its own because it does not go through run(). The
seven lookups that each read one category out of ONE dictionary --
XObject, Font, ColorSpace, Pattern, Shading, ExtGState, Properties -- go
through named(), which tries the dictionary the caller holds and then
walks outward.

named() takes that dictionary rather than reading it off the chain, and
that is not decoration: the first attempt dropped the parameter, and
TestABDCThatNamesNoLayerHidesNothing failed at once because the tests call
these functions directly. A lookup whose first dictionary is implicit is a
lookup nobody can call.

Checked in both directions. On the whole forms corpus, 2 268 documents in
eighteen populations through conformance's images harness, exactly ONE
population moves: gh-qpdf, pictures 63 -> 73, and the ten recovered land
in (samples) direct as 51 of 51 exact and 51 of 51 byte-identical with
poppler. They are not merely present; they are right. Nothing else in any
population changes a cell.

Mutation-checked: removing the walk fails the three inheritance tests,
and dropping the caller's dictionary fails the one that calls directly.
100.0% of statements.

Fixes #102.

* Images returns the inline ones, so 4 147 pictures of one corpus population can be compared at all

Images returned nothing for a picture written into the content stream with
BI, and said why: "they are named by no resource and are objects of
nothing, so there is nothing for a tool that extracts a file's objects to
hand back beside them." The first half is true and the conclusion was
never measured. pdfimages EXTRACTS inline images and lists them with an
object of 0.

Measured over the conformance corpus's fr-cerfa, 450 forms, page 1 each:
4 147 pictures the reference got out and nothing of ours was ever compared
with -- 2 614 of them stencils, 2 592 of them 16x16, one document holding
2 592. A defect in inline-image decoding was invisible to every figure in
that baseline, because a picture nobody returns is a picture nobody
compares.

ONE ENTRY PER BI, which is the only unit that can be right: an inline
image IS its draw, there is no object to collapse repeats onto, and
pdfimages lists one row per draw as well. Object 0, which is what the file
gives it and what pdfimages prints. The name is an ordinal, BI#1 upward
across the whole call including across forms, because there is no resource
name and a reader of a difference has to be able to say which picture it
was. No mask is read beside it: an inline dictionary has no abbreviation
for /SMask.

Over the WHOLE forms corpus, 2 268 documents in eighteen populations,
through conformance's images harness:

    pictures         5 917 -> 11 296
    unseen           4 243 -> 80
    paired by size       0 -> 5 368

The third line is the honest caveat and it is not a side effect: the
pairing falls back to size exactly when an object number is missing, and
until now the one class whose number is missing never reached it. That
fallback has been in the harness for weeks with a share of zero. These
5 368 pictures are matched by size and order, which is weaker than by
object, and conformance#92 made that share visible for this reason.

In fr-cerfa 2 729 of the new pictures come out exact and byte-identical
and none differs. The differing ones are three rows of ONE safedocs
fixture, Inline_Image_Abbreviations_InlineAbbreviations.pdf, whose eight
dictionaries each carry the abbreviation AND the long name with
CONFLICTING values -- /W 20 /Width 10 /H 10 /Height 40 /BPC 8
/BitsPerComponent 4, with a comment reading "COMMENT OUT THIS LINE TO SEE
EFFECT". poppler resolves those in favour of the long name, hence its
10x40 at 4 bits against our 20x10 at 8. reader's InlineImage.Expanded
already documents that choice, cites Table 93 for it, and notes that
pdf.js and MuPDF disagree with each other about it. Nothing here is a
defect; it is an ambiguity that was decided on purpose and could not be
seen until the pictures reached a comparison.

Mutation-checked: dropping the ordinal gives "BI#0", advancing it before
the decode succeeds leaves a gap a reader cannot account for, and skipping
afford() lets a page past the budget come back without an error.
100.0% of statements.

Based on #103, which the same counter found. Fixes #101.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A form XObject with INCOMPLETE /Resources shadows the page's: four of six images are not drawn

1 participant