Skip to content

Index ePub documents like PDFs - #345

Open
Sriram-PR wants to merge 1 commit into
openzim:mainfrom
Sriram-PR:epub-index-data
Open

Sriram-PR wants to merge 1 commit into
openzim:mainfrom
Sriram-PR:epub-index-data

Conversation

@Sriram-PR

Copy link
Copy Markdown

Fix #333

PyMuPDF, which scraperlib already uses for PDF indexing, also opens ePub. So instead of a new parser, get_pdf_index_data now goes through a shared _get_document_index_data(filetype=...), and get_epub_index_data is the ePub entry point on top of it. The title is built the same way from the document metadata (title - author). StaticItem auto-indexes application/epub+zip items in the content, fileobj and filepath cases, next to the PDF branch.

The filetype is now passed explicitly to pymupdf.open() for both, rather than relying on detection from the stream.

MuPDF says unknown epub version: 3.0 for every EPUB 3 file but reads it fine, so that message is added to IGNORED_MUPDF_MESSAGES. Real ePubs also make it warn about images it cannot load (html: cannot load image src='...') and CSS it does not support (syntax error: css syntax error: ..., e.g. on the img[src*="..."] selectors in Wikisource exports). Those carry the file name or the CSS, so they go in a new IGNORED_MUPDF_MESSAGE_PREFIXES matched on the start of the line. Neither matters for the text. Without these, practically every ePub item would log a "PyMuPDF issues" warning.

Tried on a real book first: Gutenberg's Pride and Prejudice EPUB 3 (24 MB with images) gives 733k characters of text and "Pride and Prejudice - Jane Austen" as title in about 0.3s.

Tests: a small EPUB 3 is built by a session fixture in tests/conftest.py (two chapters plus title/author metadata, a missing image and one such CSS selector), so there is no binary file to review. New tests cover get_epub_index_data from content, fileobj and filepath, check that none of the warnings above is logged, and check that an auto-indexed ePub item can be found in a ZIM by full-text search and by title suggestion. The item tests fail without the items.py change, and the warning check fails if the new ignored message is removed. Locally: tests/zim 311 passed (and the warning check fails without the new ignored messages), coverage of indexing.py and items.py is 100%, and ruff check, ruff format --check and pyright are clean. The only failures in the full suite are 22 gif tests, which fail the same way on main because I don't have gifsicle installed.

One thing worth flagging: since auto_index defaults to True, scrapers that already add ePub files will start indexing their content after upgrading, as happened for PDFs. That is what #333 asks for, but it means more build time and a bigger index for them. Scrapers that don't want it can pass auto_index=False. papers already does, so for openzim/papers#43 the plan is to call get_epub_index_data explicitly for the one format picked per book.

CHANGELOG entry added under Unreleased / Added.

PyMuPDF, which is already used to index PDFs, also reads ePub. So the
PDF indexing code becomes a shared helper taking the document type,
with get_pdf_index_data and a new get_epub_index_data on top of it, and
StaticItem auto-indexes application/epub+zip items the same way it does
PDFs.

MuPDF warns "unknown epub version: 3.0" on every EPUB 3 document while
reading it fine, so that message joins the ignored ones. Warnings about
images it cannot load and CSS it does not support are ignored by prefix,
since they carry document-specific details and do not affect the text.

Fix openzim#333
@Sinkleberg

Copy link
Copy Markdown

Checked the EPUB paths with synthetic books on Python 3.14. The indexing suite passed (23 tests, one skipped), plus 12 extra checks covering custom titles, custom index data, auto_index=False, Unicode metadata and a missing author. All three input forms preserved the book bytes in the resulting ZIM; search respected the indexing options.

I didn't find a problem in those checks. I couldn't validate the whole ZIM suite in this container: download-dependent tests failed with networking disabled or wget unavailable, and a permissions test failed while running as root.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add utility to index ePub documents content

2 participants