PDF Info (PyMuPDF)
The metadata and table of contents reader
- metadata_json
- toc_json
- page_count
PDF Info (PyMuPDF) is the pack's quiet metadata reader: it doesn't render or extract anything, it just tells you what a PDF is about before you decide what to do with it. Feed it a pdf socket from Load PDF (PyMuPDF) and you get the document's metadata, its table of contents, and its page count - all as data you can route on.
How it works
The mechanism is a handful of PyMuPDF document properties exposed with almost no ceremony. For each PDF the node reads:
doc.metadata- the standard PDF info dictionary: title, author, subject, keywords, creator, producer, creation/modification dates, and the like.doc.get_toc()- the table of contents as PyMuPDF returns it: a list of[level, title, page]entries, which is what the sidebar in a PDF reader is built from.doc.is_encryptedanddoc.page_count- whether the file is locked and how many pages it has.
It also returns page counts summed across the whole batch, so in batch mode this doubles as a quick inventory of a folder of documents.
What comes out
Three outputs, all straightforward:
metadata_json(STRING) - per-file metadata, one object per PDF with filename, page count, encryption flag, and the metadata dict.toc_json(STRING) - the table of contents per file.page_count(INT) - total pages across all loaded PDFs.
The two JSON strings are STRING-typed, so they'll plug into any text display, a JSON parsing node, or - more usefully - a routing decision: a page_count INT is perfect for a switch that sends short documents down one path and long ones down another. For a doc-intake workflow, this is the node that answers "is this worth processing and how big is it" before you spend time rendering anything.
Install & gotchas
Same as the rest of the pack - ComfyUI Manager or clone + pip install -r requirements.txt (a single PyMuPDF dependency), then restart. No model downloads.
Honest caveat: doc.metadata is only as good as what the producer wrote. A lot of PDFs ship with title/author empty or generic, so don't treat missing fields as a bug in the node. And if the file is encrypted, this node will still read basic info but downstream extractors may choke - that's a PyMuPDF limitation, not this node's. It's a thin, no-surprises utility: you'll mostly wire it in when a workflow needs to branch on document size or log what it processed.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| PYMUPDF_PDF | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| metadata_json | STRING | — |
| toc_json | STRING | — |
| page_count | INT | — |