Nodes/ComfyUI PyMuPDF/PDF Info (PyMuPDF)
ComfyUI Node

PDF Info (PyMuPDF)

The metadata and table of contents reader

By Stone-dielianhua·Created 3 months ago·Updated 3 months ago· 1
PDF Info (PyMuPDF)
  • pdf
  • metadata_json
  • toc_json
  • page_count

PDF Info (PyMuPDF) is the pack's quiet metadata reader: it doesn't render or extract anything, it just tells you what a PDF is about before you decide what to do with it. Feed it a pdf socket from Load PDF (PyMuPDF) and you get the document's metadata, its table of contents, and its page count - all as data you can route on.

How it works

The mechanism is a handful of PyMuPDF document properties exposed with almost no ceremony. For each PDF the node reads:

  • doc.metadata - the standard PDF info dictionary: title, author, subject, keywords, creator, producer, creation/modification dates, and the like.
  • doc.get_toc() - the table of contents as PyMuPDF returns it: a list of [level, title, page] entries, which is what the sidebar in a PDF reader is built from.
  • doc.is_encrypted and doc.page_count - whether the file is locked and how many pages it has.

It also returns page counts summed across the whole batch, so in batch mode this doubles as a quick inventory of a folder of documents.

What comes out

Three outputs, all straightforward:

  • metadata_json (STRING) - per-file metadata, one object per PDF with filename, page count, encryption flag, and the metadata dict.
  • toc_json (STRING) - the table of contents per file.
  • page_count (INT) - total pages across all loaded PDFs.

The two JSON strings are STRING-typed, so they'll plug into any text display, a JSON parsing node, or - more usefully - a routing decision: a page_count INT is perfect for a switch that sends short documents down one path and long ones down another. For a doc-intake workflow, this is the node that answers "is this worth processing and how big is it" before you spend time rendering anything.

Install & gotchas

Same as the rest of the pack - ComfyUI Manager or clone + pip install -r requirements.txt (a single PyMuPDF dependency), then restart. No model downloads.

Honest caveat: doc.metadata is only as good as what the producer wrote. A lot of PDFs ship with title/author empty or generic, so don't treat missing fields as a bug in the node. And if the file is encrypted, this node will still read basic info but downstream extractors may choke - that's a PyMuPDF limitation, not this node's. It's a thin, no-surprises utility: you'll mostly wire it in when a workflow needs to branch on document size or log what it processed.

Categorypdf/PyMuPDF

Inputs (1)

NameTypeDefaultDescription
pdfPYMUPDF_PDF—

Outputs (3)

NameTypeDescription
metadata_jsonSTRING—
toc_jsonSTRING—
page_countINT—