Nodes/ComfyUI PyMuPDF/PDF Links & Annotations (PyMuPDF)
ComfyUI Node

PDF Links & Annotations (PyMuPDF)

Read the hyperlinks and sticky notes

By Stone-dielianhua·Created 3 months ago·Updated 3 months ago· 1
PDF Links & Annotations (PyMuPDF)
  • pdf
  • links_json
  • annotations_json
◄page_rangeall►
◄max_pages0►

This is the niche one, and it knows it. PDF Links & Annotations (PyMuPDF) doesn't extract text or render pages - it reads the structure on top of them: every clickable link and every annotation (comments, highlights, sticky notes) in a PDF, exported as JSON. You won't need it for most documents, but when your PDF is a report with a table of contents full of hyperlinks, or a reviewed document covered in comments, it's the only node in the pack that can see those.

How it works

Two PyMuPDF page methods, two outputs. page.get_links() walks the clickable areas - internal links to other pages as well as external URLs - and returns them with their type, source rectangle, and destination. page.annots() walks the annotations and the node keeps the fields that matter: the annotation type (text/comment, highlight, underline, square, and the rest of PyMuPDF's enum), its rectangle, and the content/title pulled from the annotation info dict.

Inputs are the standard trio: pdf (the PYMUPDF_PDF socket from Load PDF (PyMuPDF)), page_range (default all, else 1-based lists and 1-5 ranges), and max_pages (default 0 = no cap). Batch mode works, so a folder of linked PDFs becomes one JSON blob of every URL in all of them.

What comes out

  • links_json (STRING) - one object per link: the PDF, 1-based page, and PyMuPDF's full link dict (type, rect, and the URI or page target).
  • annotations_json (STRING) - per annotation: PDF, page, type, rect, content, title.

Both are plain STRING outputs, so you can preview them, parse them with a JSON node, or hand them to an LLM for a structured crawl - "extract every URL in this PDF and tell me what each one points to" is the kind of job this node was made for. For a document-intake or audit workflow, it's the difference between knowing a PDF references things and knowing exactly what it references.

Install & gotchas

Same pack install: ComfyUI Manager or clone + pip install -r requirements.txt, then restart. No models, no downloads.

Two honest caveats. First, annotations in PDFs are a mess in the wild - many producers write them in nonstandard ways, so content comes back empty more often than you'd like; that's the file's fault, not the node's. Second, this is a pure structural reader: if a page's text is what you're after, you want Extract PDF Text (PyMuPDF) instead. It's a narrow tool, but it fills a gap nothing else in the pack does.

Categorypdf/PyMuPDF

Inputs (3)

NameTypeDefaultDescription
pdfPYMUPDF_PDF—
page_rangeSTRINGall—
max_pagesINT00–10000—

Outputs (2)

NameTypeDescription
links_jsonSTRING—
annotations_jsonSTRING—