PDF Links & Annotations (PyMuPDF)
Read the hyperlinks and sticky notes
- links_json
- annotations_json
This is the niche one, and it knows it. PDF Links & Annotations (PyMuPDF) doesn't extract text or render pages - it reads the structure on top of them: every clickable link and every annotation (comments, highlights, sticky notes) in a PDF, exported as JSON. You won't need it for most documents, but when your PDF is a report with a table of contents full of hyperlinks, or a reviewed document covered in comments, it's the only node in the pack that can see those.
How it works
Two PyMuPDF page methods, two outputs. page.get_links() walks the clickable areas - internal links to other pages as well as external URLs - and returns them with their type, source rectangle, and destination. page.annots() walks the annotations and the node keeps the fields that matter: the annotation type (text/comment, highlight, underline, square, and the rest of PyMuPDF's enum), its rectangle, and the content/title pulled from the annotation info dict.
Inputs are the standard trio: pdf (the PYMUPDF_PDF socket from Load PDF (PyMuPDF)), page_range (default all, else 1-based lists and 1-5 ranges), and max_pages (default 0 = no cap). Batch mode works, so a folder of linked PDFs becomes one JSON blob of every URL in all of them.
What comes out
links_json(STRING) - one object per link: the PDF, 1-based page, and PyMuPDF's full link dict (type, rect, and the URI or page target).annotations_json(STRING) - per annotation: PDF, page, type, rect, content, title.
Both are plain STRING outputs, so you can preview them, parse them with a JSON node, or hand them to an LLM for a structured crawl - "extract every URL in this PDF and tell me what each one points to" is the kind of job this node was made for. For a document-intake or audit workflow, it's the difference between knowing a PDF references things and knowing exactly what it references.
Install & gotchas
Same pack install: ComfyUI Manager or clone + pip install -r requirements.txt, then restart. No models, no downloads.
Two honest caveats. First, annotations in PDFs are a mess in the wild - many producers write them in nonstandard ways, so content comes back empty more often than you'd like; that's the file's fault, not the node's. Second, this is a pure structural reader: if a page's text is what you're after, you want Extract PDF Text (PyMuPDF) instead. It's a narrow tool, but it fills a gap nothing else in the pack does.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| PYMUPDF_PDF | — | ||
| page_range | STRING | all | — |
| max_pages | INT | 00–10000 | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| links_json | STRING | — |
| annotations_json | STRING | — |