Extract PDF (PyMuPDF)
Images, text, and metadata in one pass
- images
- text
- metadata_json
- page_count
- image_count
This is the kitchen sink node of the pack: give it a PDF (from Load PDF (PyMuPDF)) and it pulls out the embedded images, the text, and per-item metadata in a single run. If your document pipeline needs both "what's in this PDF visually" and "what does it say," this is the one to reach for - the more focused text-only and render-only siblings exist for when you don't want to pay for both.
How it works
Under the hood it's straight PyMuPDF. For every page it runs page.get_images(full=True) to list the images actually embedded in the file and turns each one into a tensor via fitz.Pixmap (routed through a PyMuPDF RGB conversion when the source has an alpha channel, so the batch stays clean RGB). Text comes from page.get_text("text"), chunked per page with a --- filename page N --- header so you can tell where each piece came from. It then stacks everything into a single IMAGE batch and returns the text as one big STRING.
The knobs that matter:
extract_images/extract_text- flip one off and you skip half the work. For a text-heavy document you don't need images; for an image scan you don't need text.max_pagesandmax_images- the guardrails. PDFs can be enormous, and the code is honest about it:max_pagescaps how many pages get looked at,max_imagescaps how many images get pulled total. Set them to 0 for no limit, but know that a 400-page magazine will happily try to extract hundreds of images into one tensor batch and run you out of RAM.render_pages_if_no_images- the clever fallback. When a page has no embedded images (common with scanned or vector-heavy PDFs), this renders that page to an image atrender_dpiinstead, so you still get something visual out.resize_mode-fit_firstsizes everything to the first image's dimensions,fit_smallestto the smallest,pad_largestzero-pads the smaller ones to the biggest canvas. This exists because a tensor batch must be a rectangle; the defaultfit_firstis usually fine.
What comes out
Five outputs, and they map to the three things you'll wire:
images(IMAGE) - into an img2img or ControlNet pipeline, or just a preview.text(STRING) - straight into a text display node, an LLM node for document Q&A, orSave PDF Text (PyMuPDF)for a file on disk.metadata_json(STRING) - per-image info: which PDF, which page, whether it was an embedded image or a page render, and the DPI.page_countandimage_count(INT) - handy for routing or for a text overlay.
Install & gotchas
Same story as the rest of the pack: ComfyUI Manager or clone + pip install -r requirements.txt (PyMuPDF only), then restart. If the node doesn't show up, PyMuPDF isn't importable in ComfyUI's Python environment - install pymupdf (not fitz, there's a stale decoy package by that name) into the environment ComfyUI actually runs from.
The one thing that reliably trips people: batch mode. Load PDF (PyMuPDF) set to batch hands this node all the PDFs in the input folder, and the two limits behave differently. max_pages is applied per document - each PDF in the batch gets its own 20-page cap. max_images, by contrast, is a global accumulator: once 20 images have been pulled across all files, the whole run stops, so an image-heavy folder can get you 20 images total even if you expected 20 per doc. Set both knowing that.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| PYMUPDF_PDF | — | ||
| page_range | STRING | all | — |
| extract_images | BOOLEAN | true | — |
| extract_text | BOOLEAN | true | — |
| render_pages_if_no_images | BOOLEAN | false | — |
| render_dpi | INT | 9636–300 | — |
| max_pages | INT | 200–10000 | — |
| max_images | INT | 200–10000 | — |
| resize_mode | COMBO | 3 options: fit_first, fit_smallest, pad_largest |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| text | STRING | — |
| metadata_json | STRING | — |
| page_count | INT | — |
| image_count | INT | — |