Render PDF Pages (PyMuPDF)
Pages as an IMAGE batch, ready for img2img
- images
- metadata_json
- page_count
This is the node that turns a PDF page into pixels your models can actually see. Render PDF Pages (PyMuPDF) rasterizes selected pages and hands them back as a ComfyUI IMAGE batch - the exact thing an img2img pass, a ControlNet, an upscaler, or a vision model wants. If you've ever dragged a PDF out of a viewer, screenshotted it, and manually cropped the margins, this is the node that kills that ritual.
How it works
PyMuPDF renders a page by painting it into a pixmap at a scale determined by DPI: the code builds fitz.Matrix(dpi / 72.0) and calls page.get_pixmap(matrix=matrix), so the default 96 DPI renders at about 96/72 = 1.33× the page's point size. That DPI number is the first dial you'll touch:
dpi- default96, range36–300. Crank it up for crisp text when you're feeding pages to an upscaler or a detail-hungry model; drop it for quick previews. Higher DPI means much bigger tensors, and a 300-DPI A4 page is a lot of pixels to push through a sampler.page_range- default"1", 1-based, comma-separated lists and ranges (1,3-5), orall.max_pages- caps how many pages render (default20). A 100-page PDF defaults to your first twenty, which is usually the safety you want.resize_mode- how to make a ragged batch rectangular:fit_first(resize everything to the first page's size, the default),fit_smallest, orpad_largest(zero-pad to the biggest canvas).
What comes out
images(IMAGE) - a stacked batch, one tensor per rendered page. Pages of different sizes get forced into a rectangle byresize_mode, so if a document has mixed page sizes you'll see either squished pages (fit_first) or black bars (pad_largest). For a consistent document, it's a non-issue.metadata_json(STRING) - per-page info: filename, 1-based page number, and the DPI used. Useful if you render selective pages and need to remember which is which.page_count(INT) - total pages in the document, not pages rendered.
Where it goes
The obvious wiring: images → VAE encode → img2img, or into a ControlNet preprocessor if you want structure (edge maps, depth) of a page layout to guide generation. It also plays nicely with the document-to-LLM crowd - a vision-language model reading the page instead of the extracted text is the right call for scanned PDFs that have no text layer at all.
Install & gotchas
Same one-line story: ComfyUI Manager or clone + pip install -r requirements.txt (just PyMuPDF), restart. The realistic failure modes are memory (a long page range at high DPI eats VRAM fast - respect max_pages) and the classic environment mismatch (node won't appear because PyMuPDF isn't installed into ComfyUI's own Python; install pymupdf, not the stale decoy fitz package on PyPI). Renders can't come from an encrypted PDF either - decrypt it first with something like qpdf before pointing this at it.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| PYMUPDF_PDF | — | ||
| page_range | STRING | 1 | — |
| dpi | INT | 9636–300 | — |
| max_pages | INT | 201–1000 | — |
| resize_mode | COMBO | 3 options: fit_first, fit_smallest, pad_largest |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | — |
| metadata_json | STRING | — |
| page_count | INT | — |