Nodes/ComfyUI PyMuPDF/Render PDF Pages (PyMuPDF)
ComfyUI Node

Render PDF Pages (PyMuPDF)

Pages as an IMAGE batch, ready for img2img

By Stone-dielianhua·Created 3 months ago·Updated 3 months ago· 1
Render PDF Pages (PyMuPDF)
  • pdf
  • images
  • metadata_json
  • page_count
◄page_range1►
◄dpi96►
◄max_pages20►
◄resize_mode▾►

This is the node that turns a PDF page into pixels your models can actually see. Render PDF Pages (PyMuPDF) rasterizes selected pages and hands them back as a ComfyUI IMAGE batch - the exact thing an img2img pass, a ControlNet, an upscaler, or a vision model wants. If you've ever dragged a PDF out of a viewer, screenshotted it, and manually cropped the margins, this is the node that kills that ritual.

How it works

PyMuPDF renders a page by painting it into a pixmap at a scale determined by DPI: the code builds fitz.Matrix(dpi / 72.0) and calls page.get_pixmap(matrix=matrix), so the default 96 DPI renders at about 96/72 = 1.33× the page's point size. That DPI number is the first dial you'll touch:

  • dpi - default 96, range 36–300. Crank it up for crisp text when you're feeding pages to an upscaler or a detail-hungry model; drop it for quick previews. Higher DPI means much bigger tensors, and a 300-DPI A4 page is a lot of pixels to push through a sampler.
  • page_range - default "1", 1-based, comma-separated lists and ranges (1,3-5), or all.
  • max_pages - caps how many pages render (default 20). A 100-page PDF defaults to your first twenty, which is usually the safety you want.
  • resize_mode - how to make a ragged batch rectangular: fit_first (resize everything to the first page's size, the default), fit_smallest, or pad_largest (zero-pad to the biggest canvas).

What comes out

  • images (IMAGE) - a stacked batch, one tensor per rendered page. Pages of different sizes get forced into a rectangle by resize_mode, so if a document has mixed page sizes you'll see either squished pages (fit_first) or black bars (pad_largest). For a consistent document, it's a non-issue.
  • metadata_json (STRING) - per-page info: filename, 1-based page number, and the DPI used. Useful if you render selective pages and need to remember which is which.
  • page_count (INT) - total pages in the document, not pages rendered.

Where it goes

The obvious wiring: images → VAE encode → img2img, or into a ControlNet preprocessor if you want structure (edge maps, depth) of a page layout to guide generation. It also plays nicely with the document-to-LLM crowd - a vision-language model reading the page instead of the extracted text is the right call for scanned PDFs that have no text layer at all.

Install & gotchas

Same one-line story: ComfyUI Manager or clone + pip install -r requirements.txt (just PyMuPDF), restart. The realistic failure modes are memory (a long page range at high DPI eats VRAM fast - respect max_pages) and the classic environment mismatch (node won't appear because PyMuPDF isn't installed into ComfyUI's own Python; install pymupdf, not the stale decoy fitz package on PyPI). Renders can't come from an encrypted PDF either - decrypt it first with something like qpdf before pointing this at it.

Categorypdf/PyMuPDF

Inputs (5)

NameTypeDefaultDescription
pdfPYMUPDF_PDF—
page_rangeSTRING1—
dpiINT9636–300—
max_pagesINT201–1000—
resize_modeCOMBO3 options: fit_first, fit_smallest, pad_largest

Outputs (3)

NameTypeDescription
imagesIMAGE—
metadata_jsonSTRING—
page_countINT—