Nodes/ComfyUI PyMuPDF/Save PDF Text (PyMuPDF)
ComfyUI Node

Save PDF Text (PyMuPDF)

Get the extracted text out of the graph and onto disk

By Stone-dielianhua·Created 3 months ago·Updated 3 months ago· 1
Save PDF Text (PyMuPDF)
    • path
    ◄text—►
    ◄filename_prefixpymupdf/text►
    ◄extension▾►

    Every text-extracting node in this pack ends in a STRING socket - and a string sitting on a wire is a string you eventually want as a file. Save PDF Text (PyMuPDF) is that last mile: take whatever text or JSON you extracted and write it to ComfyUI's output directory as a real, shareable, re-openable file. It's an output node, the pack's version of Save Image but for text.

    How it works

    It's gloriously boring plumbing. The node takes the incoming text STRING, picks a filename and extension, writes it to folder_paths.get_output_directory() - the same ComfyUI/output/ folder your images land in - and returns the path it wrote. Two details are worth knowing because they're where the guardrails live:

    • filename_prefix (default pymupdf/text) - the path + stem, relative to the output dir. The pymupdf/ part becomes a subfolder, so your files land in ComfyUI/output/pymupdf/. The code scrubs the filename of characters that don't belong, and it refuses to write outside the output directory entirely.
    • extension - txt, json, html, xml, or md. Pick based on what you extracted: plain text for text-mode extracts, json for the JSON/markup formats from Extract PDF Text or the _json outputs elsewhere in the pack, md if you want something readable.

    The one behavior that saves you from yourself: it never overwrites. If text.txt exists, it writes text_00001.txt, then text_00002.txt, and so on. Run a workflow twice and you get both files instead of losing the first extract - exactly the right instinct for a data-collection tool, and the reason you don't need to hand-rename things between runs.

    What comes out

    • path (STRING) - the absolute path of the file it just wrote, which you can preview, log, or route to whatever consumes files.

    Where it goes in a workflow

    Wire the text output of Extract PDF (PyMuPDF) or Extract PDF Text (PyMuPDF) into the text input, set an extension, and you've got a one-click document-to-file pipeline. It's the natural end of the document-to-LLM flow too: extract, summarize in-graph, and save the summary as .md for later. Since it's an output node, ComfyUI always runs it, so you never have to worry about the cache skipping your save.

    Install & gotchas

    The pack-standard install: ComfyUI Manager or clone + pip install -r requirements.txt (just PyMuPDF), restart. Nothing model-related to download.

    The one thing to keep in mind: the text input has forceInput, meaning it's a wire, not a text box - you have to connect something, which is fine because you're building this node to receive an extract anyway. And if you're on a workflow you downloaded, check the filename_prefix before running, since a shared workflow will write into whatever prefix it shipped with. If the output path ever tries to escape the output directory, the node raises an error rather than writing where it shouldn't - a small, sensible wall between you and a mess.

    Categorypdf/PyMuPDF

    Inputs (3)

    NameTypeDefaultDescription
    textSTRING—
    filename_prefixSTRINGpymupdf/text—
    extensionCOMBO5 options: txt, json, html, xml, md

    Outputs (1)

    NameTypeDescription
    pathSTRING—