Nodes/ComfyUI-DonutNodes/Donut Load VAE
ComfyUI Node

Donut Load VAE

The loader that stops a 2x VAE from doubling your image

By DonutsDelivery·Created 2 years ago·Updated about 19 hours ago· 26
Donut Load VAE
    • VAE
    ◄vae_name▾►

    Why a second VAE loader exists at all

    ComfyUI already has a VAE loader. Donut Load VAE is a subclass with the same single dropdown, so at first glance it's redundant - and for a normal VAE file it genuinely is, because ordinary VAEs pass straight through unchanged.

    The reason it exists is Spacepxl's Wan 2.1 VAE upscale 2x, the current community answer to a specific complaint: Krea 2 (and Qwen-Image generally) uses an autoencoder whose decoder was fine-tuned for legible small text at the cost of fine texture, so skin comes back looking airbrushed. The Qwen encoder was frozen and the Wan 2.1 encoder is identical, which is why dropping the Wan decoder into a Qwen-latent workflow works instead of producing noise - same latent space, different decoder head. Krea's own CEO has said he'd use the Flux VAE for photorealism, and the "this VAE is the worst for DiT models" thread is what pushed the swap mainstream.

    Here's the catch. That decoder doesn't output one image. It outputs 2x RGB, packing twelve channels into each decoder position, and ComfyUI's stock decode path doesn't know what to do with it. You get either a surprise 2x image or the memorable save-node error Cannot handle this data type: (1, 1, 12), |u1. Spacepxl ships an adapter for this in ComfyUI-VAE-Utils; Donut's loader wires that adapter in and hides it behind one dropdown.

    How it works

    On load, Donut sniffs the VAE rather than trusting the filename: latent_dim == 3, latent_channels == 16, conv_out_channels == 12, output_channels == 3. That's the Wan 16-channel latent with the 12-channel RGB head. If it matches, the loader applies the upstream VAEUtils_PatchWanUpscaleVAE node, then wraps only the VAE's public decode and decode_tiled methods so that after the decoder unpacks its 2x RGB, Donut runs an antialiased bilinear reduction back down to latent size × 8. Encoder, latent layout, weights, model management and offloading stay ComfyUI's.

    The practical upshot: your canvas doesn't change. The file is called "upscale 2x" but selecting it does not enlarge your sampling resolution, previews, or final image. A configured 896×1152 base stays 896×1152; a hires step at rescale 1.5 samples 1344×1728, decodes internally at 2688×3456, and is filtered immediately to 1344×1728. Spacepxl recommends filtering and downsampling if you want to keep the original resolution - the exact filter is Donut's choice, not a spec from the model card, and the decoder is trained for perceived texture rather than pixel-exact reproduction. A flag makes preparation idempotent, so the same VAE can pass through several Donut stages without being wrapped twice.

    If the VAE doesn't match the signature, nothing happens. It's a stock loader with a friendlier name.

    The one input, and the output

    • vae_name - the dropdown, listing whatever ComfyUI finds under ComfyUI/models/vae/ (subfolders included, so the catalogue's qwen-image/qwen_image_vae.safetensors shows up as a path). That's the entire input surface.

    Output is a single VAE, a drop-in for the socket any sampler, decode or detailer expects. In the pack's V5 workflow that one selection supplies base generation, hires, face detail, reference encoding and damage correction - one file, loaded once.

    Installing it

    cd ComfyUI/custom_nodes
    git clone https://github.com/DonutsDelivery/ComfyUI-DonutNodes.git donutnodes
    cd donutnodes
    python -m pip install -r requirements.txt
    

    Or just search DonutNodes in ComfyUI Manager. Either way, restart ComfyUI and refresh the browser afterwards: the author stresses that the registry package and the repository code are versioned separately, and an older registry build won't have the current node set. For the 2x VAE you also need:

    cd ComfyUI/custom_nodes && git clone https://github.com/spacepxl/ComfyUI-VAE-Utils
    # then drop the checkpoint into models/vae/
    # Wan2.1_VAE_upscale2x_imageonly_real_v1.safetensors (~507 MB, 507,684,560 bytes)
    

    Spacepxl's repo publishes a SHA-256 for that file (2413554b…6b63e06 at the pinned revision) if you want to verify it. DonutNodes' own model catalogue can fetch it for you via the workflow's "Download missing" button - and only that button; nothing downloads behind your back.

    Where people get burned

    Missing VAE-Utils. The loader raises a very literal error if the patch node isn't there: "The selected 2x VAE requires ComfyUI-VAE-Utils by spacepxl. Install/update that node pack and restart ComfyUI." That's the failure to expect if you copied someone's workflow JSON and grabbed only one of the two repositories.

    Stale VAE-Utils or stale ComfyUI. If the prepared decode doesn't return the expected RGB shape you get "The selected 2x VAE did not return the expected RGB image. Update ComfyUI and ComfyUI-VAE-Utils." Different error, different fix - update code, don't re-download the 507 MB file.

    The dropdown not listing your file. It's a filesystem listing, not a catalogue. The file must sit somewhere under ComfyUI/models/vae/, and you have to refresh the page after dropping it in.

    Assuming it's an upscaler. It isn't. If you want the 2x output, this node is deliberately the wrong tool. Also budget VRAM for the decoder's own 2x intermediate: the docs are candid that shrinking the sampling canvas after the fact is not a VRAM fix, and their recorded 4070 failure was an OOM during NAG sampling, not decode.

    The OpenCV pile-up. Donut's requirements floor opencv-python-headless; if another pack already installed a full or contrib cv2, you share one binary. Since DonutNodes refuses to pip-repair your environment on startup, that one's on you.

    Categorydonut/model

    Inputs (1)

    NameTypeDefaultDescription
    vae_nameCOMBO1 options: pixel_space

    Outputs (1)

    NameTypeDescription
    VAEVAE—