Nodes/ComfyUI-DonutNodes/Donut VAE Decode
ComfyUI Node

Donut VAE Decode

Decode your latents, then subtract the VAE's own damage

By DonutsDelivery·Created 2 years ago·Updated about 19 hours ago· 26
Donut VAE Decode
  • samples
  • vae
  • IMAGE
◄vae_damage_correctionfalse►
◄vae_damage_strength1.00►

What this actually is

Every workflow ends in a VAE decode: the diffusion model works on latents, and the VAE turns those latents back into pixels. ComfyUI ships a stock VAEDecode for that, and Donut VAE Decode is a drop-in replacement - same samples in, same VAE in, same IMAGE out. What it adds is one genuinely unusual option: subtract the damage the VAE just did to your image, without a reference.

That's the whole pitch. It's not an upscaler and it doesn't change your resolution. If you've read that the Qwen-Image VAE (what Krea 2 uses) airbrushes fine skin texture into plastic, this claws some back where the loss happens instead of bolting another sampler pass onto the end. If you're on an SDXL or Flux checkpoint where you like the decoder, you have no reason to touch this node.

How it works

Be careful with the name - "damage correction" sounds like restoration, and it isn't. There's no original to compare against. Donut does this: decode normally, then round-trip the decoded pixels back through the same VAE (encode → decode), and subtract the difference.

output = clip(image + strength * (image - roundtrip))

The round trip is the lossy step, so image - roundtrip is a map of what the VAE threw away. Adding that map back onto the image pushes back against the loss - the same intuition as unsharp masking, with the VAE as the blur kernel. Because it's derived from the VAE's own error, no reference image, mask or restoration model is involved.

Two details worth knowing. The source splits the image batch and gives each sample its own round trip, because video-capable VAEs happily read a batch of stills as a time sequence; padding handles dimensions that don't divide evenly by the VAE grid. The strength slider scales that same correction rather than adding iterations, so 2.0 isn't "twice as corrected" - the tooltip warns it amplifies artifacts.

If you load the Spacepxl Wan 2.1/Qwen 2x VAE, this node understands it: the decoder emits 2x RGB internally and Donut filters it back to the configured size before any of the above happens. That path requires ComfyUI-VAE-Utils to be installed, because the adapter comes from there.

The inputs that matter

  • samples - the LATENT, from your sampler. Wire it exactly where you'd wire stock VAEDecode.
  • vae - the VAE model. Best practice in this pack is to feed it from Donut Load VAE, especially if you're using one of the 2x decoder files.
  • vae_damage_correction - boolean, off by default. The actual feature.
  • vae_damage_strength - float, 0 to 4, default 1. 0 skips the extra pass entirely. Start at 1; treat anything above as a walk on the wild side.

One output, IMAGE, wired to your preview, save node, or a detailer. Nothing else about your graph changes.

Installing it

It's one node from a large pack, so you either install the whole pack or skip it.

ComfyUI Manager: search for DonutNodes and install. The pack keeps updating quickly, and the author warns that an older registry build may not contain the current nodes.

Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/DonutsDelivery/ComfyUI-DonutNodes.git donutnodes
cd donutnodes
python -m pip install -r requirements.txt

Use the same interpreter that launches ComfyUI, then restart ComfyUI and hard-refresh the browser. Requirements are light for a pack this size - opencv-python-headless, scipy, matplotlib, psutil, tqdm, requests - and DonutNodes deliberately won't downgrade NumPy, reinstall PyTorch, or write pip repair lists to fix your environment for you.

The 2x-VAE path additionally needs ComfyUI-VAE-Utils installed as a separate pack, and Wan2.1_VAE_upscale2x_imageonly_real_v1.safetensors (~507 MB) dropped into ComfyUI/models/vae/. Regular VAE files need neither.

Where people get burned

Correction on a VAE that can't do it. The round trip has to return the same RGB image at the same size, and the code throws VAE damage correction requires a VAE that reconstructs one RGB image at the same size. if it doesn't. In practice: don't hand this a 2x decoder loaded through stock nodes - use Donut Load VAE, which prepares it.

Stacking strengths above 1. Each panel in the Donut workflow can run correction independently (generation, both upscales, face detail), so the same pixels can get corrected at several stages, each pass sharpening the previous pass's artifacts. Turn on one, look at it, then decide. The face detailer corrects once per refined crop even after several sampling cycles.

Expecting a miracle. Reddit's take on the Krea 2 / Qwen VAE swap is worth calibrating against: the thread that kicked it off scored well over 100 upvotes, and the top-voted pushback was that the comparison was doing two different things at once. One commenter's verdict was that the 2x output just looked like an unsharp filter. Treat damage subtraction as a cheap, no-reference sharpening pass - do it once, at strength 1, and compare before deciding it's a win.

Categorydonut/image

Inputs (4)

NameTypeDefaultDescription
samplesLATENTThe latent to be decoded.
vaeVAEThe VAE model used for decoding the latent.
vae_damage_correctionoptBOOLEANfalseSubtract estimated VAE damage after decoding, using the selected VAE's encoder and decoder for one extra round trip per image or face crop. The 2x VAE filters back to the current image size before subtraction. No original reference is needed.
vae_damage_strengthoptFLOAT1.000–40 skips correction; 1 is one-pass VAE damage subtraction. Values above 1 strengthen the same correction and can amplify artifacts. Does not add more iterations.

Outputs (1)

NameTypeDescription
IMAGEIMAGEThe decoded image.