Nodes/ComfyUI-Utility-Suite/Empty Latent from VAE
ComfyUI Node

Empty Latent from VAE

The Empty Latent That Doesn't Assume You're Running SD 1.5

By tom-m-2020·Created about a month ago·Updated 7 days ago· 1
Empty Latent from VAE
  • vae
  • latent
◄width1024►
◄height1024►
◄batch_size1►

Core's Empty Latent Image node has four channels and an 8x downscale baked into it. That was true of every model when it was written and it stopped being true years ago: Flux carries 16 channels per latent pixel, Qwen-Image and Wan carry 16, SDXL is still 4 but everything after it went up, and the video VAEs add a temporal dimension on top. Hand a 4-channel latent to a 16-channel DiT and you find out at sampling time, with a shape error that mentions tensors and not models.

This node reads the numbers off the VAE you are actually using instead of guessing them.

How it works

It inspects the VAE object and pulls three things: latent_channels, latent_dim (2 or 3), and the spatial compression ratio from spacial_compression_encode(). Then it computes the latent size as width // compression by height // compression, builds an all-zeros tensor of shape [batch_size, channels, latent_height, latent_width] on ComfyUI's intermediate device and dtype, and hands it back as {"samples": tensor} - the standard LATENT dictionary every sampler expects. For a 3-D latent VAE it inserts the temporal dimension at size 1, so you get a single frame.

Nothing is encoded or decoded. No pixels, no VRAM, no VAE pass - it is pure metadata arithmetic, which is why it is instant even next to a 20 GB model.

Inputs and outputs

Four inputs: vae, plus width, height and batch_size. One output: latent.

The values are validated rather than trusted. Every dimension must be a positive integer, the VAE must expose a positive integer channel count, the latent dimensionality must be 2 or 3, and a requested size smaller than the compression ratio raises an explicit error rather than producing a zero-width latent. You get told what is wrong, which is more than most latent nodes manage.

Two practical notes on the arithmetic. The division is integer floor, so 1000 pixels at 8x compression becomes a 125-wide latent and decodes back to 1000 - but 1023 becomes 127 and decodes back to 1016. If you need an exact reconstruction, do the rounding yourself and feed multiples of 8 (or 16, if your VAE says so). And the empty latent really is zeros: the sampler is what adds noise. Do not do it by hand.

Why you'd use it instead of the core node

Three reasons come up in practice. You are running a non-SD architecture and want one node that works for all of them, so your template does not change when you swap backbones. You want the channel count to follow the VAE automatically, because that is the exact mismatch that produces the classic "why is my latent the wrong shape" evening. Or you are doing latent-space surgery - building a clean canvas to composite latents into, inpainting into a fresh region, sizing a latent to match a video VAE's expected compression - and you would rather derive the geometry from the model than hardcode it.

Pair it with Get Image Size from Latent for the round trip: build a latent, then confirm what it will decode to without paying for a decode.

Install

Part of ComfyUI-Utility-Suite. ComfyUI Manager → search ComfyUI-Utility-Suite → install → restart, or:

cd ComfyUI/custom_nodes
git clone https://github.com/tom-m-2020/ComfyUI-Utility-Suite

Restart ComfyUI. No model files and nothing to download beyond the pack itself; the declared dependency (opencv-python-headless) is unrelated to this node. The pack uses ComfyUI's V3 node API, so update the backend if the nodes do not appear at all.

Traps

Use the VAE the model was built for. The node is doing what you told it to - it reads the compression and channel count off whatever VAE you wired in. Pair a 16-channel latent with an SDXL VAE and the size maths still comes out right (both are 8x) while the channel count is wrong, which is a nastier failure than a crash because the sampler only complains later, somewhere else.

One frame is not a video. A latent from a 3-D VAE gets a temporal size of 1. That is a still image in a video model's clothing; building actual clips is the sampler's and the model's job.

Pixel-space models have no VAE. HiDream-O1-class generators encode raw pixels with no autoencoder at all, so there is no latent to build and this node has nothing to read. If your workflow for a 2026-era model has no VAE loader, that is by design and not a missing node.

CategoryUtility Suite/Latent

Inputs (4)

NameTypeDefaultDescription
vaeVAE—
widthINT10241–32768—
heightINT10241–32768—
batch_sizeINT11–4096—

Outputs (1)

NameTypeDescription
latentLATENT—