Nodes/ComfyUI-Utility-Suite/Get Image Size from Latent
ComfyUI Node

Get Image Size from Latent

How Big Will This Latent Actually Decode To?

By tom-m-2020·Created about a month ago·Updated 7 days ago· 1
Get Image Size from Latent
  • latent
  • vae
  • width
  • height
  • batch_size
  • channels

A latent tensor's dimensions are not image dimensions, and the ratio between them is not something you should be reciting from memory anymore. SD1.5 and SDXL compress 8x per side; Flux compresses 8x per side but carries 16 channels instead of 4; every video VAE adds a temporal axis. So "how big is this, actually" is a question you answer by reading the VAE, not by dividing by eight.

This node answers it without decoding anything. Feed it a latent and the matching VAE and it tells you the pixel size it will decode to.

How it works

It checks that the latent is the standard LATENT dictionary with a tensor under samples, accepts both 4-D [B, C, H, W] and 5-D [B, C, T, H, W] shapes, and reads the spatial compression ratio off the VAE via spacial_compression_decode(). Then it multiplies: width is the latent's last dimension times compression, height is the second-to-last times compression, and the channels and batch size come straight off dimensions 1 and 0.

Four outputs, all integers: width, height, batch_size, channels.

That last one is more useful than it looks. If you have inherited a workflow, found a stray latent, or are debugging a shape error, channels identifies the family - 4 means an SD-era VAE, 16 means one of the modern DiT autoencoders. It is a cheap way to answer "what am I even holding" without opening the workflow metadata.

What you use the outputs for

Mostly wiring. Width and height as INTs can drive a resize node, a filename-and-path string, a tile-count calculation in a big upscaling graph, or a comparison against the canvas you expected. It is also the correct way to size something relative to a latent you did not create - an upscaled-latent workflow, a latents-from-someone-else's-pipeline situation, a video model where you want the frames' pixel size to build a matching mask.

There are no widgets and nothing to configure: latent and vae in, four numbers out. No VAE decode, no VRAM, no time cost. Pair it with Empty Latent from VAE and you can prove your geometry end to end - build a latent, ask what it decodes to, confirm you got the size you asked for.

Install

Part of ComfyUI-Utility-Suite. ComfyUI Manager → search ComfyUI-Utility-Suite → install → restart. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/tom-m-2020/ComfyUI-Utility-Suite

Restart ComfyUI. Nothing to download beyond the pack, and the declared dependency (opencv-python-headless) is not involved. The suite is written against ComfyUI's newer V3 node API - if literally none of the Utility Suite/* nodes appear in your node search, your backend is too old, not the node.

Traps

The VAE has to match the latent. The node trusts the compression ratio off whatever VAE you hand it. Give it a latent produced by one VAE and a different VAE object and you get confidently wrong numbers instead of an error. Fortunately, the two most common cases (SD/SDXL and Flux-family) share an 8x ratio, so the size usually survives even a careless swap while channels gives the mismatch away.

Integer flooring bites in the other direction. Decoding is exact, but a size like 1023 was floored when the latent was made, so the reported width is the honest answer to "what will I get" rather than "what did I ask for". If the number looks one step off from your intended resolution, this is why.

It is a reporter, not a fixer. No resampling, no padding, no correction. If the size is wrong for your pipeline, you change the pipeline - build the latent differently or feed the width/height into a resize.

Some non-IMAGE latents are rejected. The dict has to carry samples as a tensor with 4 or 5 dimensions; other latent-shaped dictionaries that some packs pass around will raise rather than guess. That is the node telling you the thing is not a diffusion latent.

CategoryUtility Suite/Latent

Inputs (2)

NameTypeDefaultDescription
latentLATENT—
vaeVAE—

Outputs (4)

NameTypeDescription
widthINT—
heightINT—
batch_sizeINT—
channelsINT—