Get Latent Size & Count (Swwan)
Is this a video latent or just an image batch?
- latent
- latent
- batch_size
- channels
- frames
- width
- height
What it's for
A LATENT is a dictionary with a tensor in it, and everything you'd want to know about it lives inside that tensor where you can't see it. This node opens it up: it reports the batch size, channel count, frame count, width and height, and - the reason it's more useful than a logger - it tells you whether you're holding a video latent or an image batch, because it reports frames = 0 for the 4D case.
It's KJNodes' GetLatentSizeAndCount, re-registered by the Swwan pack, and it passes the latent through unchanged so you can leave it inline.
How it reads the tensor
The tensor's rank decides how it's interpreted:
- 5D
[B, C, T, H, W]- video latent. You get all five numbers. - 4D
[B, C, H, W]- image latent.framesis reported as0deliberately, because there is no time axis to count.
Anything that isn't rank 4 or 5 raises Invalid latent shape, which is a genuinely useful error: it means something upstream handed you something that isn't a latent at all, usually a model output or a raw tensor that got passed along by a node that doesn't typecheck.
It also prints a compact summary right on the node - B x C x T x H x W - so it works as a readout even with nothing wired to its outputs.
Inputs and outputs
Required: latent. Nothing to configure.
Outputs: latent (the pass-through), then batch_size, channels, frames, width, height - all INT, in that order.
Those width and height are latent-space numbers. A 1024×1024 image fed through a standard 8× VAE gives you width = 128, height = 128. Multiply by 8 and you're back in pixels. For video latents from Wan-family models it's worse than that, because time is compressed too - which is exactly why those workflows use frame counts of the form 4n+1 and why reading frames off a latent and comparing it to your source clip length is a mistake.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan
cd ComfyUI_Swwan
python -m pip install -r requirements.txt
Manager → ComfyUI Swwan, restart ComfyUI and hard-refresh. No models, no optional dependencies - it reads .shape and returns.
Why it's worth a slot in the graph
Three real uses. Debugging shape mismatches. When two branches of a graph won't merge, the error is usually torch complaining about broadcasting from dimensions you can't see. Put one of these on each branch and the answer is right there.
Detecting silent rank drift. An image batch and a video latent with one frame are not the same thing, and the nodes that consume them care. frames = 0 versus frames = 1 tells you which world you're in before a sampler does something odd with it.
Driving conditional logic. batch_size and frames are integers, which means you can feed them to the pack's maths and switch nodes and route a graph on them. That's how you build a workflow that handles "one image" and "a clip" without you changing widgets between runs - the structure of the latent decides.
The one habit to build: when a video workflow is behaving as if it has the wrong number of frames, check this node before you check the sampler. Rank confusion between an image batch and a video latent accounts for a solid share of "my model only generated one frame" reports, and this node is the cheapest way to confirm or rule it out.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| latent | LATENT | — |
Outputs (6)
| Name | Type | Description |
|---|---|---|
| latent | LATENT | — |
| batch_size | INT | — |
| channels | INT | — |
| frames | INT | — |
| width | INT | — |
| height | INT | — |