Get Latent Range From Batch (Swwan)
Slice frames before you ever decode them
- latents
- LATENT
Pixels are the wrong place to slice
The image-side version of this node - Get Image Range From Batch - decodes a batch and then throws most of it away. If you're feeding a video pipeline, that decode is wasted work: the latent is the thing the model actually consumes, and it's what a segment-extension loop needs to keep. Slicing in latent space means you never pay VAE encode/decode on the frames you were going to discard.
This is KJNodes' GetLatentRangeFromBatch, re-registered under a Swwan ID.
Shape awareness is the whole trick
The node checks the rank of latents["samples"] and behaves differently, correctly, for each:
- A 4D latent
[B, C, H, W]- that's an ordinary image batch, and the range is taken along the batch axis. Slicing 0–3 gives you three images' worth of latent. - A 5D latent
[B, C, T, H, W]- that's a video latent from AnimateDiff-style or Wan-style pipelines, and the range is taken along the time axis instead.
That second case is the one people mean when they say "frame range". It's also the one that confuses newcomers, because a 5D latent's frame count is not the number of frames you'll get out: the VAE packs time down, which is why those workflows have those odd 4n+1 frame counts everywhere. Slicing the latent gives you latent steps; the decoder expands them.
Inputs and outputs
latents- required, the only wired input.start_index--1to4096, default0.-1means "from the end": the node rewrites it tomax(0, count - num_frames), which is how you grab the last few latent steps as an overlap for the next segment. Anything else out of range raisesStart index is out of range.num_frames--1to4096, default1. Note that-1here means all the way to the end, not one frame - the count clamps to the latent's length. This asymmetry withstart_indexhas caught out more than one person.
Output: one LATENT, contiguous, ready to feed a sampler or a decode. It's a dictionary with a samples key, same shape convention as the input - the node doesn't reshape anything, and it preserves 5D latents as 5D.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan
cd ComfyUI_Swwan
python -m pip install -r requirements.txt
Manager → ComfyUI Swwan, restart, hard-refresh. No models, no optional dependencies - this is indexing a tensor.
Where it earns its place
Segment extension. You've generated a clip, you want the next one to continue from it, and the standard pattern is to hand the model a handful of trailing frames as context while starting a fresh generation - the sliding-context idea that made arbitrarily long video feasible without arbitrary VRAM. Doing that handoff in latent space means the context frames never round-trip through the VAE.
Pair it with Get Latent Size & Count when you're unsure what you've got. That node tells you the rank-derived batch/frames/channels, and frames = 0 is the giveaway that you're holding an image batch and slicing the batch axis, not a video latent and slicing time. Two nodes, ten seconds, and you avoid the classic half-hour of "why is my slice doing nothing" when the tensor was 4D all along.
And if you want to insert latents back rather than take a range out, the pack ships the counterpart nodes; this one is read-only.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| latents | LATENT | — | |
| start_index | INT | 0-1–4096 | — |
| num_frames | INT | 1-1–4096 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| LATENT | LATENT | — |