Video Latent Noise
The correctly-shaped noise your video sampler keeps asking for
- noise_latent
- shape_report
If your video sampler dies two seconds in with a shape mismatch, or hands you one decent frame followed by twenty-four black ones, the problem is usually upstream of the sampler. It's the latent.
ComfyUI's stock Empty Latent Image builds a 2D, 4-channel tensor sized for SD1.5 and SDXL. A video DiT wants something like [batch, channels, frames, H/8, W/8] - five dimensions, and the constants differ per family. LTX-Video runs 128 channels with an aggressive 32× spatial and 8× temporal VAE compression; HunyuanVideo and Wan are 16 channels at 8×4; Wan 2.2's TI2V 5B is 48 channels; Mochi-1 is 12; CogVideoX is 16. Guess wrong and the model either refuses the tensor or, worse, accepts it and denoises nonsense.
Video Latent Noise exists because that arithmetic is not something you want to do in your head, and because torch.randn with the wrong seed handling is how you get a batch of frames that are secretly the same frame.
How it actually works
The node does what the tooltips say and nothing mystical. It parses dit_config (the JSON from Video Model Info, which carries model_name, channels, spatial_compression, temporal_compression, temporal, latent_scale), merges it over the preset spec, and computes:
latent_height = height // spatial_compression, same for widthlatent_frames = ceil(frames / temporal_compression)for temporal models,1otherwise
Then it draws i.i.d. Gaussian noise at exactly that shape, seeded from a CPU generator - so the same seed and shape give you bit-identical noise. Each frame is independent, which is what every video sampler in this class expects from its starting noise.
Inputs worth touching
dit_config- wire it from Video Model Info. This is the one that matters.width/height- target pixel size, 512×512 by default. They get divided down by the compression, so you don't have to land on a multiple of 32 yourself, but staying on multiples of 8 keeps you sane.frames- pixel frames you want back. Because the latent frame count isceil(frames / temporal_compression), use a multiple of the compression plus one (25, 49, 81). That's the decoder's causal-VAE convention, and it's the difference between a 25-frame clip and a 33-frame one.seed- your reproducibility knob.noise_scale(optional) - multiplies the noise standard deviation. Leave it at 1.0. If you're reaching for this to fix a clip, something else is wrong.
Two outputs: noise_latent, which wires into Video Sampler's latent_noise, and shape_report, a string that prints the model, input resolution, resulting latent shape and the SC×TC pair. Read it once per model family. It's the fastest way to find out you've been sending 128-channel noise to a Wan model.
Install
Search Radiance in ComfyUI Manager, install, restart, refresh the browser. Manually:
cd ComfyUI/custom_nodes
git clone https://github.com/fxtd-studios/radiance.git
cd radiance
python -m pip install -r requirements.txt
Windows portable users should run that pip command with python_embeded\python.exe. No model download is needed for this node - it's pure shape math - but the pack as a whole wants OpenImageIO, OpenColorIO and the rest, and its install.py handles those.
Where people get burned
Empty dit_config. The default is "{}", and an empty config falls back to the LTX-Video 128-channel spec. Run a Wan or Hunyuan model that way and you'll get plausible-looking noise of the wrong width, then a sampler error that points nowhere near the cause. Always connect Video Model Info.
Frame counts that don't round-trip. Ask for 24 frames from a 4× temporal compression model and the latent holds 6 frames, which the VAE decodes as 21. Nothing crashes; you just get a clip that's slightly short and slightly odd at the tail.
Mixing resolutions mid-graph. This node sizes the latent from the numbers you type, not from your source video. If you load a 1080p plate and forget to change 512×512, the noise is the wrong shape for the plate you're conditioning on.
Batch size as a free win. batch_size up to 16 gives you independent samples per run and roughly linear VRAM growth on the sampler. On a 12 GB card, one at a time is not a compromise, it's the plan.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| dit_config | STRING | {} | JSON from RadianceVideoModelInfo |
| width | INT | 51264–4096 | Target frame width in pixels. Divided (rounded down) by the spec's spatial compression to size the latent. |
| height | INT | 51264–4096 | Target frame height in pixels. Divided (rounded down) by the spec's spatial compression to size the latent. |
| frames | INT | 251–512 | Target pixel frames. The latent gets ceil(frames / temporal compression) frames; use a multiple of the compression plus 1 (e.g. 25, 49) to decode back to exactly this count. |
| batch_size | INT | 11–16 | Number of independent noise samples in the batch. |
| seed | INT | 00–2147483648 | Seed for the CPU noise generator; the same seed and shape give identical noise. |
| noise_scaleopt | FLOAT | 1.000.01–4 | Multiply noise standard deviation (1.0 = unit Gaussian) |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| noise_latent | LATENT | — |
| shape_report | STRING | — |