Empty Kandinsky 6 Latent
An empty video latent with the soundtrack already in it
- LATENT
Kandinsky 6 doesn't generate a silent clip you bolt sound onto afterwards. It generates the picture and the sound together, in the same sampling pass, which is the one thing this model has that most of the open video stack doesn't. So the "empty latent" you start from isn't a stack of blank video frames any more - it's a joint video-and-audio latent, both halves the same duration, and someone has to build it. That's this node.
If you've built a Wan or LTX graph, it's EmptyLatentVideo with the audio half already wired in. If you've built nothing yet, it's your first node: everything else in a Kandinsky 6 graph hangs off it.
How it works
The node answers one zeroed latent holding video latents for the frame count you ask for, plus an audio latent covering the same stretch of time. The frame count snaps to the model's temporal grid - 4n + 1, where 121 frames is five seconds at 24 fps and 241 is ten - which is why the length widget steps in fours. The audio half is sized from length and fps together, so the sound isn't a separate decision you make later; it's derived from the clip you asked for.
One output, LATENT, and it goes straight into KSampler. The prompt side comes from Kandinsky 6 Text Encode (which takes what's seen and what's heard in two separate boxes) or core CLIP Text Encode if you only care about the picture; the model comes from Load Diffusion Model. After sampling, you decode twice: VAE Decode (Tiled) for the frames, VAE Decode Audio with the audio VAE for the waveform.
The inputs that matter
- width / height - multiples of 16. 864 x 480 is what Kandinsky 6 was trained at, and it's the default for a reason. Both of these can go much higher; that doesn't mean the model will enjoy it.
- length - frames on the 4n + 1 grid. 121 = 5 seconds at 24 fps. Kandinsky was trained on five-second clips, so 121 is where the model is happiest and 241 is already extrapolation.
- fps - 24 is Kandinsky's own rate, and it sets how much sound gets generated. Give Create Video the same number at the end, or you'll have a soundtrack and a picture that disagree about how long they are.
- batch_size - 1 for a clip, 2 for two clips from their own noise each.
Installing it
The node ships in WAS Node Suite v3, so install the pack in ComfyUI Manager (search WAS Node Suite v3) or clone it by hand and restart:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0 or newer and Python 3.10+. It installs nothing. That's a change from the pack's 2023-era v2, which dragged in OpenCV and friends and produced a lot of Import Failed threads when ComfyUI updated underneath it - v3 has no default requirements at all.
The weights are the real install. Load Diffusion Model wants a Kandinsky 6 transformer in models/diffusion_models (Lite is 3.8–4.2 GB in the pack's quantised forms, Pro 29–35 GB), DualCLIPLoader with type kandinsky5 wants qwen_2.5_vl_7b_fp8_scaled.safetensors and clip_l.safetensors in models/text_encoders, and Load VAE wants hunyuan_video_vae_bf16.safetensors. The pack publishes the whole set as a ready-laid-out folder at Hugging Face WAS/was-node-suite-weights under models/kandinsky6/.
Where it goes wrong
Could not detect model type on load means the file isn't a Kandinsky 6 transformer - a Lite file fed to a Pro workflow, or some other architecture entirely.
Pro runs out of memory. The _w6a8 checkpoint is the smallest of the Pro quants and does run on a 24 GB card; the int8 versions want more headroom.
ComfyUI dies with Fatal Python error: Aborted in VAE Decode after a Pro run. Decode with VAE Decode (Tiled) instead: tile_size 512, overlap 64, temporal_size 64, temporal_overlap 8.
Distilled checkpoint, dead negative prompt. The distill checkpoints sample at 10 steps with CFG 1, and at CFG 1 there is no unconditional pass for a negative to steer - ComfyUI skips it silently. Not specific to this pack, but it catches people every time. 50 steps at CFG 5, euler, simple is the non-distilled recipe.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| width | INT | 86416–16384 | Clip width in pixels, a multiple of 16. 864 x 480 = the size Kandinsky 6 was trained at. |
| height | INT | 48016–16384 | Clip height in pixels, a multiple of 16. |
| length | INT | 1211–16384 | Frames, 4n + 1: 121 = 5 seconds at 24 fps; 241 = 10 seconds. |
| fps | FLOAT | 241–120 | Frame rate the clip plays at, which sets how much sound is generated. 24 = Kandinsky 6's own; give Create Video the same. |
| batch_size | INT | 11–4096 | Clips generated at once: 1 = one clip; 2 = two clips, each from its own noise. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| LATENT | LATENT | Video and audio latents of one duration, for KSampler. |