Kandinsky 6 Image To Video
Start the clip on a picture, keep the sound
- positive
- negative
- vae
- start_image
- positive
- negative
- latent
You have a still and you want five seconds of moving picture with matching audio out of it. In most video models that's two or three nodes: encode the image, glue it onto the conditioning, build an empty latent, hope the sizes agree. This node does all of it - and if you leave the picture off, it quietly becomes text-to-video and you never notice the plug was there.
How it works
Two jobs in one pass. First it builds the same joint video-and-audio empty latent that Empty Kandinsky 6 Latent builds: same 4n + 1 frame grid, same rule that fps decides how much sound comes with the picture. Then, if a start_image is wired, it resizes and center-crops that picture to your canvas, encodes it with the HunyuanVideo VAE and attaches it as a reference to both the positive and negative conditioning.
That second detail is the one worth knowing: the reference is set on both sides, so the sampler steers toward it and away from it at once, the way a first-frame conditioning is supposed to work. Nothing gets pre-painted into the latent itself - the latent is still empty noise.
Inputs and outputs
Required: positive and negative conditioning (from Kandinsky 6 Text Encode, or core CLIP Text Encode if you don't want the separate audio-caption box), the vae - and note this is the HunyuanVideo VAE, the same hunyuan_video_vae_bf16.safetensors you decode with, not the audio VAE - then width, height, length, fps and batch_size, meaning the same things they mean on the empty-latent node.
Optional: start_image. An IMAGE batch; only the first picture of the batch is used. Leave it unconnected and you have text-to-video.
Three outputs, and they replace what you'd otherwise wire up yourself:
positiveandnegative- your prompts with the start image attached. Both go into KSampler.latent- the empty joint latent, into KSampler's latent input.
Then VAE Decode (Tiled) for frames, VAE Decode Audio for the waveform, and Create Video at the same fps you fed this node.
Installing it
The pack comes from ComfyUI Manager (search WAS Node Suite v3) or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
Restart, and that's the whole install - no pip step, no build, nothing downloaded on a later start. You need ComfyUI 0.14.0+ and Python 3.10+. The historical WAS Node Suite advice about reinstalling OpenCV or fighting Import Failed after a ComfyUI update belongs to the 2023 v2 pack and doesn't apply any more.
Model files: a Kandinsky 6 transformer in models/diffusion_models, qwen_2.5_vl_7b_fp8_scaled.safetensors + clip_l.safetensors in models/text_encoders, hunyuan_video_vae_bf16.safetensors and the audio VAE in models/vae. The pack's own Hugging Face mirror, WAS/was-node-suite-weights, lays them out in the right folders already.
Where it goes wrong
"encodes the start image with the HunyuanVideo VAE" - that's the node refusing to run without a VAE on the input. Wire Load VAE, not the audio VAE and not the CLIP.
The picture comes back cropped. It will: the image is resized and center-cropped to the canvas you set. If your start image is portrait and your canvas is 864 x 480, you lose the top and bottom, and no amount of prompt fixes that. Set the canvas to match the picture's rough aspect instead.
Your clip opens on the input image and then drifts. That's not this node - that's the model being five-seconds-of-training deep. Shorter clips hold the reference better. Also worth knowing: if you're running a _distill_ transformer at 10 steps / CFG 1, the negative conditioning you just wired in is computed but discarded, because there's no unconditional pass at CFG 1. Write constraints as things you want to see.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| positive | CONDITIONING | The prompt, from CLIP Text Encode or Kandinsky 6 Text Encode. | |
| negative | CONDITIONING | The negative prompt, from CLIP Text Encode or Kandinsky 6 Text Encode. | |
| vae | VAE | The HunyuanVideo VAE, the same one VAE Decode uses. | |
| width | INT | 86416–16384 | Clip width in pixels, a multiple of 16. The start image is resized and center-cropped to it. |
| height | INT | 48016–16384 | Clip height in pixels, a multiple of 16. |
| length | INT | 1211–16384 | Frames, 4n + 1: 121 = 5 seconds at 24 fps; 241 = 10 seconds. |
| fps | FLOAT | 241–120 | Frame rate the clip plays at, which sets how much sound is generated. 24 = Kandinsky 6's own; give Create Video the same. |
| batch_size | INT | 11–4096 | Clips generated at once: 1 = one clip; 2 = two clips, each from its own noise. |
| start_imageopt | IMAGE | The picture the clip opens on; the first of a batch is used. Empty = text-to-video. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | The prompt with the start image, for KSampler. |
| negative | CONDITIONING | The negative prompt with the start image, for KSampler. |
| latent | LATENT | Video and audio latents of one duration, for KSampler. |