Nodes/was-node-suite-comfyui/Kandinsky 6 Text Encode
ComfyUI Node Runs on cloud

Kandinsky 6 Text Encode

Two prompt boxes, because this model writes dialogue too

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
Kandinsky 6 Text Encode
  • clip
  • CONDITIONING
◄video_prompt—►
◄audio_prompt—►

Most text encoders give you one box and a bad habit of ignoring half of what you typed. Kandinsky 6 wants its prompt split in two - what is seen and what is heard - because it generates the picture and the soundtrack in one pass, and spoken lines have to land on specific frames. This node is that split.

How it works

The two boxes aren't concatenated and thrown at a text encoder. audio_prompt is folded into the video prompt wrapped in the model's own audio-caption markers, <AUDCAP>…<ENDAUDCAP>, then the whole thing is tokenized against a Qwen chat template and encoded by the pair of text encoders - Qwen 2.5 VL 7B doing the language understanding, CLIP-L contributing its pooled vector. Both halves arrive through one clip input from DualCLIPLoader set to type kandinsky5, which is the part people get wrong first.

Empty audio_prompt is legal and means exactly what it says: no sound description. The prompt still goes through, just without an audio section.

Inputs

Required, all three:

  • clip - qwen_2.5_vl_7b and clip_l, loaded by DualCLIPLoader with type kandinsky5. If the node complains, that's what it's complaining about; a Kandinsky 6 graph with the wrong loader type fails here.
  • video_prompt - what is seen. Multiline, and it's worth the ink: describe the shot, the subject, the camera move, the light, as a sentence or three. Spoken lines go inline, where they're said, wrapped as <S>Look there!<E> - that's how the model knows which frames the voice belongs on.
  • audio_prompt - what is heard apart from speech. Rain on a tin roof, a distant engine, a held string note. Leave it blank and you get the model's own guess from the visual prompt.

One output, CONDITIONING, straight into KSampler - or into Kandinsky 6 Image To Video if you're starting from a still.

Wire two of these: one for the positive, one for the negative. That's how the graph is meant to be built, and on the non-distilled checkpoints it's what the CFG 5 pass is steering with.

Installing it

It's part of the WAS Node Suite v3 pack. ComfyUI Manager, search WAS Node Suite v3, install, restart - or:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI 0.14.0+ and Python 3.10+. The pack installs nothing itself, which is a genuine improvement over the 2023 version everyone remembers from import-failure threads.

The text encoders are the download: qwen_2.5_vl_7b_fp8_scaled.safetensors (9.4 GB) and clip_l.safetensors (246 MB) into models/text_encoders. The pack mirrors them, in the right folder layout, at Hugging Face WAS/was-node-suite-weights.

Where it goes wrong

"Kandinsky 6 reads its prompt with Qwen 2.5 VL 7B and CLIP-L." You've handed it a single text encoder or the wrong loader type. DualCLIPLoader, type kandinsky5.

Your negative prompt does nothing. On a _distill_ Kandinsky 6 transformer you're sampling at CFG 1, and at CFG 1 ComfyUI never runs the unconditional pass, so the box renders and your text is thrown away. That's the standard behaviour for every guidance-distilled model, not a bug in this node. Run the non-distilled transformer at 50 steps / CFG 5 if you want negatives to bite, or restate the constraint positively.

Prompt weighting does nothing either. (rain:1.4) is CLIP-era syntax; here it just gets fed to the Qwen encoder as literal punctuation. Write plain sentences and put the important thing first - the encoder is an instruction-follower, not a bag of tags.

Long prompts get trimmed. The encoder caps the token count on the Qwen side. If a prompt seems to lose its ending, it did.

Audio that doesn't match the picture. Check the prompt split first: describe speech with <S>…<E> in video_prompt at the moment it happens, and keep audio_prompt for ambience and score. Sound described in the wrong box competes with the picture instead of accompanying it.

CategoryWAS Suite/Latent/Video

Inputs (3)

NameTypeDefaultDescription
clipCLIPqwen_2.5_vl_7b and clip_l, from DualCLIPLoader with type 'kandinsky5'.
video_promptSTRINGWhat is seen, as 'a red fox trots through deep snow, low tracking shot'. Spoken lines go where they are said, as <S>Look there!<E>.
audio_promptSTRINGWhat is heard besides speech, as 'rain on a tin roof, distant thunder'. Empty = no sound description.

Outputs (1)

NameTypeDescription
CONDITIONINGCONDITIONINGConditioning for KSampler, or for Kandinsky 6 Image To Video.