| clip | CLIP | | The MiniMax H3 text encoder (Qwen3-VL 32B). |
| vae | VAE | | — |
| panel | IMAGE | | The clip pinned beside the video (copied in sync), or a picture. |
| task | COMBO | character swap | What to do. With a VLM connected it writes the prompt for this task. |
| instruction | STRING | | What changes, in seen words: 'a 1990s anime cel style', 'a sunny beach at sunset', 'an elderly woman with short grey hair', or for custom anything you want done. Empty is fine for character swap (the person comes from the pictures). |
| prompt | STRING | | Your own prompt (wins over the VLM and the draft). How to refer to the side panel in the prompt: the panel has NO tag (it is not <Picture n> or <Video n>; the text encoder never sees it). Name it by its place: 'the LEFT half is the kept footage' (write {layout} to insert that sentence for the current size and side). Describe only the generated part: never describe the panel's performer, clothes or room, even to contrast them; what you leave undescribed is copied from the panel. Restate the new identity in every shot ('her face from <Picture 1>' plus two or three face, hair or outfit words). Give exact times for cuts ('[Shot 2] At 00:03.708, both halves cut together to ...'). |
| width | INT | 44832–4096 | — |
| height | INT | 80032–4096 | — |
| length | INT | 00–3600 | Frames (17k+5). 0 = the panel clip's length. |
| position | COMBO | left | 4 options: top, left, right, bottom |
| size | FLOAT | 1.000.1–1.5 | — |
| panel_noise | FLOAT | 0.000–1 | 0 pins the panel exactly; 0.1-0.2 for big changes. |
| rope_mode | COMBO | canvas | canvas (one wide grid) or shifted (connect the model and use the model output). |
| fit | COMBO | contain | 3 options: contain, cover, stretch |
| gap | INT | 00–8 | — |
| hold | COMBO | all frames | 2 options: all frames, first latent frame |
| rope_gap | FLOAT | 00–256 | — |
| ref_image_size | COMBO | match | 2 options: match, max |
| vlm_max_tokens | INT | 1024128–4096 | — |
| vlmopt | CLIP | | Optional VLM (CLIPLoader with a Qwen3-VL text encoder) that writes the prompt. |
| modelopt | MODEL | | Needed for rope_mode = shifted. |
| audio_vaeopt | VAE | | — |
| guideopt | IMAGE | | Optional aligned latent guide in the video area. |
| keep_maskopt | MASK | | Inpainting inside the duet: 1 = regenerate (e.g. the person, dilated), 0 = keep. The video area starts from keep_video (or the panel clip) and only the masked region is generated, so the background and its light stay pixel-exact. |
| keep_videoopt | IMAGE | | The video kept outside the mask (default: the panel clip). |
| ref_imagesopt | COMFY_AUTOGROW_V3 | | — |