ComfyUI Node
BFS Shot H3 Conditioning
Build MiniMax H3 conditioning for one shot with the native nodes: Reference to Video (prompt, references, length) plus the shot as an aligned guide and/or reference video.
BFS Shot H3 Conditioning
- shot
- clip
- vae
- audio_vae
- model
- vlm
- setting_mask
- positive
- latent
- model
- prompt
◄guide_modealigned guide (Add Guide)►
◄use_ref_2true►
◄first_framenone►
◄ref_image_sizematch►
◄with_audiofalse►
◄duetoff►
◄panel_positionleft►
◄panel_size1.00►
◄panel_noise0.00►
◄rope_gap0►
◄taskplanner prompt►
◄instruction►
◄setting_refoff►
CategoryBFS/shot loop
Inputs (20)
| Name | Type | Default | Description |
|---|---|---|---|
| shot | BFS_SHOT | — | |
| clip | CLIP | — | |
| vae | VAE | — | |
| guide_mode | COMBO | aligned guide (Add Guide) | aligned guide: the shot's frames sit on the generated frames (MiniMax H3 Add Guide), frame by frame. native reference video: the shot enters as <Video 1> like the reference video input. both: the two together. |
| use_ref_2 | BOOLEAN | true | Also pass the second reference. |
| first_frame | COMBO | none | Anchor the shot's own first frame as an extra aligned image guide. |
| ref_image_size | COMBO | match | 2 options: match, max |
| audio_vaeopt | VAE | — | |
| with_audioopt | BOOLEAN | false | Attach the shot's soundtrack to the guide / reference video (needs audio_vae). |
| duetopt | COMBO | off | Pin the shot's own clip in a side panel and generate in sync with it (training-free duet). canvas: panel and video share one wide grid. shifted RoPE: the video keeps its own RoPE positions and the panel sits past its edge (connect the model and use the model output). BFS Shot Join cuts the panel off by itself. In the prompt the panel has no tag: call it 'the kept footage' by its side ('the LEFT half'). |
| modelopt | MODEL | Needed for 'shifted RoPE': route the model through this node. | |
| panel_positionopt | COMBO | left | 4 options: left, right, top, bottom |
| panel_sizeopt | FLOAT | 1.000.1–1.5 | Panel size against the video (1.0 = two equal halves). |
| panel_noiseopt | FLOAT | 0.000–1 | 0 pins the panel exactly; 0.05-0.2 loosens it for bigger changes. |
| rope_gapopt | FLOAT | 00–256 | shifted RoPE only: empty RoPE steps (2x2 patches) between video and panel. Keep it small against the video width (0-2 at low resolution): a gap close to the video's width makes the model draw its own split screen. |
| taskopt | COMBO | planner prompt | Optional prompt writer for the duet. 'planner prompt' (default) uses the shot's prompt from the planner as it is. A task writes the prompt of every shot in the duet format instead: with a VLM connected it looks at the shot and its references and writes it; without one, a template for the task. Needs duet on (canvas or shifted RoPE). See the written text on the prompt output. |
| instructionopt | STRING | What changes, in seen words, for the task: 'a 1990s anime cel style', 'a sunny beach at sunset', 'an elderly woman with short grey hair'. Empty is fine for character swap (the person comes from the references). | |
| vlmopt | CLIP | Optional VLM (CLIPLoader with a Qwen3-VL text encoder) that writes the prompt when a task is chosen. Without it the task's template is used. | |
| setting_refopt | COMBO | off | TSC's trick: one more reference picture, the shot's middle frame with the person covered in TV static, so the model sees the place in full detail (the panel / guide is often small). It is the last <Picture n>; a sentence about it is added to subject_definitions (or write {setting} where you want its tag). Mask: setting_mask, else the shot's SAM 3 crop mask, else SAM 3 'person' on that frame. 'source size' uses the video's own resolution (up to 2048 short edge): sharper, slower. |
| setting_maskopt | MASK | Optional mask of the person to cover in the setting picture. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | — |
| latent | LATENT | — |
| model | MODEL | — |
| prompt | STRING | — |