Nodes/ComfyUI-BFSNodes/BFS H3 Duet Conditioning (prompt writer)
ComfyUI Node

BFS H3 Duet Conditioning (prompt writer)

MiniMax H3 duet conditioning for your own sampler. The panel (the source clip) is pinned beside the video. The prompt: yours if you write one; otherwise a connected VLM writes it in the duet format from the task and instruction, looking at the clip and the references; otherwise a draft. Outputs positive + latent (with the panel's noise mask) + model (RoPE shift applied in shifted mode) + panel_info for BFS H3 Side Panel Crop. How to refer to the side panel in the prompt: the panel has NO tag (it is not <Picture n> or <Video n>; the text encoder never sees it). Name it by its place: 'the LEFT half is the kept footage' (write {layout} to insert that sentence for the current size and side). Describe only the generated part: never describe the panel's performer, clothes or room, even to contrast them; what you leave undescribed is copied from the panel. Restate the new identity in every shot ('her face from <Picture 1>' plus two or three face, hair or outfit words). Give exact times for cuts ('[Shot 2] At 00:03.708, both halves cut together to ...'). Credits: the initial idea for this H3 node came from TSC's latent-pin duet (the source pinned beside the video with a noise mask). BFS had already used the same principle on LTX (a green side panel holding the reference) and implemented the virtual sidecar approach (reference tokens placed beside the frame in RoPE). New here: the shifted RoPE layout (the video keeps its own grid and the panel sits past its edge, with an optional gap), dynamic references, task prompts and the shot-loop node.

By alisson-anjos·Created 7 months ago·Updated about 8 hours ago· 115
BFS H3 Duet Conditioning (prompt writer)
  • clip
  • vae
  • panel
  • vlm
  • model
  • audio_vae
  • guide
  • keep_mask
  • keep_video
  • ref_images
  • positive
  • latent
  • model
  • panel_info
  • prompt
  • canvas_preview
◄taskcharacter swap►
◄instruction►
◄prompt►
◄width448►
◄height800►
◄length0►
◄positionleft►
◄size1.00►
◄panel_noise0.00►
◄rope_modecanvas►
◄fitcontain►
◄gap0►
◄holdall frames►
◄rope_gap0►
◄ref_image_sizematch►
◄vlm_max_tokens1024►
CategoryBFS/MiniMax H3

Inputs (26)

NameTypeDefaultDescription
clipCLIPThe MiniMax H3 text encoder (Qwen3-VL 32B).
vaeVAE—
panelIMAGEThe clip pinned beside the video (copied in sync), or a picture.
taskCOMBOcharacter swapWhat to do. With a VLM connected it writes the prompt for this task.
instructionSTRINGWhat changes, in seen words: 'a 1990s anime cel style', 'a sunny beach at sunset', 'an elderly woman with short grey hair', or for custom anything you want done. Empty is fine for character swap (the person comes from the pictures).
promptSTRINGYour own prompt (wins over the VLM and the draft). How to refer to the side panel in the prompt: the panel has NO tag (it is not <Picture n> or <Video n>; the text encoder never sees it). Name it by its place: 'the LEFT half is the kept footage' (write {layout} to insert that sentence for the current size and side). Describe only the generated part: never describe the panel's performer, clothes or room, even to contrast them; what you leave undescribed is copied from the panel. Restate the new identity in every shot ('her face from <Picture 1>' plus two or three face, hair or outfit words). Give exact times for cuts ('[Shot 2] At 00:03.708, both halves cut together to ...').
widthINT44832–4096—
heightINT80032–4096—
lengthINT00–3600Frames (17k+5). 0 = the panel clip's length.
positionCOMBOleft4 options: top, left, right, bottom
sizeFLOAT1.000.1–1.5—
panel_noiseFLOAT0.000–10 pins the panel exactly; 0.1-0.2 for big changes.
rope_modeCOMBOcanvascanvas (one wide grid) or shifted (connect the model and use the model output).
fitCOMBOcontain3 options: contain, cover, stretch
gapINT00–8—
holdCOMBOall frames2 options: all frames, first latent frame
rope_gapFLOAT00–256—
ref_image_sizeCOMBOmatch2 options: match, max
vlm_max_tokensINT1024128–4096—
vlmoptCLIPOptional VLM (CLIPLoader with a Qwen3-VL text encoder) that writes the prompt.
modeloptMODELNeeded for rope_mode = shifted.
audio_vaeoptVAE—
guideoptIMAGEOptional aligned latent guide in the video area.
keep_maskoptMASKInpainting inside the duet: 1 = regenerate (e.g. the person, dilated), 0 = keep. The video area starts from keep_video (or the panel clip) and only the masked region is generated, so the background and its light stay pixel-exact.
keep_videooptIMAGEThe video kept outside the mask (default: the panel clip).
ref_imagesoptCOMFY_AUTOGROW_V3—

Outputs (6)

NameTypeDescription
positiveCONDITIONING—
latentLATENT—
modelMODEL—
panel_infoBFS_H3_PANEL—
promptSTRING—
canvas_previewIMAGE—