Nodes/ComfyUI-BFSNodes/BFS H3 Duet (pinned panel, all in one)
ComfyUI Node

BFS H3 Duet (pinned panel, all in one)

MiniMax H3 duet in one node: a panel (the source clip or a picture) is pinned beside the video, the video is generated in sync with it, and only the video comes out. For any edit that keeps the source's motion and timing: character swap, style, setting, appearance, light. Optional aligned guide. How to refer to the side panel in the prompt: the panel has NO tag (it is not <Picture n> or <Video n>; the text encoder never sees it). Name it by its place: 'the LEFT half is the kept footage' (write {layout} to insert that sentence for the current size and side). Describe only the generated part: never describe the panel's performer, clothes or room, even to contrast them; what you leave undescribed is copied from the panel. Restate the new identity in every shot ('her face from <Picture 1>' plus two or three face, hair or outfit words). Give exact times for cuts ('[Shot 2] At 00:03.708, both halves cut together to ...'). Credits: the initial idea for this H3 node came from TSC's latent-pin duet (the source pinned beside the video with a noise mask). BFS had already used the same principle on LTX (a green side panel holding the reference) and implemented the virtual sidecar approach (reference tokens placed beside the frame in RoPE). New here: the shifted RoPE layout (the video keeps its own grid and the panel sits past its edge, with an optional gap), dynamic references, task prompts and the shot-loop node.

By alisson-anjos·Created 7 months ago·Updated about 8 hours ago· 115
BFS H3 Duet (pinned panel, all in one)
  • model
  • clip
  • vae
  • audio_vae
  • panel
  • guide
  • ref_images
  • ref_videos
  • ref_video_audios
  • ref_audios
  • images
  • audio
  • canvas
  • layout_text
  • latent
  • prompt
◄taskcharacter swap►
◄instruction►
◄prompt►
◄width448►
◄height800►
◄length0►
◄positionleft►
◄size1.00►
◄fitcontain►
◄gap0►
◄panel_noise0.00►
◄holdall frames►
◄rope_modecanvas►
◄rope_gap0►
◄ref_image_sizematch►
◄steps20►
◄sampler_nameeuler►
◄schedulerbeta►
◄seed42►
◄decode_canvasfalse►
◄guide_frame_idx0►
CategoryBFS/MiniMax H3

Inputs (31)

NameTypeDefaultDescription
modelMODEL—
clipCLIP—
vaeVAE—
taskCOMBOcharacter swapWrites a DRAFT prompt when the prompt box is empty: character swap (the person from the pictures), style, setting, appearance or lighting / weather (the instruction says what). The draft cannot see the video, so it is generic: a prompt written for the clip (see the example in the prompt tooltip, and the 'prompt' output for the draft) gives much better results. A text in the prompt box always wins.
instructionSTRINGWhat changes, in a few seen words: 'a 1990s anime cel style', 'a beach at sunset', 'an elderly woman with grey hair', 'night lit by pink and blue neon', or the new person's look.
promptSTRINGYour own REF2VA prompt (overrides the task). <Picture n> / <Video n> / <Audio n> are the references in order. How to refer to the side panel in the prompt: the panel has NO tag (it is not <Picture n> or <Video n>; the text encoder never sees it). Name it by its place: 'the LEFT half is the kept footage' (write {layout} to insert that sentence for the current size and side). Describe only the generated part: never describe the panel's performer, clothes or room, even to contrast them; what you leave undescribed is copied from the panel. Restate the new identity in every shot ('her face from <Picture 1>' plus two or three face, hair or outfit words). Give exact times for cuts ('[Shot 2] At 00:03.708, both halves cut together to ...'). Example (character swap, source clip pinned on the left, one face picture and one full-body picture): subject_definitions: <Subject 1> is the woman whose appearance comes from <Picture 1> and <Picture 2>: fair skin, a narrow oval face, grey-green eyes, light brown hair in a low ponytail, wearing a white long-sleeved top under a navy denim apron. summary: [reference generation] The target video is a split screen: the kept footage beside <Subject 1>, who moves in sync with it, in the same bedroom. retention_analysis: <Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - her face, ponytail, white top and navy apron are retained. The kept footage: fully_preserved - the panel is kept exactly. detailed_description: The target video is in a realistic style, as handheld vertical smartphone footage under soft daylight. {layout} [Shot 1] A medium close-up at chest height, the phone steady at eye level. <Subject 1>, her face from <Picture 1> with grey-green eyes and soft pink lips, her light brown ponytail and navy apron, talks to the camera with small nods. [Shot 2] At 00:03.708, both halves cut together to a closer framing. <Subject 1>, her face from <Picture 1>, the white sleeves and apron straps visible, raises one open hand beside her mouth and smiles. overall_soundscape: Her voice speaking to the camera in a quiet bedroom. non_diegetic_music: None.
widthINT44832–4096Generated video width (the output).
heightINT80032–4096—
lengthINT00–3600Frames at 24 fps, snapped to 17k+5. 0 = the guide's or panel clip's length.
positionCOMBOleft4 options: top, left, right, bottom
sizeFLOAT1.000.1–1.5Panel size against the video (1.0 = two equal halves).
fitCOMBOcontaincontain keeps the whole clip (smaller, with grey around) so nothing is cropped; cover fills the panel and crops; stretch distorts.
gapINT00–8Grey separator in 32 px patches (pixels, held).
panel_noiseFLOAT0.000–10 copies the panel exactly (swaps, light). 0.1-0.2 gives room for big changes (style, creatures).
holdCOMBOall frames2 options: all frames, first latent frame
rope_modeCOMBOcanvascanvas: panel and video share one wide grid (TSC). shifted (BFS): the video keeps the RoPE positions of a render without the panel and the panel sits past its edge.
rope_gapFLOAT00–256shifted only: empty RoPE steps (2x2 patches) between video and panel. Keep it small against the video width (0-2 at low resolution): a large gap makes the model draw its own split screen.
ref_image_sizeCOMBOmatch2 options: match, max
stepsINT201–200—
sampler_nameCOMBOeuler44 options: euler, euler_cfg_pp, euler_ancestral, euler_ancestral_cfg_pp, heun, heunpp2, +38
schedulerCOMBObeta9 options: simple, sgm_uniform, karras, exponential, ddim_uniform, beta, +3
seedINT420–18446744073709550000—
decode_canvasBOOLEANfalseAlso decode the whole canvas, to check the sync.
audio_vaeoptVAE—
paneloptIMAGEPinned beside the video: the source clip (copied in sync) or a picture.
guideoptIMAGEOptional aligned latent guide in the video area (for body-swap LoRAs).
guide_frame_idxoptINT0-9999–9999—
ref_imagesoptCOMFY_AUTOGROW_V3—
ref_videosoptCOMFY_AUTOGROW_V3—
ref_video_audiosoptCOMFY_AUTOGROW_V3—
ref_audiosoptCOMFY_AUTOGROW_V3—

Outputs (6)

NameTypeDescription
imagesIMAGE—
audioAUDIO—
canvasIMAGE—
layout_textSTRING—
latentLATENT—
promptSTRING—