| model | MODEL | | — |
| clip | CLIP | | — |
| vae | VAE | | — |
| task | COMBO | character swap | Writes a DRAFT prompt when the prompt box is empty: character swap (the person from the pictures), style, setting, appearance or lighting / weather (the instruction says what). The draft cannot see the video, so it is generic: a prompt written for the clip (see the example in the prompt tooltip, and the 'prompt' output for the draft) gives much better results. A text in the prompt box always wins. |
| instruction | STRING | | What changes, in a few seen words: 'a 1990s anime cel style', 'a beach at sunset', 'an elderly woman with grey hair', 'night lit by pink and blue neon', or the new person's look. |
| prompt | STRING | | Your own REF2VA prompt (overrides the task). <Picture n> / <Video n> / <Audio n> are the references in order. How to refer to the side panel in the prompt: the panel has NO tag (it is not <Picture n> or <Video n>; the text encoder never sees it). Name it by its place: 'the LEFT half is the kept footage' (write {layout} to insert that sentence for the current size and side). Describe only the generated part: never describe the panel's performer, clothes or room, even to contrast them; what you leave undescribed is copied from the panel. Restate the new identity in every shot ('her face from <Picture 1>' plus two or three face, hair or outfit words). Give exact times for cuts ('[Shot 2] At 00:03.708, both halves cut together to ...').
Example (character swap, source clip pinned on the left, one face picture and one full-body picture):
subject_definitions:
<Subject 1> is the woman whose appearance comes from <Picture 1> and <Picture 2>: fair skin, a narrow oval face, grey-green eyes, light brown hair in a low ponytail, wearing a white long-sleeved top under a navy denim apron.
summary:
[reference generation] The target video is a split screen: the kept footage beside <Subject 1>, who moves in sync with it, in the same bedroom.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - her face, ponytail, white top and navy apron are retained.
The kept footage: fully_preserved - the panel is kept exactly.
detailed_description:
The target video is in a realistic style, as handheld vertical smartphone footage under soft daylight. {layout}
[Shot 1] A medium close-up at chest height, the phone steady at eye level. <Subject 1>, her face from <Picture 1> with grey-green eyes and soft pink lips, her light brown ponytail and navy apron, talks to the camera with small nods.
[Shot 2] At 00:03.708, both halves cut together to a closer framing. <Subject 1>, her face from <Picture 1>, the white sleeves and apron straps visible, raises one open hand beside her mouth and smiles.
overall_soundscape:
Her voice speaking to the camera in a quiet bedroom.
non_diegetic_music:
None. |
| width | INT | 44832–4096 | Generated video width (the output). |
| height | INT | 80032–4096 | — |
| length | INT | 00–3600 | Frames at 24 fps, snapped to 17k+5. 0 = the guide's or panel clip's length. |
| position | COMBO | left | 4 options: top, left, right, bottom |
| size | FLOAT | 1.000.1–1.5 | Panel size against the video (1.0 = two equal halves). |
| fit | COMBO | contain | contain keeps the whole clip (smaller, with grey around) so nothing is cropped; cover fills the panel and crops; stretch distorts. |
| gap | INT | 00–8 | Grey separator in 32 px patches (pixels, held). |
| panel_noise | FLOAT | 0.000–1 | 0 copies the panel exactly (swaps, light). 0.1-0.2 gives room for big changes (style, creatures). |
| hold | COMBO | all frames | 2 options: all frames, first latent frame |
| rope_mode | COMBO | canvas | canvas: panel and video share one wide grid (TSC). shifted (BFS): the video keeps the RoPE positions of a render without the panel and the panel sits past its edge. |
| rope_gap | FLOAT | 00–256 | shifted only: empty RoPE steps (2x2 patches) between video and panel. Keep it small against the video width (0-2 at low resolution): a large gap makes the model draw its own split screen. |
| ref_image_size | COMBO | match | 2 options: match, max |
| steps | INT | 201–200 | — |
| sampler_name | COMBO | euler | 44 options: euler, euler_cfg_pp, euler_ancestral, euler_ancestral_cfg_pp, heun, heunpp2, +38 |
| scheduler | COMBO | beta | 9 options: simple, sgm_uniform, karras, exponential, ddim_uniform, beta, +3 |
| seed | INT | 420–18446744073709550000 | — |
| decode_canvas | BOOLEAN | false | Also decode the whole canvas, to check the sync. |
| audio_vaeopt | VAE | | — |
| panelopt | IMAGE | | Pinned beside the video: the source clip (copied in sync) or a picture. |
| guideopt | IMAGE | | Optional aligned latent guide in the video area (for body-swap LoRAs). |
| guide_frame_idxopt | INT | 0-9999–9999 | — |
| ref_imagesopt | COMFY_AUTOGROW_V3 | | — |
| ref_videosopt | COMFY_AUTOGROW_V3 | | — |
| ref_video_audiosopt | COMFY_AUTOGROW_V3 | | — |
| ref_audiosopt | COMFY_AUTOGROW_V3 | | — |