Nodes/radiance/I2V Pipeline
ComfyUI Node

I2V Pipeline

One reference image in, video latent out — plus the widgets that only rewrite your prompt

By FXTD-Studios·Created 8 months ago·Updated about 18 hours ago· 246
I2V Pipeline
  • model
  • clip
  • vae
  • reference_image
  • character_conditioning
  • clip_vision_output
  • video_latent
  • preview_frames
  • pipeline_report
◄positive_promptsmooth camera motion, cinematic HDR, 4K►
◄negative_promptwatermark, blurry, flickering, sdr►
◄frames25►
◄seed0►
◄dit_config{}►
◄cfg_schedule_json►
◄i2v_strategyauto►
◄image_strength0.85►
◄motion_strength0.50►
◄steps0►
◄cfg0.0►
◄sampler_nameeuler►
◄schedulernormal►
◄peak_nits1000►
◄target_gamutBT.2020►
◄hdr_eotfPQ (ST.2084)►

In plain ComfyUI, image-to-video means a spider's web: load the video model, load the matching text encoder, run CLIP Vision Encode if the model wants it, hand-build the conditioning, encode the reference image, splice it into the latent, sample, decode. Doable. Annoying to do five times in an afternoon.

I2V Pipeline collapses that into one node. Model, CLIP and VAE in; a reference image, two prompts, a frame count and a seed; a video latent and a batch of decoded preview frames out.

How it works: five strategies, one of which is usually right

Models differ in how they expect the start frame to be attached, and the node knows about five ways:

  • first_frame_lock and prepend_latent - blend the encoded image into the first latent frame. Work with any video model, and are the fallback when nothing smarter applies.
  • concat_channels - writes the concatenated image latent and mask keys that Wan I2V models read. This is what ComfyUI's own WanImageToVideo produces.
  • clip_vision_inject - adds clip_vision_output to both conditionings, which Wan 2.1 I2V reads as clip_fea. Needs a CLIP Vision Encode actually connected.
  • auto (default) - uses concat_channels when the model has an image-concat input (Wan 2.1 I2V, Wan 2.2 I2V 14B), otherwise first_frame_lock.

image_strength and motion_strength only apply to the two blend-based strategies - they're ignored by concat_channels and clip_vision_inject, which is worth knowing before you spend time tuning a widget that isn't in the path. image_strength blends the image latent into the first frame and lowers denoise to 1 - 0.4 × strength; motion_strength scales the starting noise.

The widgets that don't touch pixels

peak_nits, target_gamut and hdr_eotf are prompt-text only. Selecting BT.2020 or PQ (ST.2084) appends a descriptor to your positive prompt. No pixel conversion happens. That's a legitimate design - diffusion models respond to "shot on X" phrasing - but if you expected the output to be an actual HDR video because you picked 4000 nits, that isn't what this does. peak_nits below or at 100 also skips the append entirely, for an SDR look.

Also worth knowing: cfg_schedule_json only uses its first value, as a static CFG override. There's no per-step schedule. And when a dit_config string from RadianceVideoModelInfo carries a model name, the model's defaults replace your steps, cfg, sampler_name and scheduler - set those in the config and leave the widgets alone.

Inputs and outputs that matter

Required: model (a video model - image models are rejected outright), clip, vae, reference_image (its width and height set the video size, rounded down to the VAE's spatial compression), positive_prompt, negative_prompt, frames, seed.

The frames tooltip is the one people trip on: the latent holds ceil(frames / temporal compression) frames, so you want a count of multiple of the compression, plus one - 25, 49, 81 - if you want exactly that many frames back. Asking for 24 gets you 25 or 17 depending on the model; you'll spend an evening wondering why your loop doesn't close.

Optional: steps and cfg default to 0, which means "use the model default" (the LTX-Video preset's 25 and 3.5 when nothing else is connected), plus i2v_strategy, clip_vision_output, character_conditioning (its tokens get appended to the positive prompt at weight 0.75), and the prompt-text HDR widgets above.

Outputs: video_latent (to a video decoder), preview_frames (a decoded IMAGE batch - handy, but it's a quick VAE pass, not your final render), and pipeline_report, a text diagnostic. Read the report when output looks wrong; it tells you which strategy ran and what defaults were applied.

Installing Radiance

Manager → search Radiance → install → restart → refresh. Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/fxtd-studios/radiance.git
cd radiance
python -m pip install -r requirements.txt

Windows portable: python_embeded\python.exe for the pip line. Radiance is GPL-3.0, ~147 visible nodes, and installs OpenEXR, OpenImageIO, OpenColorIO, diffusers and accelerate. No extra weights for the pipeline itself beyond the video model you already need.

Troubleshooting

Frame count is off by a few. See above - round to a multiple of the temporal compression plus one.

Image models get rejected. By design. This is a DiT video pipeline.

Reference image looks ignored. Check the strategy in pipeline_report. On a model with no image-concat input, auto falls back to blending, and low image_strength means a loose first frame. Note that Wan 2.2 TI2V 5B always takes first_frame_lock - the pack documents that as a known limitation, since ComfyUI's own path for that model is a latent-with-noise-mask approach the pipeline doesn't implement yet.

Sampler widgets changed nothing. A dit_config with a model name is overriding them, or you're on defaults of 0 and the preset is driving.

Video flickers frame to frame. Expected for per-frame processing elsewhere in the pack; on this node, keep the seed fixed and check you're not accidentally re-rolling per frame in a batch setup.

CategoryFXTD STUDIOS/Radiance/Video

Inputs (22)

NameTypeDefaultDescription
modelMODELVideo diffusion model. Image models are rejected; a model with an image-concat input (Wan I2V) enables concat_channels.
clipCLIPText encoder matching the model, used for both prompts.
vaeVAEVAE matching the model. Encodes the reference image, sets the latent shape and decodes preview_frames.
reference_imageIMAGEStart frame (display-referred sRGB, one image). Its width and height set the video size, rounded down to the VAE's spatial compression.
positive_promptSTRINGsmooth camera motion, cinematic HDR, 4KWhat should happen in the shot. ', <n> nits HDR, <gamut>' is appended when peak_nits is above 100.
negative_promptSTRINGwatermark, blurry, flickering, sdrWhat to steer away from, encoded with the same text encoder.
framesINT251–512Requested frames. The latent holds ceil(frames / temporal compression) frames; use a multiple of the compression plus 1 (e.g. 25, 49, 81) to get exactly this count back.
seedINT00–2147483648Seed for the initial noise and the sampler.
dit_configoptSTRING{}JSON from RadianceVideoModelInfo. When it carries a model_name, that model's defaults replace steps, cfg, sampler_name and scheduler.
character_conditioningoptCONDITIONINGOptional conditioning whose tokens are appended to the positive prompt at weight 0.75. Skipped if its embedding width differs.
cfg_schedule_jsonoptSTRINGJSON float array. Only the first value is used, as a static CFG override; CFG does not vary per step.
i2v_strategyoptCOMBOautoauto: concat_channels when the model has an image-concat input (Wan 2.1 I2V, Wan 2.2 I2V 14B), else first_frame_lock. clip_vision_inject needs clip_vision_output connected.
clip_vision_outputoptCLIP_VISION_OUTPUTFrom CLIP Vision Encode. Added to both conditionings as clip_vision_output, which Wan 2.1 I2V reads (clip_fea).
image_strengthoptFLOAT0.850–1first_frame_lock and prepend_latent only: blend weight of the image latent into the first latent frame; also lowers denoise to 1 - 0.4 x strength. Ignored by concat_channels and clip_vision_inject.
motion_strengthoptFLOAT0.500–1first_frame_lock and prepend_latent only: scales the starting noise to 0.2 + 0.8 x strength of its unit amplitude. Ignored by concat_channels and clip_vision_inject.
stepsoptINT00–200Sampling steps. 0 uses the model default (the LTX-Video preset's 25 when no dit_config is connected). Ignored when dit_config carries a model_name.
cfgoptFLOAT0.00–30Guidance scale. 0 uses the model default (the LTX-Video preset's 3.5 when no dit_config is connected). cfg_schedule_json overrides it.
sampler_nameoptCOMBOeulerComfyUI sampler. Ignored when dit_config carries a model_name.
scheduleroptCOMBOnormalComfyUI sigma scheduler. Ignored when dit_config carries a model_name.
peak_nitsoptCOMBO1000Adds the selected peak brightness to the prompt; 100 requests an SDR look. No pixel change.
target_gamutoptCOMBOBT.2020Adds the selected gamut descriptor to the prompt at every peak brightness. No pixel conversion.
hdr_eotfoptCOMBOPQ (ST.2084)Adds the selected transfer-function descriptor to the prompt. No pixel encoding.

Outputs (3)

NameTypeDescription
video_latentLATENT—
preview_framesIMAGE—
pipeline_reportSTRING—