Nodes/ComfyUI_Swwan/Image Prep For IC Lora (Swwan)
ComfyUI Node

Image Prep For IC Lora (Swwan)

Side-by-side the reference and the canvas

By aining2022·Created 10 months ago·Updated about 18 hours ago· 33
Image Prep For IC Lora (Swwan)
  • reference_image
  • latent_image
  • latent_mask
  • reference_mask
  • IMAGE
  • MASK
◄output_width1024►
◄output_height1024►
◄border_width0►

IC-LoRA - in-context LoRA - conditions a model by putting the reference material inside the same image the model is generating, rather than through a separate adapter or encoder path. You hand the model one wide picture that reads "here's the reference on the left, here's the empty region on the right, continue." The model can't tell where one ends and the other begins, and that's the trick.

This node builds that composite. It's the preprocessor for the layout, and it's still relevant in 2026 because in-context conditioning is where video control moved: LTX-2's trainer ships IC-LoRA control adapters as a conditioning mode of the base model, and Lightricks' Creative Lab collection is a set of IC-LoRAs stacked on LTX-2.3. The ControlNet panel's summary of the shift is blunt - video control is not the ControlNet vocabulary any more; it's VACE for Wan and IC-LoRA adapters for LTX.

Inputs and outputs

Required: reference_image (IMAGE), output_width (INT, default 1024, 1–4096), output_height (INT, default 1024, 1–4096), border_width (INT, default 0, 0–4096).

Optional: latent_image (IMAGE), latent_mask (MASK), reference_mask (MASK).

Outputs: IMAGE and MASK.

What it assembles

It takes the reference and resizes it to output_height with its aspect ratio preserved (lanczos), so the reference keeps its proportions and only its height is normalized. Then it builds the right-hand side - the part the model is supposed to generate:

  • No latent_image: the right side is a black canvas of output_width × output_height.
  • With latent_image: that image is resized to fill the output size and used as the right side instead. This is the "start from an existing frame, not from nothing" case - the shape of I2V and V2V conditioning.
  • With border_width over 0: a solid black strip of that width is inserted between them, so the reference and the canvas don't touch.

The two halves are concatenated horizontally into one wide IMAGE.

The MASK output is the parallel structure: 0 over the reference (and the border), 1 over the region the model should generate. If you supply a latent_mask, that region's mask is resized and copied through instead of being a flat 1. So the mask output is the "regenerate here" map for whatever sampler you're driving, in the same convention the pack's outpaint nodes use.

Two smaller behaviors: if reference_mask is provided it's nearest-resized to the reference and multiplied into the reference image - masking areas of the reference to black so the model doesn't condition on them. And if reference_mask is entirely black, the node logs a warning and treats it as absent, same as the outpaint nodes. If the reference and latent_image batch sizes differ, the reference is repeated to match.

The sizing intuition

The math to keep in your head: new_width = (reference_width / reference_height) * output_height. So a square reference at output_height=1024 becomes 1024 wide, and a portrait 3:4 reference becomes 768 wide. The total canvas is new_width + border_width + output_width, which means the actual width you hand the model depends on your reference's aspect ratio, not just on output_width. Set output_width to the width of the region you want generated, then check the preview before you commit to a resolution-heavy run.

Install

cd /path/to/ComfyUI/custom_nodes
git clone https://github.com/aining2022/ComfyUI_Swwan.git
cd ComfyUI_Swwan
python -m pip install -r requirements.txt

Windows portable:

.\python_embeded\python.exe -m pip install .\ComfyUI\custom_nodes\ComfyUI_Swwan\requirements.txt

Restart ComfyUI, hard-refresh the browser, search Swwan; it's filed under Swwan/Advanced/Image. It needs no model download itself - the IC-LoRA weights (and the base video or image model they attach to) come from wherever you got them, and this node just does the canvas math. It's pure tensor work on your existing torch.

One honest caveat: this node is the geometry half of an IC-LoRA workflow, and IC-LoRA conditioning formats differ between model families. Match your output_width/output_height and border conventions to the specific LoRA's own instructions - the node builds the canvas you ask for, and it will not tell you if that canvas isn't the shape your model was trained to expect.

CategorySwwan/Advanced/Image

Inputs (7)

NameTypeDefaultDescription
reference_imageIMAGE—
output_widthINT10241–4096—
output_heightINT10241–4096—
border_widthINT00–4096—
latent_imageoptIMAGE—
latent_maskoptMASK—
reference_maskoptMASK—

Outputs (2)

NameTypeDescription
IMAGEIMAGE—
MASKMASK—