ComfyUI Node

Reference Latent+

A ComfyUI node in advanced/conditioning with 56 inputs and 1 output.

By shootthesound·Created 3 months ago·Updated 2 months ago· 29
Reference Latent+
  • conditioning
  • vae
  • image_1
  • image_2
  • image_3
  • image_4
  • mask_1
  • mask_2
  • mask_3
  • mask_4
  • conditioning
max_megapixels1.0
mask_fill_modepixel_grey
image1_strength0.85
image1_facefalse
image1_hairfalse
image1_bodyfalse
image1_clothesfalse
image1_backgroundfalse
image1_ignore_areanone
image1_feather0
image1_grow0
image1_start_percent0.000
image1_end_percent1.000
image2_strength0.85
image2_facefalse
image2_hairfalse
image2_bodyfalse
image2_clothesfalse
image2_backgroundfalse
image2_ignore_areanone
image2_feather0
image2_grow0
image2_start_percent0.000
image2_end_percent1.000
image3_strength0.85
image3_facefalse
image3_hairfalse
image3_bodyfalse
image3_clothesfalse
image3_backgroundfalse
image3_ignore_areanone
image3_feather0
image3_grow0
image3_start_percent0.000
image3_end_percent1.000
image4_strength0.85
image4_facefalse
image4_hairfalse
image4_bodyfalse
image4_clothesfalse
image4_backgroundfalse
image4_ignore_areanone
image4_feather0
image4_grow0
image4_start_percent0.000
image4_end_percent1.000
Categoryadvanced/conditioning

Inputs (56)

NameTypeDefaultDescription
conditioningCONDITIONING
vaeVAE
image_1IMAGE
max_megapixelsFLOAT1.00.1–4Caps each ref image by area. Downscales only — small images pass through. Lower = faster (fewer attention tokens).
mask_fill_modeCOMBOpixel_greyHow masked-out regions are neutralised before the model attends to the ref. pixel_grey: replace masked pixels with 0.5 grey before VAE encode. In-distribution; the masked latent has soft uniform signal. latent_zero: encode full image, then zero the latent at masked positions. Constant K/V at those tokens (= layer bias) so attention distributes uniformly over them. Spiritual equivalent of 'transparent'. latent_noise: same as latent_zero but fills with Gaussian noise instead of zeros. Different attention shape — random K/V means no Q matches strongly, model weakly ignores those tokens.
image_2optIMAGE
image_3optIMAGE
image_4optIMAGE
mask_1optMASKOptional explicit mask for image 1. Overrides this image's auto-mask config below.
mask_2optMASKOptional explicit mask for image 2 (overrides auto-mask).
mask_3optMASKOptional explicit mask for image 3 (overrides auto-mask).
mask_4optMASKOptional explicit mask for image 4 (overrides auto-mask).
image1_strengthoptFLOAT0.85-5–50Per-image strength (signed). 1.0 = stock ReferenceLatent strength. 0.85 = empirical sweet spot. 0 = bypass this image entirely. Negative = anti-reference (model pulled away from this ref's features).
image1_faceoptBOOLEANfalseFace skin
image1_hairoptBOOLEANfalseHair
image1_bodyoptBOOLEANfalseBody skin
image1_clothesoptBOOLEANfalseClothes
image1_backgroundoptBOOLEANfalseBackground (tick to use the background only — handy for environment refs)
image1_ignore_areaoptCOMBOnoneExclude a side strip from the auto-mask. Useful when multiple subjects are in frame.
image1_featheroptINT00–200Soften mask edges by N pixels (0 = sharp). 10-30 is usually enough.
image1_growoptINT0-200–200Grow (>0) or shrink (<0) the mask by N pixels. Useful for tucking or extending the masked area beyond what MediaPipe detected.
image1_start_percentoptFLOAT0.0000–1Fraction of denoising at which this image's ref starts applying. 0 = active from the very first step.
image1_end_percentoptFLOAT1.0000–1Fraction of denoising at which this image's ref stops applying. 1 = active through the last step. For Flux2/Klein 9B, style/identity signal concentrates at the clean end of denoising — try start=0.6, end=1.0 to apply the ref only at the late (clean) timesteps and let the noisy/early timesteps freely follow your prompt's composition.
image2_strengthoptFLOAT0.85-5–50Per-image strength (signed). 1.0 = stock ReferenceLatent strength. 0.85 = empirical sweet spot. 0 = bypass this image entirely. Negative = anti-reference (model pulled away from this ref's features).
image2_faceoptBOOLEANfalseFace skin
image2_hairoptBOOLEANfalseHair
image2_bodyoptBOOLEANfalseBody skin
image2_clothesoptBOOLEANfalseClothes
image2_backgroundoptBOOLEANfalseBackground (tick to use the background only — handy for environment refs)
image2_ignore_areaoptCOMBOnoneExclude a side strip from the auto-mask. Useful when multiple subjects are in frame.
image2_featheroptINT00–200Soften mask edges by N pixels (0 = sharp). 10-30 is usually enough.
image2_growoptINT0-200–200Grow (>0) or shrink (<0) the mask by N pixels. Useful for tucking or extending the masked area beyond what MediaPipe detected.
image2_start_percentoptFLOAT0.0000–1Fraction of denoising at which this image's ref starts applying. 0 = active from the very first step.
image2_end_percentoptFLOAT1.0000–1Fraction of denoising at which this image's ref stops applying. 1 = active through the last step. For Flux2/Klein 9B, style/identity signal concentrates at the clean end of denoising — try start=0.6, end=1.0 to apply the ref only at the late (clean) timesteps and let the noisy/early timesteps freely follow your prompt's composition.
image3_strengthoptFLOAT0.85-5–50Per-image strength (signed). 1.0 = stock ReferenceLatent strength. 0.85 = empirical sweet spot. 0 = bypass this image entirely. Negative = anti-reference (model pulled away from this ref's features).
image3_faceoptBOOLEANfalseFace skin
image3_hairoptBOOLEANfalseHair
image3_bodyoptBOOLEANfalseBody skin
image3_clothesoptBOOLEANfalseClothes
image3_backgroundoptBOOLEANfalseBackground (tick to use the background only — handy for environment refs)
image3_ignore_areaoptCOMBOnoneExclude a side strip from the auto-mask. Useful when multiple subjects are in frame.
image3_featheroptINT00–200Soften mask edges by N pixels (0 = sharp). 10-30 is usually enough.
image3_growoptINT0-200–200Grow (>0) or shrink (<0) the mask by N pixels. Useful for tucking or extending the masked area beyond what MediaPipe detected.
image3_start_percentoptFLOAT0.0000–1Fraction of denoising at which this image's ref starts applying. 0 = active from the very first step.
image3_end_percentoptFLOAT1.0000–1Fraction of denoising at which this image's ref stops applying. 1 = active through the last step. For Flux2/Klein 9B, style/identity signal concentrates at the clean end of denoising — try start=0.6, end=1.0 to apply the ref only at the late (clean) timesteps and let the noisy/early timesteps freely follow your prompt's composition.
image4_strengthoptFLOAT0.85-5–50Per-image strength (signed). 1.0 = stock ReferenceLatent strength. 0.85 = empirical sweet spot. 0 = bypass this image entirely. Negative = anti-reference (model pulled away from this ref's features).
image4_faceoptBOOLEANfalseFace skin
image4_hairoptBOOLEANfalseHair
image4_bodyoptBOOLEANfalseBody skin
image4_clothesoptBOOLEANfalseClothes
image4_backgroundoptBOOLEANfalseBackground (tick to use the background only — handy for environment refs)
image4_ignore_areaoptCOMBOnoneExclude a side strip from the auto-mask. Useful when multiple subjects are in frame.
image4_featheroptINT00–200Soften mask edges by N pixels (0 = sharp). 10-30 is usually enough.
image4_growoptINT0-200–200Grow (>0) or shrink (<0) the mask by N pixels. Useful for tucking or extending the masked area beyond what MediaPipe detected.
image4_start_percentoptFLOAT0.0000–1Fraction of denoising at which this image's ref starts applying. 0 = active from the very first step.
image4_end_percentoptFLOAT1.0000–1Fraction of denoising at which this image's ref stops applying. 1 = active through the last step. For Flux2/Klein 9B, style/identity signal concentrates at the clean end of denoising — try start=0.6, end=1.0 to apply the ref only at the late (clean) timesteps and let the noisy/early timesteps freely follow your prompt's composition.

Outputs (1)

NameTypeDescription
conditioningCONDITIONING