ComfyUI Node
Qwen 2.1 Ref Head Detailer
Re-renders the head of ONE reference person with Qwen Image 2.1 Edit. The head crop (with context) becomes <image1>, the person's reference photos <image2>..<imageN>; the conditioning is built per crop like TextEncodeQwenImage21 (reference latents included), sampled at the crop's size and stitched back feathered. Chain one node per reference: image, pipe and person_data pass through. Prompt placeholders: {canvas} {refs} {ref_count} {hair} {expression} {extra}.
Qwen 2.1 Ref Head Detailer
- image
- pipe
- person_data
- reference_images
- model_override
- clip_override
- vae_override
- image
- pipe
- person_data
- refined_crops
- canvas_crops
- ref_sheet
- info
◄ref_index1►
◄prompt_templateReplace the face of the person in the center of {canvas} with the face of the person shown in {refs}, so that it is unmistakably this person. Keep the facial expression, mouth shape, open or closed mouth, visible teeth, gaze and head pose of the person in {canvas} exactly as they are; only the identity changes. {refs} show the same person ({ref_count}); take the identity and face shape{hair} from them, but not their facial expression, skin tone, lighting or colors. {canvas} is the canvas: keep its lighting, skin tone, color grading, exposure and camera perspective, and keep everything else in {canvas} unchanged. {expression} {extra}►
◄extra_prompt►
◄negative_prompt►
◄seed0►
◄steps30►
◄cfg1.0►
◄sampler_nameeuler►
◄schedulersimple►
◄denoise1.00►
◄latent_modemasked►
◄ref_modeseparate►
◄max_refs4►
◄ref_cropface►
◄ref_crop_factor2.2►
◄canvas_resolution1024►
◄ref_resolution512►
◄mask_typehead►
◄mask_expand_percent0.20►
◄mask_blend_pixels32►
◄context_expand_factor2.00►
◄output_padding32►
◄delta_clamp0.35►
◄feather_directionboth►
◄take_hairstyleyes►
◄skin_tone_match1.00►
◄candidates3►
◄expression_hintauto►
CategoryFVM Tools/Face
Inputs (35)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | The picture whose head gets re-rendered. First node: the same image that went into the Person Selector as current_image. Later nodes: the image output of the previous detailer. Must keep the Person Selector's size - PERSON_DATA masks are in its pixels (a different size is resized with a warning). Batches work; alpha is dropped. | |
| pipe | FVM_QWEN_PIPE | model / clip / vae of Qwen Image 2.1 - from the Qwen 2.1 Pipe node or the pipe output of the previous detailer (passed through unchanged). | |
| person_data | PERSON_DATA | From Person Selector SAM3 (Native), or the person_data output of the previous detailer (passed through unchanged). Supplies the head / face / hair / neck masks of every matched reference person. | |
| ref_index | INT | 11–10 | Which person this node re-renders: N = the person the selector matched to its reference_N input. Must fit reference_images (reference_3 set -> 3). If the selector found nobody for this reference, the image passes through unchanged and info says why. |
| reference_images | IMAGE | Photos of this ONE person as a batch (e.g. Image Batch Multiple). Every photo is used, up to max_refs - unlike TextEncodeQwenImage21, which only takes the first image of a batch. Best: 2-4 sharp photos, different angles, face clearly visible. Mixed sizes are fine after batching; each photo is cropped to the face (ref_crop). | |
| prompt_template | STRING | Replace the face of the person in the center of {canvas} with the face of the person shown in {refs}, so that it is unmistakably this person. Keep the facial expression, mouth shape, open or closed mouth, visible teeth, gaze and head pose of the person in {canvas} exactly as they are; only the identity changes. {refs} show the same person ({ref_count}); take the identity and face shape{hair} from them, but not their facial expression, skin tone, lighting or colors. {canvas} is the canvas: keep its lighting, skin tone, color grading, exposure and camera perspective, and keep everything else in {canvas} unchanged. {expression} {extra} | Edit instruction for Qwen. Placeholders are filled per run: {canvas} -> <image1>, the head crop being edited {refs} -> <image2>, <image3> ... the reference photos {ref_count} -> 'these 4 photos' / 'this photo grid' / 'this photo' {hair} -> hairstyle wording, see take_hairstyle {expression} -> mouth description, see expression_hint {extra} -> extra_prompt The default is tuned: identity from the references, expression / pose / light / skin tone from the picture. Only change it on purpose. |
| extra_prompt | STRING | Your own addition, inserted at {extra} at the end of the prompt, e.g. 'wearing sunglasses', 'wet hair', 'light stubble'. Refer to the picture as <image1>. Leave empty for a plain identity swap. | |
| negative_prompt | STRING | What to avoid, e.g. 'blurry, passport photo, neutral expression'. Only has an effect with cfg above 1.0 - at cfg 1.0 it is ignored. | |
| seed | INT | 00–18446744073709550000 | Noise seed. With candidates > 1 the node renders seed, seed+1, ... and keeps the best. In a batch each image gets its own seeds. Results vary a lot between seeds - that is what candidates is for. |
| steps | INT | 301–100 | Sampling steps per candidate. 30 is the tested value (~10 s per candidate on an RTX 5090). The official Qwen pipeline uses 40-50; below ~20 the face gets less detailed. |
| cfg | FLOAT | 1.01–10 | Guidance. 1.0 (Qwen Image 2.1 default) = negative prompt off, fastest. 1.2-1.5 turns the negative prompt on and can sharpen the identity a little, but costs ~1.7x the time; in testing it gave no reliable gain. |
| sampler_name | COMBO | euler | Sampler. euler is what Qwen Image 2.1 is tuned and tested with. |
| scheduler | COMBO | simple | Noise schedule. simple is the Qwen Image 2.1 default. |
| denoise | FLOAT | 1.000–1 | How much of the head is re-generated. Keep 1.0. Tested: at 0.92 and below Qwen just returns the original face (no identity change at all); 0.95-0.97 flips unpredictably between original and new face. Expression is preserved via the prompt instead. |
| latent_mode | COMBO | masked | masked: only the (expanded) head mask is re-noised; everything else in the crop, including neighbouring people, is held fixed in latent space. Safe for groups. full: the whole crop is re-generated, then only the head is pasted back. Lets the head grow past the mask (longer hair), but risks shifts and seams at the edge. |
| ref_mode | COMBO | separate | How the reference photos reach Qwen. separate: each photo is its own image (<image2>, <image3>, ...) - best identity. grid: all photos tiled into one image - same token cost, the model may read it as a collage. first_only: only the first photo - what TextEncodeQwenImage21 does with a batch (for comparison). |
| max_refs | INT | 41–9 | Upper limit of reference photos used (the rest of the batch is ignored). Qwen takes at most 10 images, the head crop counts as one. More photos = more stable identity, but more tokens: time and VRAM grow with every photo (x ref_resolution²). 4 is the tested value. |
| ref_crop | COMBO | face | face: each reference photo is cut to a square around its largest face (InsightFace), so the pixel budget goes to the face instead of body and background. none: use the photos as they are (only for photos that are already head shots). Photos without a detectable face are always used uncropped. |
| ref_crop_factor | FLOAT | 2.21.2–4 | Size of the reference face crop: square side = this x the face box, shifted up a little for the hair. ~1.5: face only (hair gets cut); 2.2: head with hair and a bit of neck (tested default); 3+: head and shoulders, less face detail. Only with ref_crop = face. |
| canvas_resolution | INT | 1024256–2048 | Pixel budget of the head crop that is rendered (side of a square of the same area; the crop keeps its aspect ratio, sizes snap to multiples of 32). 1024 = ~1 MP, tested default. Higher = more face detail, but slower and more VRAM; pointless if the head is small in the source. |
| ref_resolution | INT | 512256–2048 | Pixel budget per reference photo (grid mode: per grid cell). 512 is the tested default; 768 gave no reliable gain. Each step up costs tokens for every photo - with 4 photos at 1024 the model's prefix cache may no longer fit. |
| mask_type | COMBO | head | Which region of the person is re-rendered (masks from PERSON_DATA). head: face + hair (default). head+neck: also the neck - helps when the new face's skin tone meets the neck with a visible edge. face: face only, the original hair stays. face+hair: face and hair masks combined, for when the head mask is poor. |
| mask_expand_percent | FLOAT | 0.200–0.5 | Grows the mask before rendering, as a fraction of the head size (0.20 = 20% of the larger side of the head box). Gives the new head room for a different face shape / hairstyle. Too small: identity stays stuck to the old outline. Too large: background around the head gets re-rendered too. |
| mask_blend_pixels | INT | 320–128 | Width of the soft transition where the new head is blended into the picture, in pixels of the rendered crop (converted to image pixels automatically). Larger = softer, less visible seam; 0 = hard edge. See feather_direction for which side of the mask edge the ramp sits on. |
| context_expand_factor | FLOAT | 2.001–4 | How much surrounding picture the head crop includes, relative to the head mask (2.0 = crop twice as big as the head). More context gives Qwen the scene's light and pose, but in groups also pulls neighbouring faces into the crop (they stay protected in masked mode). 1.0 = tight crop around the head. |
| output_padding | INT | 320–256 | Extra margin in image pixels added around the crop on every side, on top of context_expand_factor. |
| delta_clamp | FLOAT | 0.350.05–1 | Maximum change per pixel at the mask edge when the new head is blended back (0.35 = at most 35% of the black-to-white range). The cap relaxes to unlimited ~24 px inside the mask, so the face itself changes fully while the seam stays calm. Raise towards 1.0 if a visible ring or halo of the OLD head remains at the edge; lower if the seam shows a colour step. |
| feather_direction | COMBO | both | Where the soft blend (mask_blend_pixels) sits relative to the mask edge. both: half inside, half outside - default. inward: entirely inside the mask - nothing outside the head is touched, but the rim of the new head is only partly applied. outward: the head is applied at full strength, the transition reaches into the surroundings. |
| take_hairstyle | COMBO | yes | Fills {hair} in the prompt. yes: hairstyle and hair color come from the reference photos (tested default - helps the identity). no: the person keeps the hair from the picture; only the face changes. Combine with mask_type = face for a stricter result. |
| skin_tone_match | FLOAT | 1.000–1 | Qwen tends to copy the skin tone and white balance of the reference photos, which looks pale or cold in warm scene light. After rendering, the new face is shifted (Lab colour) to the original face's cheek colour and brightness. Only changed skin pixels are adjusted; sky, clothes and hair stay as rendered. 1.0 = full match (tested default), 0.5 = halfway, 0 = off. |
| candidates | INT | 31–8 | Number of variants rendered per person (seed, seed+1, ...). The node keeps the one whose face is most similar to the reference photos (InsightFace), with a penalty if the mouth opened/closed or the head turned compared to the original. Scores are listed in info. 1 = fast (~15 s per person) but some seeds miss the identity; 3 = tested default (~33 s); each extra variant ~10 s. The prompt is encoded only once. |
| expression_hint | COMBO | auto | Qwen likes to copy the mouth of the reference photos (neutral portraits -> the laughing person comes out with a closed mouth). This measures the mouth in the picture and writes it into the prompt at {expression} (laugh / toothy smile / closed smile / neutral). auto: only when the reference photos mostly show the other mouth state - recommended, because the hint costs a little identity. always: every time. off: never. |
| model_overrideopt | MODEL | Use a different model for THIS person only, e.g. with a character LoRA applied. The pipe output still passes the original model on to the next node. | |
| clip_overrideopt | CLIP | Use a different text encoder for this person only (e.g. LoRA-patched CLIP). The pipe output stays unchanged. | |
| vae_overrideopt | VAE | Use a different VAE for this person only. Rarely needed; the pipe output stays unchanged. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| pipe | FVM_QWEN_PIPE | — |
| person_data | PERSON_DATA | — |
| refined_crops | IMAGE | — |
| canvas_crops | IMAGE | — |
| ref_sheet | IMAGE | — |
| info | STRING | — |