Qwen Edit Difference Mask + Composite
Keep only what Qwen actually changed — and keep the rest of the frame untouched
- source
- edited
- limit_mask
- composite
- mask
- preview
- difference
This is the node you reach for when you like what Qwen edited but not what it did to the other 95% of the frame. Point it at the original and the edit, and it figures out which pixels actually changed, cleans that up into a usable mask, and pastes only those onto your original.
It's the difference between "the tattoo is right, but the face re-rendered" and a real deliverable. It's also the one node in this trio you can run blind on a fresh workflow and get something sane out of.
How it works
Mechanically it's the same engine as the Align Composite node plus a difference pass. If align_first is on (default), it estimates the translation drift between the pair and warps the edit back before comparing - because as the author's own tooltip puts it, without that, a 3 px shift lights up every edge in the image. Then it blurs both images by pre_blur_px (comparison only - compositing still uses the sharp edit), computes a difference, thresholds it, and cleans the mask up.
diff_mode decides what "different" means: color (max RGB) fires on any single-channel change, luminance on brightness, chroma on colour relative to brightness. If both images are RGBA it also picks up alpha changes.
Cleanup runs in a sensible order: drop connected blobs under min_region_px, optionally fill_holes for gaps inside regions, then grow_px (negative shrinks), then feather_px for the blend. limit_mask is the veto - anywhere it's black, the source always wins, even after feathering.
Where the final mask is exactly zero, the output is the source pixel, unchanged. Not "blended back", not "VAE-decoded back". That's the whole point.
The inputs that matter
source/edited- original, then Qwen's output. Order matters.threshold(0.08) - the one knob you'll actually tune. Lower = bigger mask = more of the edit comes through. Start here and go down until the mask covers the object.align_first(true) - leave it on for Qwen. Turn it off if you already ran the edit through Align Composite'saligned_editoutput.pre_blur_px(2) - your anti-noise. VAE round-trips leave low-level fuzz everywhere; blurring the comparison kills it without softening your composite.min_region_px(64) - the escape hatch for speckle. Set 0 if you're chasing genuinely tiny edits.limit_mask- for removals and edits near other busy content. Removals specifically: the object is replaced by background from the edited image.
Outputs: composite, mask, preview (a red overlay on the source - watch this one live), and difference, a grayscale view scaled around your threshold. Wire composite to Save Image and mask wherever you need it downstream.
Install
ComfyUI Manager → search WepeNerd, or:
cd ComfyUI/custom_nodes
git clone https://github.com/WepeNerd/ComfyUI-WepeNerd
python -m pip install -r ComfyUI-WepeNerd/requirements.txt
Restart, hard-refresh, look under WepeNerd/Qwen Edit Align. No model downloads, no GPU wheels. Don't skip the pip line - the pack's modules import SciPy and PyAV at import time, so a bare clone shows you zero WepeNerd nodes and a ModuleNotFoundError in the console rather than a useful error from this node.
Where people get burned
Expecting semantics. This detects visual difference, not objects. It's in the node's own description and it bites both ways: a lighting change Qwen made by accident gets selected, while the parts of your new object that blend into the background can be missed. Preview the mask, then tune.
Threshold too low. Half the frame comes through because the whole image was re-encoded. If your composite looks "the same but worse", the mask is leaking - raise threshold, keep pre_blur_px at 2, don't set min_region_px to 0 "just in case".
Affine drift. align_first only corrects translation. If the edit came back slightly scaled or rotated, run it through Qwen Edit Align Composite in affine mode first, then feed that node's aligned_edit plus the original here with align_first off.
Qwen's own misalignment. Even a correctly-sized input comes back slightly offset or soft - that's a known, widely-hit property of the 2509 line, not something you configured wrong. Which is also why "composite the original back" became standard practice, and why this node exists.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| source | IMAGE | — | |
| edited | IMAGE | — | |
| threshold | FLOAT | 0.0800–1 | How different a pixel must be to count as edited. Lower = bigger mask. |
| align_first | BOOLEAN | true | Undo the Qwen drift before diffing. Without it a 3 px shift lights up every edge. |
| diff_mode | COMBO | 3 options: color (max RGB), luminance, chroma | |
| pre_blur_px | INT | 20–32 | Blur both images before diffing to ignore VAE noise / fine texture changes. |
| min_region_px | INT | 640–1000000 | Drop isolated blobs smaller than this many pixels. |
| fill_holes | BOOLEAN | true | — |
| grow_px | INT | 2-128–256 | Expand (or shrink, if negative) the mask. |
| feather_px | INT | 20–256 | Soft edge for the blend. |
| limit_maskopt | MASK | Optional: only allow changes inside this region. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| composite | IMAGE | — |
| mask | MASK | — |
| preview | IMAGE | — |
| difference | IMAGE | — |