Qwen Image Edit (qwen-image-edit-2511)
Instruct-editing 2511 without the 20B download
- image1
- image2
- image3
- image
Qwen-Image-Edit is the open editing standard of the last year. The KB's summary of why: it "took the work masks and adapters used to do" - object removal, garment swaps, relighting, re-posing, camera angles - and it won on licensing rather than raw quality, because Apache 2.0 let the community build a LoRA library on it while Flux Kontext's licence didn't. Rev 2511 is the current revision.
Running it locally means an 8B-plus-class model in your VRAM budget. This node doesn't run it at all: it calls Qwen/Qwen-Image-Edit-2511 through ModelScope's API-Inference and hands the result back as a normal IMAGE. Same instructions, none of the weights.
If you're already on the local Qwen edit stack and happy, you don't need this. If you're on a small card, or you want the edit-2511 behaviour in a graph on a laptop, this is the shortcut.
The inputs
- image1 - required, and it's the main image. In multi-image mode it's also the first reference.
- prompt - the edit instruction. Natural language: "replace the background with a rainy street at night, keep the subject lit the same." With more than one image, name each one's role or the model is guessing.
- ratio -
autoby default, which reads the aspect off image1 and keeps your framing. The other options (1:1,2:3,3:2,3:4,4:3,9:16,16:9,21:9) re-frame the output instead of preserving it. - seed - 0 lets the API choose; non-zero is sent and also busts ComfyUI's cache.
Optional:
- image2, image3 - additional reference images for multi-image editing. Three total is the node's limit, and only image1 is required.
Output: one image (IMAGE).
There's no resolution knob here, unlike the pack's Qwen Image 2.1 node. The node maps your ratio to a fixed size table capped at a 1664px long edge - that's the Qwen-Image series ceiling on ModelScope's endpoint, which the pack documents in its source because 2000 or 2048 sizes get rejected outright. So "can I get more pixels" is answered for you: no, 1664 max edge, and 1:1 means 1664×1664.
How it works
Encode, POST, poll, download. Each input image becomes an uncompressed PNG data URI (only the first frame of a batch is used per socket), the task goes out with a size derived from ratio, and the node polls ModelScope's task endpoint with progress text until it reports success.
The pack sets a 600-second polling ceiling on this node - versus 30 minutes on the 2.1 node - because 2511 jobs are shorter. If you see a timeout, it's real: the platform is backed up, and re-queuing is the answer rather than more patience. Network-level failures during create get three retries; business errors don't.
The queue blocks while this runs. Plan your graph accordingly: put cheap local nodes after it, not expensive ones you'll need to re-run if the edit disappoints.
Install and key it
cd ComfyUI/custom_nodes
git clone https://github.com/uiiiaiii/UIIIAIII_Toolkit.git
# restart ComfyUI
ComfyUI Manager users: search UIIIAIII Toolkit. Dependencies are requests>=2.28.0 and nothing more - no weights, no CUDA anything.
The token is ModelScope's, not Agnes': generate it at modelscope.cn/my/myaccesstoken, then paste it into Settings → UIIIAIII Toolkit → ① Node API. Stored plaintext in config.json inside the pack, and that file beats a MODELSCOPE_API_TOKEN environment variable for these nodes. Worth knowing that this panel is shared with the pack's Agnes nodes - one settings screen, two different providers, two different keys.
Where it goes wrong
"Input image encoding failed." A MASK or odd tensor on an image socket. All three sockets want IMAGE.
The whole frame moved when you only wanted one thing changed. This is the model's structural behaviour, not a bug in the node, and it's the reason the KB notes that the standard 2026 workflow "bolts a mask back on around it." Pixels you didn't ask about move; a chain of edits drifts faces. Keep edits few and specific, and if identity matters, composite your original back over the untouched regions downstream.
Faces drift across a series of edits. Same mechanism. Re-using an unedited reference as image1 each time beats feeding the previous output forward - error compounds when you chain.
Timeout at ten minutes. That's the node's polling ceiling, not your graph's fault. Retry - midday on a free hosted tier is exactly when everyone else is also editing.
"Prompt cannot be empty." Edit instructions are required even though the images are supplied. Describe the change in words, not just "make it better."
You wanted text rendering or a specific glyph. 2511 is strong at instruction edits; if your edit is really a re-generation with precise layout, the pack's Qwen Image 2.1 node is the text-to-image route.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | Edit instruction/prompt describing the desired change. For multi-image editing, specify each image's role/position | |
| image1 | IMAGE | Input image 1 (main, required). First reference image in multi-image editing | |
| seed | INT | 00–18446744073709550000 | Random seed. 0 means the API picks a random seed. Used to break the ComfyUI execution cache |
| ratio | COMBO | auto | Output aspect ratio. 'auto' detects the input ratio (based on image1) |
| image2opt | IMAGE | Input image 2 (optional). Second reference image in multi-image editing | |
| image3opt | IMAGE | Input image 3 (optional). Third reference image in multi-image editing |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |