Nodes/UIIIAIII Toolkit/Qwen Image 2.1 (qwen-image-2.1)
ComfyUI Node

Qwen Image 2.1 (qwen-image-2.1)

One node that's both a generator and an editor

By uiiiaiii·Created 23 days ago·Updated 10 days ago· 0
Qwen Image 2.1 (qwen-image-2.1)
  • image1
  • image2
  • image3
  • image4
  • image
◄prompt►
◄seed0►
◄ratioauto►
◄resolution1K►

Qwen-Image 2.1 is Alibaba's API-only successor to the 20B family everyone's been running locally - the KB's read on 2.0 was "cut 20B to 7B, unified generation with editing, launched API-only with no open weights yet." No weights means no local option at any VRAM. So this node is the door: ModelScope's API-Inference, from your graph, no download.

The neat part is the unified design. Leave the image sockets empty and it's a text-to-image model. Connect one and it's an instruction editor. Connect up to four and it's a multi-reference composition tool - "put the jacket from image 2 on the person from image 1, keep the pose from image 3." That's a workflow that used to need an IP-Adapter stack and a ControlNet, condensed into one node and a sentence.

The inputs

  • prompt - either a generation prompt or an edit instruction. With several reference images, say what each one is for. The author's tooltip is explicit about this, and it's the single biggest quality lever on this node: "the woman in image 1 wearing the coat from image 2" beats a prompt that assumes the model guesses roles.
  • seed - 0 means "let the API pick." Non-zero gets sent, so you can lock a look. It also doubles as the ComfyUI cache-buster.
  • ratio - auto (default), 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9, 21:9. auto reads the ratio off your first connected image; with none connected it falls back to 1:1.
  • resolution - 1K or 2K. Read the tooltip before you assume bigger is better: 1K is "fast and stable and recommended for multi-image editing"; 2K is the model's native quality, "much slower and may fail with 4 reference images." That's the author reporting a platform limit, and it matches the source: even the 2K sizes are clamped to a 2048px max edge rather than the model's headline dimensions, because ModelScope's endpoint won't accept wider.
  • image1 … image4 - optional. All empty means text-to-image.

Output is a single image (IMAGE). Everything downstream still works normally.

How it works

ModelScope API-Inference is an async job API, not a squeeze-and-return endpoint. The node encodes each reference at full resolution as an uncompressed PNG data URI, POSTs a task, then polls until it succeeds, then downloads the image. The pack retries the create call up to three times on network failures (business errors fail fast), closes connections to dodge the SSL EOF flakiness people hit on that endpoint, and gives 2.1 a 30-minute polling ceiling because - per its own comment - 600 seconds "causes tasks to time out before they finish."

While it runs you get progress text in the node, from encoding through to download. Note the queue is blocked the entire time: nothing downstream executes until this node returns.

There's a locally enforced ceiling of four reference images. The model itself reportedly handles ten; the platform's practical limit is four, and the node raises a clear error rather than letting the API fail obscurely.

Install and key it

cd ComfyUI/custom_nodes
git clone https://github.com/uiiiaiii/UIIIAIII_Toolkit.git
# restart ComfyUI

Or install UIIIAIII Toolkit from ComfyUI Manager. One dependency (requests>=2.28.0), no model files, no GPU requirement.

The credential is a ModelScope token, not an Agnes key: grab one at modelscope.cn/my/myaccesstoken and paste it into Settings → UIIIAIII Toolkit → ① Node API. It's stored in config.json in the pack folder, plaintext, and that file outranks the MODELSCOPE_API_TOKEN environment variable for these nodes. ModelScope's hosting and online testing are free and underwritten by Alibaba Cloud - this isn't Comfy's metered Partner Node credits - but you're still sharing a public inference tier, which is where the slowness comes from.

Common problems

"Image generation task failed" after a long wait. The classic 2K-with-four-references case. Drop to 1K; the author's own note records 4-image 1K runs finishing in about twenty seconds versus 2K runs failing outright.

"Too many input images." Five or more wired in. The node refuses locally rather than wasting a job.

"Prompt cannot be empty." Yes, even for a pure edit - this node always wants an instruction.

A reference image is being ignored. Uncompressed PNG data URIs of large inputs make a big request body, and multi-image prompts are underspecified more often than they're refused. Name roles and positions explicitly, and try 1K while iterating.

Nothing changes when you re-roll. With seed: 0 the API picks a new one each call, so a cached result is really "you changed nothing." Bump the seed to force a re-run.

Output shape is off. auto snaps your first image's aspect to the nearest supported ratio, so a 1900×1000 input becomes 16:9. Set the ratio explicitly when framing matters.

CategoryUIIIAIII Toolkit/Modelscope

Inputs (8)

NameTypeDefaultDescription
promptSTRINGGeneration prompt, or edit instruction describing the desired change. With multiple images, specify each image's role/position
seedINT00–18446744073709550000Random seed. 0 means the API picks a random seed. Used to break the ComfyUI execution cache
ratioCOMBOautoOutput aspect ratio. 'auto' detects the ratio of the first connected image, and falls back to 1:1 for text-to-image
resolutionCOMBO1KOutput resolution. 1K (1024px max edge) is fast and stable and recommended for multi-image editing; 2K (2048px max edge) is the native quality but much slower and may fail with 4 reference images
image1optIMAGEReference image 1 (optional). Connect images for editing; leave all empty for text-to-image
image2optIMAGEReference image 2 (optional). Up to 4 reference images per run
image3optIMAGEReference image 3 (optional). Up to 4 reference images per run
image4optIMAGEReference image 4 (optional). Up to 4 reference images per run

Outputs (1)

NameTypeDescription
imageIMAGE—