SAX Qwen Image Prompt
Hand Qwen-Image up to 10 images and edit by sentence
- pipe
- images
- PIPE
- POPULATED_TEXT
Multi-image Qwen-Image editing in vanilla ComfyUI is a pile of Load Image nodes, a VAE Encode per reference, a TextEncodeQwenImage21, a latent you resize by hand, and a graph you can't read any more. SAX Qwen Image Prompt folds that into one node: drop your prompt in, optionally attach up to ten images, pass the PIPE straight to SAX KSampler. It also makes the t2i/edit switch free - no images connected, it's a text-to-image encoder; one or more, it's the editing path.
What it actually does
The important thing to know: it doesn't reimplement Qwen's encoder. It delegates to ComfyUI's own TextEncodeQwenImage21, imported lazily at execution time - so if your ComfyUI predates Qwen-Image 2.1 support you get a readable "please update ComfyUI" error instead of a node that fails to load. Deliberate, and correct; anything else would rot the moment ComfyUI changes its template.
Around that encoder the node does four jobs. It pulls model, clip and vae out of the incoming PIPE_LINE, expands wildcards and extracts <lora:name:weight> tags from your positive text (applying those to model and clip), encodes positive and negative together, and - in the edit path - swaps the empty latent for one built from your reference image. That last bit is the one that trips people up, because the latent size stops being yours to set: with images attached it comes out at image_1's aspect and size, and the loader's width/height stop mattering. The node also re-broadens the latent to the loader's batch_size, since the core encoder returns batch 1 and a batch of four would otherwise silently collapse to one. Slots sort numerically, so image_1, image_2, image_10 map in the right order no matter which port you plugged in first.
The inputs you'll actually touch
wildcard_text is your prompt or edit instruction - despite the name, wildcards are optional and only work if Impact Pack is installed. Refer to your inputs as <image1>, <image2>, and so on. Because the encoder is an LLM (Qwen3-VL), write a sentence, not tag soup: "change her dress to blue" beats a keyword list, and weighting syntax like (dress:1.4) is inert.
negative_text exists and does nothing at the settings you should be using. At cfg = 1 ComfyUI skips the unconditional pass entirely - the negative isn't weighted to zero, it simply isn't computed. If your prompt isn't landing, restate the constraint as a presence ("clean studio backdrop, sharp focus") instead of moving it down here.
resolution defaults to 1024 and controls how big your reference images get resized to - aspect kept, multiples of 32. Set it to 0 to leave each image at native size. The output follows image_1, so feeding a wildly different-sized first image is the usual cause of an edit that looks shifted.
images is an autogrow input: start at image_1 and add ports up to ten. image_1 is the edit target; the rest are composition material you reference with <image2>, <image3>. Leave them all empty for plain text-to-image at the loader's dimensions.
Two outputs. PIPE goes to SAX KSampler; POPULATED_TEXT is the expanded prompt text, handy for a text display or for logging what actually went in.
Install
ComfyUI Manager → "Install via Git URL" → https://github.com/so16tm/SAX_Bridge. Or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/so16tm/SAX_Bridge
Restart ComfyUI. Then note what you don't need: the pack ships no runtime pip dependencies, and it's meant to be driven from SAX Diffusion Loader with a split distribution - UNET in diffusion_models, text encoder in text_encoders, VAE in vae. The real prerequisite is a current ComfyUI. Impact Pack is only for wildcards; sam3 and triton only matter for the SAM3 nodes. Starting point from the author's own workflow, set on the loader: cfg = 1, euler, simple, 25–50 steps.
Where people get burned
"Pipe input is empty." Something upstream is muted or bypassed, so the node got nothing. Un-bypass the loader.
"Pipe does not contain a model" / "a CLIP model" / "a VAE". The VAE error only fires when reference images are attached - t2i is happy without one - but the other two mean the pipe never had them: a wrong loader line, or a hand-made dict.
Wildcards just don't expand. Impact Pack wasn't found; the text passes through as written with a console warning. Not a bug.
Edits drift, faces wander. That's the Qwen-Image-Edit family, not this node - the model re-emits the whole frame and the error compounds across chained edits. It's why the standard 2026 workflow bolts a mask back on for anything you need held still, and why resolution control matters more than prompt wording when people complain about offset.
One honest caveat about the pack: PIPE_LINE is SAX's own bus, and these context types don't interoperate across packs. Adopt it and the workflow is married to SAX_Bridge from loader to output. Go in with the whole line - SAX Diffusion Loader → this node → SAX KSampler → SAX Output - rather than sprinkling it into an rgthree-based graph.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| pipe | PIPE_LINE | — | |
| wildcard_text | STRING | Prompt or edit instruction. Refer to reference images as <image1>, <image2>, ... Wildcard and LoRA syntax are supported. | |
| negative_text | STRING | Negative prompt. Has no effect while cfg is 1 (the official setting). | |
| resolution | INT | 10240–4096 | Reference images are resized to about resolution x resolution pixels (aspect ratio kept, multiples of 32). 0 keeps each image at its own size. The output size follows image_1. |
| images | COMFY_AUTOGROW_V3 | Reference images for editing. image_1 is the edit target. Leave all empty for text-to-image. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| PIPE | PIPE_LINE | — |
| POPULATED_TEXT | STRING | — |