Qwen VL List Encode Rebalance
Many prompts, four reference images, one list of conditioning
- clip
- image1
- image2
- image3
- image4
- conditioning
The stock CLIP Text Encode takes one prompt and one encoder. The moment you want four reference images plus three variations of the instruction, you're either duplicating nodes or giving up. This node is the other option: up to four image slots, a text field that can carry a list, and a list of conditionings out the back - one per prompt, per image set.
It arrived in the pack's 26/09 update, listed as "streamlined Qwen VL encoder and many QoL nodes," which is why there's essentially no community write-up of it yet. It's plumbing in the good sense: the encoder work every Qwen-VL edit workflow needs, exposed properly.
Why a vision-language encoder needs its own node
Krea 2's text encoder is Qwen3-VL, chosen because Krea wanted a model that could see, not just read - so reference images get tokenized into the prompt as vision tokens, encoded by the vision tower, and pooled across twelve tap layers. That's a different operation from text conditioning, and it's why hand-rolling it with a generic text encode node gives you refusal-adjacent mush. The pack hardcodes the Krea 2 chat template (the "describe the key features of the input image, then explain how the instruction should alter it" system prompt) and does the image plumbing itself.
How the pass structure works
For each prompt in text, for each image set, it does the following:
- Prefixes your instruction with
Picture 1: <|vision_start|><|image_pad|><|vision_end|>per image, which is the interleaved format the encoder expects. - Scales each reference to its token tier's longest side.
- Tokenizes with the Qwen3-VL template and encodes with
encode_from_tokens_scheduled.
"Image set" is the largest image count across your four slots. Wire four images into image1 and one into image2 and you get four passes, with image2's single image reused in each - so a fixed style reference can sit alongside four different character references without you duplicating anything.
The tiers are resolutions, not token counts: low 384, normal 768, high 1024, max 1280 on the longest side, and focal 32. More pixels means more vision tokens, more time, more VRAM. focal exists because sometimes you want an image to contribute position and colour rather than detail.
image1_pos–image4_pos take "x,y" as fractions of a virtual canvas, empty meaning centre. Under the hood the node patches the Qwen3-VL encoder's position-id builder so each image claims to sit at that spot, without paying tokens for a canvas that isn't there - the corner-of-black-canvas trick, token-free. It's a no-op on encoders that aren't Qwen3-VL. That this works at all is a side effect of Qwen3-VL's training data including bounding boxes, which is why Krea 2 picks up Ideogram-style positional prompting zero-shot.
The output shape is the whole point
conditioning comes out as a list. Length is prompts × image sets. A sampler wants a single conditioning, so this node pairs directly with the pack's Conditioning Merge (List) - encode the sets, merge them down, sample. Order is prompt-major, so a list of three prompts across four sets is three groups of four.
clip must be a Qwen-VL-family encoder supplied by ComfyUI as CLIP - Krea 2's stock loader (krea2) is the target case. Pair it with anything that expects a different chat template and you're feeding the model a badly framed question; the template is hardcoded, not configurable.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/nova452/Rebalance-Pack.git
# restart ComfyUI, then Ctrl+Shift+R in the browser
Or ComfyUI Manager → Rebalance Pack. No requirements.txt, no downloads, no extra pip packages - it uses the ComfyUI torch you already have. The pack's web JS needs that hard refresh. And if you still have the pre-merge ComfyUI-ConditioningKrea2Rebalance folder installed, delete it; the classes overlap.
What goes wrong
OOM with four references. Four images at high or max, on a twelve-tap encoder, is the single most likely way to kill a run. Drop tiers before you drop references - normal for the references that only need to be recognised, low for the ones that are basically colour and pose.
Nothing comes out. All four slots empty and an empty text field gives you a list with one text-only conditioning, but an empty list upstream gives you nothing to merge and a complaint further down. Check your wiring.
A list into text plus {a|b} wildcards. The field is multiline with dynamic prompts enabled, so either route can multiply your runs. Doing both at once gets confusing fast - pick one while you're learning the node.
The list output scares downstream nodes. That's expected; anything that takes one conditioning needs the merge step, not a bigger wire.
Inputs (14)
| Name | Type | Default | Description |
|---|---|---|---|
| text | STRING | — | |
| clip | CLIP | — | |
| image1opt | IMAGE | — | |
| image1_tokensopt | COMBO | normal | 5 options: low, normal, high, max, focal |
| image2opt | IMAGE | — | |
| image2_tokensopt | COMBO | normal | 5 options: low, normal, high, max, focal |
| image3opt | IMAGE | — | |
| image3_tokensopt | COMBO | normal | 5 options: low, normal, high, max, focal |
| image4opt | IMAGE | — | |
| image4_tokensopt | COMBO | normal | 5 options: low, normal, high, max, focal |
| image1_posopt | STRING | — | |
| image2_posopt | STRING | — | |
| image3_posopt | STRING | — | |
| image4_posopt | STRING | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| conditioning | CONDITIONING | — |