🔡 SDVN CLIP Text Encode Qwen Image 2.1
Your Reference Photos Go In Through the Text Encoder
- clip
- images
- vae
- positive
- latent
- prompt
If you learned ComfyUI on SDXL, "text encode" means one thing: a box you type into. This node is not that. It has an image input, and connecting one doesn't quietly convert your picture into latents and hope - the pictures get handed to the text encoder itself, as part of the prompt.
That's possible because Qwen Image 2.1 doesn't encode text with CLIP. It uses a vision-language model - an 8B Qwen3-VL-class file, not the Qwen2.5-VL 7B the first Qwen Image line loads. When the encoder reads images and text in one pass, your reference photos can steer the generation through the same channel as your words. SDVN's node is the plumbing for that, plus the pack's usual prompt conveniences.
What actually happens when you hit run
The order of operations matters more than any single widget. Your positive prompt gets the chosen style appended to it. Then the pack's dynamic-prompt generator expands wildcards and picks one variant using seed. Then translate fires a live Google Translate call if you set it to anything but None. Only then does the finished string get tokenized - with the reference images riding along - and encoded into the conditioning.
With a VAE connected, each reference image is also encoded and attached to the conditioning as a reference latent. That's the same dual-encoding split the Qwen Edit line uses: the VL encoder carries the visual meaning, the latent carries the visual appearance. Note the switch in the code - if there are reference latents, the node drops the vision tokens. You're choosing one path, not stacking both.
The inputs worth touching
- clip - the Qwen Image 2.1 encoder, loaded with ComfyUI's CLIPLoader and its Qwen Image type (the pack's 📥 CLIP Download node reads that list off ComfyUI, so whatever your build calls it is what you pick).
- positive - an instruction. On an LLM-encoded model,
(face:1.4)is not ignored, it's fed to the encoder as literal punctuation, andmasterpiece, best qualityburns tokens that your actual description needs. - images - an autogrow socket. Connect one and the next appears, up to
image_1…image_16. Order counts: the first image decides the size of thelatentoutput. - mode / resolution -
max resolutionkeeps total pixel area equal toresolution²with the aspect ratio intact (mirroring ComfyUI's original Qwen Image 2.1 node);max sizecaps the long edge Kontext-style;resolution = 0leaves the source alone. - vae (optional) - the switch described above.
- style / translate / seed - the pack's sugar.
Outputs are positive (CONDITIONING, into your sampler), latent (a zeros LATENT sized to the first reference - this is your empty latent, not an img2img start), and prompt (the final translated text as a STRING, handy for dumping into a text save node or a filename).
Installing it
It's on the Comfy Registry, so ComfyUI Manager → search "SDVN" and install. Manually:
cd <ComfyUI>/custom_nodes
git clone https://github.com/StableDiffusionVN/SDVN_Comfy_node
cd .. # back to the ComfyUI root
pip install -r custom_nodes/SDVN_Comfy_node/requirements.txt
Be warned, that requirements file is a whole toolkit's worth of dependencies for one text box: ultralytics==8.3.107 pinned, opencv-python, scipy, pandas, gallery-dl, instaloader, googletrans-py, google-genai, openai, dynamicprompts. The README also asks Windows and macOS users to install the aria2c binary separately for the pack's download nodes.
Then the models, which is the real cost. Comfy-Org publishes the 2.1 trio under Qwen-Image-2.1: a diffusion model (bf16 or int8_convrot), a text encoder (Qwen3-VL 8B in bf16 / int8_convrot / w4a8, plus two 9B builds tagged t2i and i2i), and qwen_image_2.1_vae_bf16. The encoder is 8-9B on its own, so it's a second VRAM budget - on a 30-series card grab the int8_convrot build, which ComfyUI has supported natively since v0.27 and which runs faster than fp8. The pack's 📥 AnyDownload List node already carries these files if you'd rather not hunt.
Where people get burned
- Seed 0 looks like broken wildcards. The dynamic-prompt picker does
seed % combinations, so a seed of 0 always resolves to the first combination. Move the seed and your{red|blue}starts behaving. - Silent translation.
translatehits Google's endpoint at run time: no network, no run, and even when it works it rewrites your prompt. The node prints the translation - read it once. - Transparent PNGs go black. Before the image reaches the encoder, alpha is composited against black. Cutouts become black-background images. Flatten onto white first if that matters.
- Don't mix latents. The
latentoutput is 64-channel, spaced 16× - Qwen Image 2.1's VAE is not the 8×, 16-channel one the v1/Edit line uses. Feeding the wrong one to a sampler fails loudly, which is the good kind of failure. - There's no negative field, and you don't need one. Qwen Image 2.1 runs at the low-CFG guidance-distilled settings where the unconditional pass isn't computed at all, so anything you'd put there would be discarded. State constraints as presence: "plain seamless backdrop, sharp focus".
- Update ComfyUI before filing a bug. This node is written against the newer V3 authoring API (
io.ComfyNode, autogrow inputs), so an older build can fail to register it entirely.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| clip | CLIP | CLIP Qwen Image 2.1 dùng để mã hóa prompt. | |
| positive | STRING | Prompt tích cực mô tả nội dung bạn muốn sinh ra. | |
| style | COMBO | None | Chọn style mẫu có sẵn để thêm vào prompt. |
| translate | COMBO | Ngôn ngữ dịch prompt. | |
| mode | COMBO | max resolution | Max size giới hạn cạnh dài như KontextReference; max resolution giữ tổng diện tích như node Qwen Image 2.1 gốc. |
| resolution | INT | 10240–4096 | Kích thước ảnh tham chiếu theo chế độ đã chọn. 0 giữ kích thước ảnh gốc. |
| seed | INT | 00–18446744073709550000 | Seed ngẫu nhiên cho prompt. |
| images | COMFY_AUTOGROW_V3 | Ảnh tham chiếu cho Qwen Image 2.1. Nối ảnh để hiện thêm đầu vào ảnh. | |
| vaeopt | VAE | VAE dùng để mã hóa ảnh tham chiếu thành reference latent. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | — |
| latent | LATENT | Latent rỗng có kích thước khớp với ảnh tham chiếu đầu tiên. |
| prompt | STRING | — |