Nodes/SDVN_Comfy_node/🔡 SDVN CLIP Text Encode Qwen Image 2.1
ComfyUI Node

🔡 SDVN CLIP Text Encode Qwen Image 2.1

Your Reference Photos Go In Through the Text Encoder

By StableDiffusionVN·Created 2 years ago·Updated 4 days ago· 126
🔡 SDVN CLIP Text Encode Qwen Image 2.1
  • clip
  • images
  • vae
  • positive
  • latent
  • prompt
◄positive—►
◄styleNone►
◄translate▾►
◄modemax resolution►
◄resolution1024►
◄seed0►

If you learned ComfyUI on SDXL, "text encode" means one thing: a box you type into. This node is not that. It has an image input, and connecting one doesn't quietly convert your picture into latents and hope - the pictures get handed to the text encoder itself, as part of the prompt.

That's possible because Qwen Image 2.1 doesn't encode text with CLIP. It uses a vision-language model - an 8B Qwen3-VL-class file, not the Qwen2.5-VL 7B the first Qwen Image line loads. When the encoder reads images and text in one pass, your reference photos can steer the generation through the same channel as your words. SDVN's node is the plumbing for that, plus the pack's usual prompt conveniences.

What actually happens when you hit run

The order of operations matters more than any single widget. Your positive prompt gets the chosen style appended to it. Then the pack's dynamic-prompt generator expands wildcards and picks one variant using seed. Then translate fires a live Google Translate call if you set it to anything but None. Only then does the finished string get tokenized - with the reference images riding along - and encoded into the conditioning.

With a VAE connected, each reference image is also encoded and attached to the conditioning as a reference latent. That's the same dual-encoding split the Qwen Edit line uses: the VL encoder carries the visual meaning, the latent carries the visual appearance. Note the switch in the code - if there are reference latents, the node drops the vision tokens. You're choosing one path, not stacking both.

The inputs worth touching

  • clip - the Qwen Image 2.1 encoder, loaded with ComfyUI's CLIPLoader and its Qwen Image type (the pack's 📥 CLIP Download node reads that list off ComfyUI, so whatever your build calls it is what you pick).
  • positive - an instruction. On an LLM-encoded model, (face:1.4) is not ignored, it's fed to the encoder as literal punctuation, and masterpiece, best quality burns tokens that your actual description needs.
  • images - an autogrow socket. Connect one and the next appears, up to image_1 … image_16. Order counts: the first image decides the size of the latent output.
  • mode / resolution - max resolution keeps total pixel area equal to resolution² with the aspect ratio intact (mirroring ComfyUI's original Qwen Image 2.1 node); max size caps the long edge Kontext-style; resolution = 0 leaves the source alone.
  • vae (optional) - the switch described above.
  • style / translate / seed - the pack's sugar.

Outputs are positive (CONDITIONING, into your sampler), latent (a zeros LATENT sized to the first reference - this is your empty latent, not an img2img start), and prompt (the final translated text as a STRING, handy for dumping into a text save node or a filename).

Installing it

It's on the Comfy Registry, so ComfyUI Manager → search "SDVN" and install. Manually:

cd <ComfyUI>/custom_nodes
git clone https://github.com/StableDiffusionVN/SDVN_Comfy_node
cd ..                                  # back to the ComfyUI root
pip install -r custom_nodes/SDVN_Comfy_node/requirements.txt

Be warned, that requirements file is a whole toolkit's worth of dependencies for one text box: ultralytics==8.3.107 pinned, opencv-python, scipy, pandas, gallery-dl, instaloader, googletrans-py, google-genai, openai, dynamicprompts. The README also asks Windows and macOS users to install the aria2c binary separately for the pack's download nodes.

Then the models, which is the real cost. Comfy-Org publishes the 2.1 trio under Qwen-Image-2.1: a diffusion model (bf16 or int8_convrot), a text encoder (Qwen3-VL 8B in bf16 / int8_convrot / w4a8, plus two 9B builds tagged t2i and i2i), and qwen_image_2.1_vae_bf16. The encoder is 8-9B on its own, so it's a second VRAM budget - on a 30-series card grab the int8_convrot build, which ComfyUI has supported natively since v0.27 and which runs faster than fp8. The pack's 📥 AnyDownload List node already carries these files if you'd rather not hunt.

Where people get burned

  • Seed 0 looks like broken wildcards. The dynamic-prompt picker does seed % combinations, so a seed of 0 always resolves to the first combination. Move the seed and your {red|blue} starts behaving.
  • Silent translation. translate hits Google's endpoint at run time: no network, no run, and even when it works it rewrites your prompt. The node prints the translation - read it once.
  • Transparent PNGs go black. Before the image reaches the encoder, alpha is composited against black. Cutouts become black-background images. Flatten onto white first if that matters.
  • Don't mix latents. The latent output is 64-channel, spaced 16× - Qwen Image 2.1's VAE is not the 8×, 16-channel one the v1/Edit line uses. Feeding the wrong one to a sampler fails loudly, which is the good kind of failure.
  • There's no negative field, and you don't need one. Qwen Image 2.1 runs at the low-CFG guidance-distilled settings where the unconditional pass isn't computed at all, so anything you'd put there would be discarded. State constraints as presence: "plain seamless backdrop, sharp focus".
  • Update ComfyUI before filing a bug. This node is written against the newer V3 authoring API (io.ComfyNode, autogrow inputs), so an older build can fail to register it entirely.
Category📂 SDVN

Inputs (9)

NameTypeDefaultDescription
clipCLIPCLIP Qwen Image 2.1 dùng để mã hóa prompt.
positiveSTRINGPrompt tích cực mô tả nội dung bạn muốn sinh ra.
styleCOMBONoneChọn style mẫu có sẵn để thêm vào prompt.
translateCOMBONgôn ngữ dịch prompt.
modeCOMBOmax resolutionMax size giới hạn cạnh dài như KontextReference; max resolution giữ tổng diện tích như node Qwen Image 2.1 gốc.
resolutionINT10240–4096Kích thước ảnh tham chiếu theo chế độ đã chọn. 0 giữ kích thước ảnh gốc.
seedINT00–18446744073709550000Seed ngẫu nhiên cho prompt.
imagesCOMFY_AUTOGROW_V3Ảnh tham chiếu cho Qwen Image 2.1. Nối ảnh để hiện thêm đầu vào ảnh.
vaeoptVAEVAE dùng để mã hóa ảnh tham chiếu thành reference latent.

Outputs (3)

NameTypeDescription
positiveCONDITIONING—
latentLATENTLatent rỗng có kích thước khớp với ảnh tham chiếu đầu tiên.
promptSTRING—