ComfyUI Node: CLIP Text Image Encode
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.
Category
conditioning
Inputs
vl_model ZIMAGE_VL_MODEL
text STRING
image1 IMAGE
image1_mask MASK
image2 IMAGE
image2_mask MASK
image3 IMAGE
image3_mask MASK
Outputs
CONDITIONING
CONDITIONING
Extension: comfyui-zimage-vl
ComfyUI plugin providing vision-language model (VLM) conditioning via Qwen3-VL for Z-Image-Turbo video generation with multi-modal concept consistency.
Authored by yaofeng
Looking for a different node?
More nodes in comfyui-zimage-vl
Other conditioning nodes
- ACE-Step 1.5 Task Text Encode ⚡🅡🅞🅣🅘
- Add Artist To CSV
- Advanced Tiling
- Apply ConDelta
- Apply ConDelta AutoScale
- Apply Cosmos Reference Latent
- BlehBlendConditioning
- CLIP Text Encode with Caching
- CachingCLIPTextEncode|ARZUMATA
- CFG-less Negative Prompt
- Character Selector (positive only)
- Smooth Clamp ConDelta between -1 and 1
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.