ComfyUI Node
ArchAi3D_Qwen_Encoder
A ComfyUI node in ArchAi3d/Qwen with 18 inputs and 3 outputs.
ArchAi3D_Qwen_Encoder
- clip
- vae
- image1_vl
- image2_vl
- image3_vl
- image1_latent
- image2_latent
- image3_latent
- conditioning
- latent
- formatted_prompt
◄prompt—►
◄system_prompt►
◄image1_labelImage 1►
◄image2_labelImage 2►
◄image3_labelImage 3►
◄conditioning_strength1.00►
◄image1_latent_strength1.00►
◄image2_latent_strength1.00►
◄image3_latent_strength1.00►
◄debug_modefalse►
CategoryArchAi3d/Qwen
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| clip | CLIP | Qwen-VL CLIP model for encoding text and vision tokens | |
| prompt | STRING | Text prompt (vision tokens inserted automatically in ChatML format) | |
| system_prompt | STRING | Optional system prompt (wrapped in ChatML <|im_start|>system block) | |
| image1_label | STRING | Image 1 | Custom label for Image 1 (e.g., 'Image 1 (target)', 'Image 1 (room)') |
| image2_label | STRING | Image 2 | Custom label for Image 2 (e.g., 'Image 2 (style ref)', 'Image 2 (material)') |
| image3_label | STRING | Image 3 | Custom label for Image 3 (e.g., 'Image 3 (color ref)', 'Image 3 (lighting)') |
| conditioning_strength | FLOAT | 1.000–2 | Global strength for text+vision embeddings (1.0=normal, <1.0=weaker, >1.0=stronger) |
| image1_latent_strength | FLOAT | 1.000–2 | Image1 latent strength (1.0=normal, <1.0=weaker, >1.0=stronger) |
| image2_latent_strength | FLOAT | 1.000–2 | Image2 latent strength (1.0=normal, <1.0=weaker, >1.0=stronger) |
| image3_latent_strength | FLOAT | 1.000–2 | Image3 latent strength (1.0=normal, <1.0=weaker, >1.0=stronger) |
| debug_mode | BOOLEAN | false | Enable console logging (shows strengths, shapes, and formatted prompt) |
| vaeopt | VAE | VAE for encoding reference latents (required if using latent images) | |
| image1_vlopt | IMAGE | Image 1 for vision encoder (RGB only, expects correct size) | |
| image2_vlopt | IMAGE | Image 2 for vision encoder (RGB only, expects correct size) | |
| image3_vlopt | IMAGE | Image 3 for vision encoder (RGB only, expects correct size) | |
| image1_latentopt | IMAGE | Image 1 for reference latent (RGB only, expects correct size) | |
| image2_latentopt | IMAGE | Image 2 for reference latent (RGB only, expects correct size) | |
| image3_latentopt | IMAGE | Image 3 for reference latent (RGB only, expects correct size) |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| conditioning | CONDITIONING | Text+vision embeddings with reference latents metadata attached |
| latent | LATENT | Image1 latent in standard format (for VAEDecode or other latent nodes) |
| formatted_prompt | STRING | Final ChatML-formatted prompt with vision tokens (for debugging) |