ComfyUI Node
Smart TextEncodeEditAdvanced (6 images)
A ComfyUI node in SmartHelperNodes with 12 inputs and 1 output.
Smart TextEncodeEditAdvanced (6 images)
- clip
- vae
- image1
- image2
- image3
- image4
- image5
- image6
- CONDITIONING
◄prompt—►
◄use_vl_encodingtrue►
◄vl_megapixels0.50►
◄max_images_allowed6►
CategorySmartHelperNodes
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| clip | CLIP | — | |
| prompt | STRING | — | |
| use_vl_encoding | BOOLEAN | true | Enable VL image feeding: prepend Picture N vision tokens, pass downscaled images to clip.tokenize, and apply the edit-style llama_template. Turn off to behave like plain text encode + reference_latents. |
| vl_megapixels | FLOAT | 0.500–4 | Target megapixels for Vision-Language model. Set to 0 to disable VL image feeding. Recommended: 0.2-1.0 MP. Qwen2.5-VL trained range: 0.2-1.0 MP |
| max_images_allowed | COMBO | 6 | Maximum number of images to process. Images are processed in order: image1..image6 |
| vaeopt | VAE | — | |
| image1opt | IMAGE | — | |
| image2opt | IMAGE | — | |
| image3opt | IMAGE | — | |
| image4opt | IMAGE | — | |
| image5opt | IMAGE | — | |
| image6opt | IMAGE | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| CONDITIONING | CONDITIONING | — |