ComfyUI Node
Vision Prompt Assistant
Generates text locally with a compatible multimodal CLIP such as Qwen3-VL. Supports separate system/user prompts and up to three reference images. For fuller results, end the user prompt with the desired approximate token count, keeping it slightly below max_length (for example: 'Write about 220 tokens' with max_length set to 256).
Vision Prompt Assistant
- image_0
- image_1
- image_2
- generated_text
◄clip_name►
◄clip_typeltxv►
◄load_devicedefault►
◄user_promptAnalyze the reference images and write a detailed generation prompt.►
◄system_promptYou write production-ready prompts for MiniMax H3 Reference to Video. Use the exact supplied <Picture n> tags, clearly assigning identity, appearance, style, motion, and camera. Return only the final generation prompt.►
◄max_length256►
◄samplingtrue►
◄temperature0.70►
◄top_k40►
◄top_p0.90►
◄min_p0.05►
◄repetition_penalty1.05►
◄seed0►
Categorytext
Inputs (16)
| Name | Type | Default | Description |
|---|---|---|---|
| clip_name | COMBO | 0 options: | |
| clip_type | COMBO | ltxv | 28 options: stable_diffusion, stable_cascade, sd3, stable_audio, mochi, ltxv, +22 |
| load_device | COMBO | default | 2 options: default, cpu |
| user_prompt | STRING | Analyze the reference images and write a detailed generation prompt. | — |
| system_prompt | STRING | You write production-ready prompts for MiniMax H3 Reference to Video. Use the exact supplied <Picture n> tags, clearly assigning identity, appearance, style, motion, and camera. Return only the final generation prompt. | — |
| max_length | INT | 2561–4096 | Hard generation limit. For a fuller prompt, also request an approximate token count near the end of user_prompt, slightly below this value. |
| sampling | BOOLEAN | true | — |
| temperature | FLOAT | 0.700.01–2 | — |
| top_k | INT | 400–1000 | — |
| top_p | FLOAT | 0.900–1 | — |
| min_p | FLOAT | 0.050–1 | — |
| repetition_penalty | FLOAT | 1.050–5 | — |
| seed | INT | 00–18446744073709550000 | — |
| image_0opt | IMAGE | — | |
| image_1opt | IMAGE | — | |
| image_2opt | IMAGE | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| generated_text | STRING | — |