ComfyUI Node
Qwen-VL Vision Language Model
THIS NODE IS NO LONGER SUPPORTED! Use "Simple Qwen-VL Vision Language Model"
Qwen-VL Vision Language Model
- image
- image2
- image3
- text
- conditioning
◄system_promptYou are a highly accurate vision-language assistant. Provide detailed, precise, and well-structured image descriptions.►
◄user_promptDescribe this image.►
◄model_path►
◄mmproj_path►
◄output_max_tokens2048►
◄image_max_tokens4096►
◄ctx8192►
◄n_batch512►
◄gpu_layers-1►
◄temperature0.70►
◄seed42►
◄unload_all_modelsfalse►
◄top_p0.92►
◄repeat_penalty1.20►
◄top_k0►
◄pool_size4194304►
◄script►
Category🌐 SimpleQwenVL
Inputs (20)
| Name | Type | Default | Description |
|---|---|---|---|
| system_prompt | STRING | You are a highly accurate vision-language assistant. Provide detailed, precise, and well-structured image descriptions. | — |
| user_prompt | STRING | Describe this image. | — |
| model_path | STRING | — | |
| mmproj_path | STRING | — | |
| output_max_tokens | INT | 204864–4096 | — |
| image_max_tokens | INT | 40961024–1024000 | — |
| ctx | INT | 81921024–1024000 | — |
| n_batch | INT | 51264–1024000 | — |
| gpu_layers | INT | -1-1–100 | — |
| temperature | FLOAT | 0.700–2 | — |
| seed | INT | 42 | — |
| unload_all_models | BOOLEAN | false | — |
| top_p | FLOAT | 0.920–1 | — |
| repeat_penalty | FLOAT | 1.201–2 | — |
| top_k | INT | 00–32768 | — |
| pool_size | INT | 41943041048576–10485760 | — |
| imageopt | IMAGE | — | |
| image2opt | IMAGE | — | |
| image3opt | IMAGE | — | |
| scriptopt | STRING | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| text | STRING | — |
| conditioning | CONDITIONING | — |