ComfyUI Node
QwenVL
A ComfyUI node in 🧠aistudynow/QwenVL with 10 inputs and 1 output.
QwenVL
- image
- video
- text
â—„model_nameQwen3-VL-4B-Instructâ–º
â—„quantization8-bit (Balanced)â–º
â—„preset_promptDescribe this image in detail.â–º
â—„custom_promptâ–º
â—„max_tokens1024â–º
â—„keep_model_loadedtrueâ–º
â—„seed1â–º
â—„attention_modeautoâ–º
Category🧠aistudynow/QwenVL
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| model_name | COMBO | Qwen3-VL-4B-Instruct | 18 options: Qwen3-VL-2B-Instruct, Qwen3-VL-2B-Thinking, Qwen3-VL-2B-Instruct-FP8, Qwen3-VL-2B-Thinking-FP8, Qwen3-VL-4B-Instruct, Qwen3-VL-4B-Thinking, +12 |
| quantization | COMBO | 8-bit (Balanced) | 3 options: 4-bit (VRAM-friendly), 8-bit (Balanced), None (FP16) |
| preset_prompt | COMBO | Describe this image in detail. | 13 options: Describe this image in detail., Describe this video in detail., Summarize the key events in this video., Generate 5 descriptive keywords for this content., Create a detailed text-to-image prompt from this image., Generate a detailed Stable Diffusion prompt that includes subject, background, lighting, and style., +7 |
| custom_prompt | STRING | — | |
| max_tokens | INT | 102464–2048 | — |
| keep_model_loaded | BOOLEAN | true | — |
| seed | INT | 11–18446744073709550000 | — |
| attention_mode | COMBO | auto | 4 options: auto, sage, flash_attention_2, sdpa |
| imageopt | IMAGE | — | |
| videoopt | IMAGE | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| text | STRING | — |