ComfyUI Node
QwenVL (Advanced)
A ComfyUI node in 🧠aistudynow/QwenVL with 17 inputs and 1 output.
QwenVL (Advanced)
- image
- video
- text
â—„model_nameQwen3-VL-4B-Instructâ–º
â—„quantization8-bit (Balanced)â–º
â—„preset_promptDescribe this image in detail.â–º
â—„custom_promptâ–º
â—„max_tokens1024â–º
â—„temperature0.6â–º
â—„top_p0.90â–º
â—„num_beams1â–º
â—„repetition_penalty1.20â–º
â—„frame_count16â–º
â—„deviceautoâ–º
â—„use_torch_compilefalseâ–º
â—„keep_model_loadedtrueâ–º
â—„seed1â–º
â—„attention_modeautoâ–º
Category🧠aistudynow/QwenVL
Inputs (17)
| Name | Type | Default | Description |
|---|---|---|---|
| model_name | COMBO | Qwen3-VL-4B-Instruct | 18 options: Qwen3-VL-2B-Instruct, Qwen3-VL-2B-Thinking, Qwen3-VL-2B-Instruct-FP8, Qwen3-VL-2B-Thinking-FP8, Qwen3-VL-4B-Instruct, Qwen3-VL-4B-Thinking, +12 |
| quantization | COMBO | 8-bit (Balanced) | 3 options: 4-bit (VRAM-friendly), 8-bit (Balanced), None (FP16) |
| preset_prompt | COMBO | Describe this image in detail. | 13 options: Describe this image in detail., Describe this video in detail., Summarize the key events in this video., Generate 5 descriptive keywords for this content., Create a detailed text-to-image prompt from this image., Generate a detailed Stable Diffusion prompt that includes subject, background, lighting, and style., +7 |
| custom_prompt | STRING | — | |
| max_tokens | INT | 102464–2048 | — |
| temperature | FLOAT | 0.60.1–1 | — |
| top_p | FLOAT | 0.900–1 | — |
| num_beams | INT | 11–10 | — |
| repetition_penalty | FLOAT | 1.200–2 | — |
| frame_count | INT | 161–64 | — |
| device | COMBO | auto | 4 options: auto, cuda, cpu, mps |
| use_torch_compile | BOOLEAN | false | — |
| keep_model_loaded | BOOLEAN | true | — |
| seed | INT | 11–18446744073709550000 | — |
| attention_mode | COMBO | auto | 4 options: auto, sage, flash_attention_2, sdpa |
| imageopt | IMAGE | — | |
| videoopt | IMAGE | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| text | STRING | — |