ComfyUI Node
ialhabbal VLLM
A ComfyUI node in π§ͺialhabbal_VLLM with 10 inputs and 1 output.
ialhabbal VLLM
- image
- video
- RESPONSE
βmodel_nameQwen3-VL-2B-InstructβΊ
βquantizationNone (FP16)βΊ
βattention_modeautoβΊ
βpreset_promptπΌοΈ Detailed DescriptionβΊ
βcustom_promptβΊ
βmax_tokens512βΊ
βkeep_model_loadedtrueβΊ
βseed1βΊ
Categoryπ§ͺialhabbal_VLLM
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| model_name | COMBO | Qwen3-VL-2B-Instruct | Pick the Qwen-VL checkpoint. First run downloads weights into models/LLM/Qwen-VL, so leave disk space. |
| quantization | COMBO | None (FP16) | Precision vs VRAM. FP16 gives the best quality if memory allows; 8-bit suits 8β16 GB GPUs; 4-bit fits 6 GB or lower but is slower. |
| attention_mode | COMBO | auto | auto tries flash-attn v2 when installed and falls back to SDPA. Only override when debugging attention backends. |
| preset_prompt | COMBO | πΌοΈ Detailed Description | Built-in instruction describing how Qwen-VL should analyze the media input. |
| custom_prompt | STRING | Optional overrideβwhen filled it completely replaces the preset template. | |
| max_tokens | INT | 51264β2048 | Maximum number of new tokens to decode. Larger values yield longer answers but consume more time and memory. |
| keep_model_loaded | BOOLEAN | true | Keeps the model resident in VRAM/RAM after the run so the next prompt skips loading. |
| seed | INT | 11β4294967295 | Seed controlling sampling and frame picking; reuse it to reproduce results. |
| imageopt | IMAGE | β | |
| videoopt | IMAGE | β |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| RESPONSE | STRING | β |