Nodes/ComfyUI-ialhabbal/ialhabbal VLLM Advanced
ComfyUI Node

ialhabbal VLLM Advanced

A ComfyUI node in πŸ§ͺialhabbal_VLLM with 17 inputs and 1 output.

By ialhabbalΒ·Created 4 months agoΒ·Updated about a month agoΒ· 7
ialhabbal VLLM Advanced
  • image
  • video
  • RESPONSE
β—„model_nameQwen3-VL-2B-Instructβ–Ί
β—„quantizationNone (FP16)β–Ί
β—„attention_modeautoβ–Ί
β—„use_torch_compilefalseβ–Ί
β—„deviceautoβ–Ί
β—„preset_promptπŸ–ΌοΈ Detailed Descriptionβ–Ί
β—„custom_promptβ–Ί
β—„max_tokens512β–Ί
β—„temperature0.60β–Ί
β—„top_p0.90β–Ί
β—„num_beams1β–Ί
β—„repetition_penalty1.20β–Ί
β—„frame_count16β–Ί
β—„keep_model_loadedtrueβ–Ί
β—„seed1β–Ί
CategoryπŸ§ͺialhabbal_VLLM

Inputs (17)

NameTypeDefaultDescription
model_nameCOMBOQwen3-VL-2B-InstructPick the Qwen-VL checkpoint. First run downloads weights into models/LLM/Qwen-VL, so leave disk space.
quantizationCOMBONone (FP16)Precision vs VRAM. FP16 gives the best quality if memory allows; 8-bit suits 8–16 GB GPUs; 4-bit fits 6 GB or lower but is slower.
attention_modeCOMBOautoauto tries flash-attn v2 when installed and falls back to SDPA. Only override when debugging attention backends.
use_torch_compileBOOLEANfalseEnable torch.compile('reduce-overhead') on supported CUDA/Torch 2.1+ builds for extra throughput after the first compile.
deviceCOMBOautoChoose where to run the model: auto, cpu, mps, or cuda:x for multi-GPU systems.
preset_promptCOMBOπŸ–ΌοΈ Detailed DescriptionBuilt-in instruction describing how Qwen-VL should analyze the media input.
custom_promptSTRINGOptional overrideβ€”when filled it completely replaces the preset template.
max_tokensINT51264–4096Maximum number of new tokens to decode. Larger values yield longer answers but consume more time and memory.
temperatureFLOAT0.600.1–1Sampling randomness when num_beams == 1. 0.2–0.4 is focused, 0.7+ is creative.
top_pFLOAT0.900–1Nucleus sampling cutoff when num_beams == 1. Lower values keep only top tokens; 0.9–0.95 allows more variety.
num_beamsINT11–8Beam-search width. Values >1 disable temperature/top_p and trade speed for more stable answers.
repetition_penaltyFLOAT1.200.5–2Values >1 (e.g., 1.1–1.3) penalize repeated phrases; 1.0 leaves logits untouched.
frame_countINT161–64Number of frames extracted from video inputs before prompting Qwen-VL. More frames provide context but cost time.
keep_model_loadedBOOLEANtrueKeeps the model resident in VRAM/RAM after the run so the next prompt skips loading.
seedINT11–4294967295Seed controlling sampling and frame picking; reuse it to reproduce results.
imageoptIMAGEβ€”
videooptIMAGEβ€”

Outputs (1)

NameTypeDescription
RESPONSESTRINGβ€”