Nodes/ComfyUI_Simple_Qwen3-VL-gguf/Qwen-VL Vision Language Model
ComfyUI Node

Qwen-VL Vision Language Model

THIS NODE IS NO LONGER SUPPORTED! Use "Simple Qwen-VL Vision Language Model"

By KLL535·Created 9 months ago·Updated about a month ago· 80
Qwen-VL Vision Language Model
  • image
  • image2
  • image3
  • text
  • conditioning
system_promptYou are a highly accurate vision-language assistant. Provide detailed, precise, and well-structured image descriptions.
user_promptDescribe this image.
model_path
mmproj_path
output_max_tokens2048
image_max_tokens4096
ctx8192
n_batch512
gpu_layers-1
temperature0.70
seed42
unload_all_modelsfalse
top_p0.92
repeat_penalty1.20
top_k0
pool_size4194304
script
Category🌐 SimpleQwenVL

Inputs (20)

NameTypeDefaultDescription
system_promptSTRINGYou are a highly accurate vision-language assistant. Provide detailed, precise, and well-structured image descriptions.
user_promptSTRINGDescribe this image.
model_pathSTRING
mmproj_pathSTRING
output_max_tokensINT204864–4096
image_max_tokensINT40961024–1024000
ctxINT81921024–1024000
n_batchINT51264–1024000
gpu_layersINT-1-1–100
temperatureFLOAT0.700–2
seedINT42
unload_all_modelsBOOLEANfalse
top_pFLOAT0.920–1
repeat_penaltyFLOAT1.201–2
top_kINT00–32768
pool_sizeINT41943041048576–10485760
imageoptIMAGE
image2optIMAGE
image3optIMAGE
scriptoptSTRING

Outputs (2)

NameTypeDescription
textSTRING
conditioningCONDITIONING