Nodes/ComfyUI-Youtu-VL/Youtu-VL (Advanced)
ComfyUI Node

Youtu-VL (Advanced)

A ComfyUI node in 🧪AILab/YoutuVL with 13 inputs and 1 output.

By 1038lab·Created 6 months ago·Updated 6 months ago· 13
Youtu-VL (Advanced)
  • image
  • text
modelYoutu-VL-4B-Instruct
quantizationNone (FP16)
attention_modeauto
deviceauto
preset_prompt🖼️ Describe Image
custom_prompt
max_tokens512
temperature0.10
top_p0.001
repetition_penalty1.05
keep_model_loadedtrue
seed1
Category🧪AILab/YoutuVL

Inputs (13)

NameTypeDefaultDescription
modelCOMBOYoutu-VL-4B-InstructSelect the Youtu-VL model. First run downloads weights to models/LLM/Youtu-VL.
quantizationCOMBONone (FP16)Precision vs VRAM. FP16 gives best quality; 8-bit suits 8-16GB GPUs; 4-bit fits 6GB or less.
attention_modeCOMBOautoauto tries flash-attn v2 when available, falls back to SDPA.
deviceCOMBOauto4 options: auto, cuda, cpu, mps
preset_promptCOMBO🖼️ Describe ImageBuilt-in instruction for how Youtu-VL should analyze the input.
custom_promptSTRINGOptional override - replaces preset template when filled.
max_tokensINT51264–32768Maximum number of new tokens to generate.
temperatureFLOAT0.100.01–2Sampling randomness. 0.1-0.4 is focused, 0.7+ is creative.
top_pFLOAT0.0010.001–1Nucleus sampling cutoff. Lower values keep only top tokens.
repetition_penaltyFLOAT1.050.5–2Values >1 penalize repeated phrases.
keep_model_loadedBOOLEANtrueKeep model in VRAM after run for faster subsequent inference.
seedINT11–4294967295Seed for reproducible results.
imageoptIMAGE

Outputs (1)

NameTypeDescription
textSTRING