ComfyUI Node

Youtu-VL

A ComfyUI node in 🧪AILab/YoutuVL with 9 inputs and 1 output.

By 1038lab·Created 6 months ago·Updated 6 months ago· 13
Youtu-VL
  • image
  • text
modelYoutu-VL-4B-Instruct
quantizationNone (FP16)
attention_modeauto
preset_prompt🖼️ Describe Image
custom_prompt
max_tokens512
keep_model_loadedtrue
seed1
Category🧪AILab/YoutuVL

Inputs (9)

NameTypeDefaultDescription
modelCOMBOYoutu-VL-4B-InstructSelect the Youtu-VL model. First run downloads weights to models/LLM/Youtu-VL.
quantizationCOMBONone (FP16)Precision vs VRAM. FP16 gives best quality; 8-bit suits 8-16GB GPUs; 4-bit fits 6GB or less.
attention_modeCOMBOautoauto tries flash-attn v2 when available, falls back to SDPA.
preset_promptCOMBO🖼️ Describe ImageBuilt-in instruction for how Youtu-VL should analyze the input.
custom_promptSTRINGOptional override - replaces preset template when filled.
max_tokensINT51264–4096Maximum number of new tokens to generate.
keep_model_loadedBOOLEANtrueKeep model in VRAM after run for faster subsequent inference.
seedINT11–4294967295Seed for reproducible results.
imageoptIMAGE

Outputs (1)

NameTypeDescription
textSTRING