Nodes/ComfyUI-VisionPromptAssistant/Vision Prompt Assistant
ComfyUI Node

Vision Prompt Assistant

Generates text locally with a compatible multimodal CLIP such as Qwen3-VL. Supports separate system/user prompts and up to three reference images. For fuller results, end the user prompt with the desired approximate token count, keeping it slightly below max_length (for example: 'Write about 220 tokens' with max_length set to 256).

By elgalardi·Created 21 days ago·Updated a day ago· 1
Vision Prompt Assistant
  • image_0
  • image_1
  • image_2
  • generated_text
clip_name
clip_typeltxv
load_devicedefault
user_promptAnalyze the reference images and write a detailed generation prompt.
system_promptYou write production-ready prompts for MiniMax H3 Reference to Video. Use the exact supplied <Picture n> tags, clearly assigning identity, appearance, style, motion, and camera. Return only the final generation prompt.
max_length256
samplingtrue
temperature0.70
top_k40
top_p0.90
min_p0.05
repetition_penalty1.05
seed0
Categorytext

Inputs (16)

NameTypeDefaultDescription
clip_nameCOMBO0 options:
clip_typeCOMBOltxv28 options: stable_diffusion, stable_cascade, sd3, stable_audio, mochi, ltxv, +22
load_deviceCOMBOdefault2 options: default, cpu
user_promptSTRINGAnalyze the reference images and write a detailed generation prompt.
system_promptSTRINGYou write production-ready prompts for MiniMax H3 Reference to Video. Use the exact supplied <Picture n> tags, clearly assigning identity, appearance, style, motion, and camera. Return only the final generation prompt.
max_lengthINT2561–4096Hard generation limit. For a fuller prompt, also request an approximate token count near the end of user_prompt, slightly below this value.
samplingBOOLEANtrue
temperatureFLOAT0.700.01–2
top_kINT400–1000
top_pFLOAT0.900–1
min_pFLOAT0.050–1
repetition_penaltyFLOAT1.050–5
seedINT00–18446744073709550000
image_0optIMAGE
image_1optIMAGE
image_2optIMAGE

Outputs (1)

NameTypeDescription
generated_textSTRING