Nodes/Prompt Assistant/✨Image Caption (VLM)
ComfyUI Node

✨Image Caption (VLM)

Extract text prompt from image using Vision-Language Models

By yawiii·Created about a year ago·Updated 4 months ago· 2,119
✨Image Caption (VLM)
  • image
  • caption_text
  • caption_list
rule像素级描述(by:阿丹)
custom_rulefalse
custom_rule_content
user_prompt
vlm_service智谱/glm-4.6V-Flash
ollama_auto_unloadtrue
seed0
Category✨Prompt Assistant

Inputs (8)

NameTypeDefaultDescription
imageIMAGEThe image to analyze. Supports single image or IMAGE batch (processes each frame independently)
ruleCOMBO像素级描述(by:阿丹)Choose a preset rule for analysis
custom_ruleBOOLEANfalseEnable custom rule input
custom_rule_contentSTRINGCustom rule content, only used when Custom Rule is enabled
user_promptSTRINGEnter additional prompts here, sent with the rule
vlm_serviceCOMBO智谱/glm-4.6V-FlashSelect VLM Service
ollama_auto_unloadBOOLEANtrueAuto unload Ollama model after generation
seedINT00–18446744073709550000Controls randomness. Set to non-fixed mode to force re-execution

Outputs (2)

NameTypeDescription
caption_textSTRING
caption_listSTRING