Nodes/ComfyUI-ZhiHui/🎨 智绘_图片反推
ComfyUI Node

🎨 智绘_图片反推

Image-to-prompt reversal with BLIP, Qwen2-VL and full generation controls

By zhuyungen·Created 8 months ago·Updated 4 months ago· 0
🎨 智绘_图片反推
  • 🖼️ 图像1
  • 🖼️ 图像2
  • 🖼️ 图像3
  • 🖼️ 图像4
  • 🎥 视频
  • 🎯 Qwen3VL额外选项
  • 描述文本
◄🤖 模型选择BLIP-Large (推荐-精度高)►
◄⚙️ 量化级别None (FP16)►
◄💭 预设提示词提示词风格 - 详细►
◄✏️ 自定义提示词►
◄🔢 最大令牌数1234►
◄🌡️ 采样温度0.6►
◄🎯 核采样参数0.90►
◄🔍 束搜索数量1►
◄🚫 重复惩罚1.20►
◄🎬 视频帧数16►
◄💻 设备选择auto►
◄🚀 开启TF32加速false►
◄🔄 保持模型加载true►
◄🎲 随机种子-1►
◄🎯 种子控制随机►

Image "反推" - reverse-engineering a prompt from an image - is the workhorse move of the SD ecosystem, and ZH_ImageCaptioning (🎨 智绘_图片反推) is the 智绘灵箱 pack's serious version of it. Not a toy: five model choices from BLIP up to Qwen2-VL-7B, 4-bit/8-bit quantization, custom prompt templates, and a full bank of generation controls (temperature, top-p, beam search, repetition penalty, seed modes) that you simply don't get from the one-shot captioners most packs ship. It also takes up to four reference images and a video-frame sequence, which puts it well past "describe this one picture."

How it works

The node maintains a persistent model session - load once, keep warm (there's a 保持模型加载 (keep loaded) toggle, default on) - and runs your chosen VLM over the input frames. The model tier decides everything:

  • BLIP-Large (default) / BLIP-Base - the old reliable captioners. Fast, light, single-image only, and honestly what most "tag reversal" tasks still use.
  • BLIP2-2.7B - stronger, needs ~3GB+ VRAM, still single-image.
  • Qwen2-VL-2B / 7B - the current quality frontier here. Multi-image and video support, real reasoning, and the 7B is the one to reach for when the caption needs to understand the scene, not just list it. Expect it to be slow and heavy.

Models live in ComfyUI/models/prompt_generator/ and download on first use - the pack's own requirements notes 1–14GB depending on model, so budget the disk. Qwen3-VL also works if you connect the optional Qwen3VL额外选项 node; just know it demands transformers >= 4.57.0 per the requirements file.

Generation controls mirror what you'd set in an LLM sampler: 采样温度 (temperature, 0.6 default - higher = more creative), 核采样参数 (top-p, 0.9), 束搜索数量 (beam count, 1–10 - more beams, more accurate, slower), 重复惩罚 (repetition penalty, 1.2 default), 最大令牌数 (max tokens, 1234 default). The seed controls (随机种子 with 随机/固定/递增 modes) let you lock a caption or roll a different phrasing each run. There's even 开启TF32加速 for Ampere+ GPUs and a 设备选择 (auto/cuda/cpu/mps).

The inputs that matter

For a beginner, the one setting that changes everything is 🤖 模型选择 - start on BLIP-Large, move to Qwen2-VL-2B when the captions get too dumb. 💭 预设提示词 picks the output style (标签/tag-flavored, 简单, 详细, 极致详细, 电影感) - for LoRA training you want the tag style; for general use the detailed one reads better. ✍️ 自定义提示词 overrides the template entirely, which is how you make it ask the model specific questions ("describe only the lighting"). Inputs: up to four 🖼️ 图像 sockets plus a 🎥 视频 socket (Qwen models only). Output: 描述文本 (STRING) - straight into a text box, a prompt node, or a save.

Install

Part of the 智绘灵箱 (ComfyUI-ZhiHui) pack:

cd ComfyUI/custom_nodes
git clone https://github.com/zhuyungen/ComfyUI-ZhiHui.git
pip install -r requirements.txt

The heavy deps are unavoidable here: transformers, accelerate, bitsandbytes (for 4/8-bit), modelscope (download mirror). If you use an older ComfyUI install, update transformers first - pip install --upgrade transformers - or the node errors on import. ComfyUI Manager search "智绘灵箱" / "ComfyUI-ZhiHui" also works, then install the requirements manually.

Where people get burned

  • First run is a model download, not a crash. Give it 1–14GB of patience, and check ComfyUI/models/prompt_generator/ if it seems stuck. For users in China, install modelscope so the download uses the domestic mirror.
  • Quantization isn't free. 4-bit saves VRAM but visibly degrades caption quality; 8-bit is the usual balance. On a 7B model it's the difference between fitting and not.
  • BLIP can't see your multi-image setup. Multi-image and video inputs only work with the Qwen models - the node's own docs say it plainly. Feed BLIP a batch and it just won't behave.
  • "反推" is descriptive, not instructional. The caption is a description of what's in the frame; if your goal is LoRA training captions, use the tag preset and keep prompts short - that's a dataset-format decision, not a node setting.

If you're already inside this pack, this is the captioner that scales with you: free and fast when you need tags, heavy and smart when you need a real description. Just bring a big-enough drive and a first-run coffee.

Category智绘灵箱/图片

Inputs (21)

NameTypeDefaultDescription
🤖 模型选择COMBOBLIP-Large (推荐-精度高)选择图片理解模型
⚙️ 量化级别COMBONone (FP16)量化级别,降低显存占用
💭 预设提示词COMBO提示词风格 - 详细选择预设的提示词模板
✏️ 自定义提示词STRING—
🔢 最大令牌数INT123464–4096生成文本的最大令牌数
🌡️ 采样温度FLOAT0.60.1–1创造性控制,值越高越随机
🎯 核采样参数FLOAT0.900–1核采样参数
🔍 束搜索数量INT11–10束搜索数量,值越大越准确但越慢
🚫 重复惩罚FLOAT1.200–2重复惩罚,防止生成重复内容
🎬 视频帧数INT161–64视频帧数(用于视频输入的采样)
💻 设备选择COMBOauto设备选择
🚀 开启TF32加速BOOLEANfalse启用TF32加速(仅支持Ampere及以上架构显卡,如30/40/50系,能显著提升速度)
🔄 保持模型加载BOOLEANtrue是否保持模型加载(推荐开启)
🎲 随机种子INT-1-1–18446744073709550000随机种子,-1为随机
🎯 种子控制COMBO随机种子控制模式
🖼️ 图像1optIMAGE输入图片1
🖼️ 图像2optIMAGE输入图片2
🖼️ 图像3optIMAGE输入图片3
🖼️ 图像4optIMAGE输入图片4
🎥 视频optIMAGE视频输入(作为图像序列)
🎯 Qwen3VL额外选项optQWEN3VL_EXTRA_OPTIONS可选的Qwen3VL额外选项,连接Qwen3VL额外选项节点

Outputs (1)

NameTypeDescription
描述文本STRING—