🎨 智绘_图片反推
Image-to-prompt reversal with BLIP, Qwen2-VL and full generation controls
- 🖼️ 图像1
- 🖼️ 图像2
- 🖼️ 图像3
- 🖼️ 图像4
- 🎥 视频
- 🎯 Qwen3VL额外选项
- 描述文本
Image "反推" - reverse-engineering a prompt from an image - is the workhorse move of the SD ecosystem, and ZH_ImageCaptioning (🎨 智绘_图片反推) is the 智绘灵箱 pack's serious version of it. Not a toy: five model choices from BLIP up to Qwen2-VL-7B, 4-bit/8-bit quantization, custom prompt templates, and a full bank of generation controls (temperature, top-p, beam search, repetition penalty, seed modes) that you simply don't get from the one-shot captioners most packs ship. It also takes up to four reference images and a video-frame sequence, which puts it well past "describe this one picture."
How it works
The node maintains a persistent model session - load once, keep warm (there's a 保持模型加载 (keep loaded) toggle, default on) - and runs your chosen VLM over the input frames. The model tier decides everything:
- BLIP-Large (default) / BLIP-Base - the old reliable captioners. Fast, light, single-image only, and honestly what most "tag reversal" tasks still use.
- BLIP2-2.7B - stronger, needs ~3GB+ VRAM, still single-image.
- Qwen2-VL-2B / 7B - the current quality frontier here. Multi-image and video support, real reasoning, and the 7B is the one to reach for when the caption needs to understand the scene, not just list it. Expect it to be slow and heavy.
Models live in ComfyUI/models/prompt_generator/ and download on first use - the pack's own requirements notes 1–14GB depending on model, so budget the disk. Qwen3-VL also works if you connect the optional Qwen3VL额外选项 node; just know it demands transformers >= 4.57.0 per the requirements file.
Generation controls mirror what you'd set in an LLM sampler: 采样温度 (temperature, 0.6 default - higher = more creative), 核采样参数 (top-p, 0.9), 束搜索数量 (beam count, 1–10 - more beams, more accurate, slower), 重复惩罚 (repetition penalty, 1.2 default), 最大令牌数 (max tokens, 1234 default). The seed controls (随机种子 with 随机/固定/递增 modes) let you lock a caption or roll a different phrasing each run. There's even 开启TF32加速 for Ampere+ GPUs and a 设备选择 (auto/cuda/cpu/mps).
The inputs that matter
For a beginner, the one setting that changes everything is 🤖 模型选择 - start on BLIP-Large, move to Qwen2-VL-2B when the captions get too dumb. 💭 预设提示词 picks the output style (标签/tag-flavored, 简单, 详细, 极致详细, 电影感) - for LoRA training you want the tag style; for general use the detailed one reads better. ✍️ 自定义提示词 overrides the template entirely, which is how you make it ask the model specific questions ("describe only the lighting"). Inputs: up to four 🖼️ 图像 sockets plus a 🎥 视频 socket (Qwen models only). Output: 描述文本 (STRING) - straight into a text box, a prompt node, or a save.
Install
Part of the 智绘灵箱 (ComfyUI-ZhiHui) pack:
cd ComfyUI/custom_nodes
git clone https://github.com/zhuyungen/ComfyUI-ZhiHui.git
pip install -r requirements.txt
The heavy deps are unavoidable here: transformers, accelerate, bitsandbytes (for 4/8-bit), modelscope (download mirror). If you use an older ComfyUI install, update transformers first - pip install --upgrade transformers - or the node errors on import. ComfyUI Manager search "智绘灵箱" / "ComfyUI-ZhiHui" also works, then install the requirements manually.
Where people get burned
- First run is a model download, not a crash. Give it 1–14GB of patience, and check
ComfyUI/models/prompt_generator/if it seems stuck. For users in China, installmodelscopeso the download uses the domestic mirror. - Quantization isn't free. 4-bit saves VRAM but visibly degrades caption quality; 8-bit is the usual balance. On a 7B model it's the difference between fitting and not.
- BLIP can't see your multi-image setup. Multi-image and video inputs only work with the Qwen models - the node's own docs say it plainly. Feed BLIP a batch and it just won't behave.
- "反推" is descriptive, not instructional. The caption is a description of what's in the frame; if your goal is LoRA training captions, use the tag preset and keep prompts short - that's a dataset-format decision, not a node setting.
If you're already inside this pack, this is the captioner that scales with you: free and fast when you need tags, heavy and smart when you need a real description. Just bring a big-enough drive and a first-run coffee.
Inputs (21)
| Name | Type | Default | Description |
|---|---|---|---|
| 🤖 模型选择 | COMBO | BLIP-Large (推荐-精度高) | 选择图片理解模型 |
| ⚙️ 量化级别 | COMBO | None (FP16) | 量化级别,降低显存占用 |
| 💭 预设提示词 | COMBO | 提示词风格 - 详细 | 选择预设的提示词模板 |
| ✏️ 自定义提示词 | STRING | — | |
| 🔢 最大令牌数 | INT | 123464–4096 | 生成文本的最大令牌数 |
| 🌡️ 采样温度 | FLOAT | 0.60.1–1 | 创造性控制,值越高越随机 |
| 🎯 核采样参数 | FLOAT | 0.900–1 | 核采样参数 |
| 🔍 束搜索数量 | INT | 11–10 | 束搜索数量,值越大越准确但越慢 |
| 🚫 重复惩罚 | FLOAT | 1.200–2 | 重复惩罚,防止生成重复内容 |
| 🎬 视频帧数 | INT | 161–64 | 视频帧数(用于视频输入的采样) |
| 💻 设备选择 | COMBO | auto | 设备选择 |
| 🚀 开启TF32加速 | BOOLEAN | false | 启用TF32加速(仅支持Ampere及以上架构显卡,如30/40/50系,能显著提升速度) |
| 🔄 保持模型加载 | BOOLEAN | true | 是否保持模型加载(推荐开启) |
| 🎲 随机种子 | INT | -1-1–18446744073709550000 | 随机种子,-1为随机 |
| 🎯 种子控制 | COMBO | 随机 | 种子控制模式 |
| 🖼️ 图像1opt | IMAGE | 输入图片1 | |
| 🖼️ 图像2opt | IMAGE | 输入图片2 | |
| 🖼️ 图像3opt | IMAGE | 输入图片3 | |
| 🖼️ 图像4opt | IMAGE | 输入图片4 | |
| 🎥 视频opt | IMAGE | 视频输入(作为图像序列) | |
| 🎯 Qwen3VL额外选项opt | QWEN3VL_EXTRA_OPTIONS | 可选的Qwen3VL额外选项,连接Qwen3VL额外选项节点 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| 描述文本 | STRING | — |