XB-BOX - 🖼️ Qwen2.1提示词预设
The text encoder you already loaded is a working LLM
- images
- 提示词
- 空latent
The "text encoder" on modern image models isn't a little CLIP sidecar any more. On the Qwen-Image line it's a full vision-language model, and when ComfyUI loads it as CLIP it still holds the whole thing in VRAM. It can write.
This node exploits that. Click a few style/view/subject chips, type "a girl in a coat on a pier", hit run - and the encoder rewrites it into a long structured prompt, locally, offline, no API key, no Ollama, no second download. Mechanically it's ComfyUI's native Generate Text path, minus the assembly work.
What you're actually getting
Two halves in one node. The preset half is the same panel the pack's XB_ImagePromptPreset uses: 15 preset modes, split into a text-to-image group (三视图/四视图/五视图 character sheets, 背景纯透明, 图文版面, 信息图, 多格分镜, 广告分镜板) and an image-to-image group (保持主体换场景, 局部编辑, 老照片修复, 整图风格化, 360°全景, 多图指认合成) - plus an eight-category element palette (风格 / 视角 / 主体 / 姿态 / 装扮 / 道具 / 光影 / 背景).
The generator half is the official Load CLIP and Generate Text path, copied field-for-field.
The pairing is the point. prompt-engineering.md argues that on an LLM-encoded model your prompt is an instruction, not a token bag, so having a language model compose it is the right shape. The preset half keeps the composer honest: the output is the preset's 设定词 placed verbatim at the top, with the generated body underneath. That fixed header is your anchor against enhancer drift.
The mechanism, briefly
Your text goes in as 【正文】, the preset text as 【参考设定】; both are tokenized together with the encoder. Generation runs at your sampling settings, the result is decoded (a <think> block, if any, is stripped), and the node assembles 设定词 (untouched) + generated body with the right framing clauses. Two outputs:
- 提示词 (STRING) - wire this into your text encode node, or into an edit encoder on the image side.
- 空latent (LATENT) - wire this into the sampler. On the image-to-image path, don't: the official edit chain is input image → VAEEncode → sampler.
No negative prompt here, and that's correct: negatives are inert on LLM encoders.
The fields worth touching
Most inputs hide behind two popup buttons (🤖 LLM设置, ✨ 预设参数). The ones you'll actually set:
- clip_name / clip_type / device - the Load CLIP trio. Drop your text encoder in
ComfyUI/models/text_encoders, pick it, leave clip_type onqwen_image, device ondefault.cputrades speed for card space. A mismatched type is the classic load failure. - preset_mode and io_mode - pick a mode;
自动already checks whether an image is connected and picks the right framing (and the right SKILL file). Forget the reference image on a 图生图 mode and it logs a warning and falls back to plain text. - latent_kind / aspect_ratio / width / height / batch_size - nine latent kinds (Anima, Boogu, Flux2, Hunyuan, Krea2, Qwen-image, SD3, SDXL, Z-image), each with its own official size step, so width snaps to a legal value.
- max_length - the 512 default is thin for 图文版面, 信息图, 多格分镜 and 广告分镜板; the author's own guidance is 1024–2048 or your layout rules get truncated mid-sentence.
- sampling_mode / temperature / seed -
offis greedy decoding and ignores the sampler settings entirely. Want the same prompt twice? Useoff, or fix the seed.
The optional text input overrides and locks the on-node prompt box; images takes a batch as N reference images (up to 10, referenced by number inside the prompt).
Installing it
ComfyUI Manager, search XB_ToolBox, or:
cd ComfyUI/custom_nodes
git clone https://github.com/WJLUOXIAO/XB_ToolBox.git
Restart, then point it at your Qwen-Image text encoder. No models ship with this node, and no key is needed.
One caveat: the README claims "NO extra pip dependencies required! Plug and play." The repo's requirements.txt disagrees (PyAV, easyocr, opencv, onnxruntime, pythonnet), and nodes_audio_slicer.py imports av at module top, pulled in unguarded by __init__.py. If the pack never appears in your node list on a bare environment, that's usually why:
pip install -r ComfyUI/custom_nodes/XB_ToolBox/requirements.txt
Where people get burned
There's no bypass. Execute it and it generates, every queue, even if all you wanted was the latent. On a slow card that's real seconds - minutes with the long output modes.
The encoder cache holds one entry keyed on name/type/device, so alternating between two text encoders reloads each time. Selecting a SKILL file does nothing unless skill_mode is 手动, and SKILL does nothing at all when use_default_template is off - same as the official node ignoring system_prompt. JSON output with a rewritten_prompt field is the bundled skill files' contract, not a break. And on AMD/APU boxes where the iGPU enumerates first, HIP error: device kernel image is invalid means ComfyUI landed on the wrong device - relaunch with HIP_VISIBLE_DEVICES=1.
Inputs (30)
| Name | Type | Default | Description |
|---|---|---|---|
| latent_kind | COMBO | Qwen-image | 空latent类型:选你正在用的模型即可(自动适配形状/下采样/步长) |
| output_lang | COMBO | 中文 [ZH] | 输出语言:决定词表与设定词按哪种语言加载、生成文本的语言 |
| io_mode | COMBO | 自动 | 模版:自动 = 按有没有接参考图判断;文生图 / 图生图 = 手动指定。每个预设模式都有两套设定词,按它自动选用(模板切换会直接把设定词换成对应那套) |
| preset_mode | COMBO | 无预设 | 预设模式:无预设=不前置任何设定词,只输出正文;其余档位把该档设定词置顶到最终提示词最顶端(设定词分文生图 / 图生图两套) |
| three_view_text | STRING | 生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。 | 设定词:预设模式对应的设定文本;改过的按「模式 + 语言」记进节点,换模式 / 换语言都不会丢 |
| skill_mode | COMBO | 自动 | SKILL 模式:自动 = 按预设模式适配(图生图档 → prompt_edit,其余 → prompt_t2i);手动 = 用下面选中的那一个;不用 = SKILL 完全不生效(哪怕选了也不生效) |
| skill_name | COMBO | 不使用 | SKILL选择:support_llama/skills 里的 txt,整段作为模型自带的系统提示词(官方 Generate Text 的 system_prompt);仅在「SKILL 模式 = 手动」时生效;「使用内置模板」关闭时也不生效 |
| aspect_ratio | COMBO | Free | 画幅比例:Free=自由(仅按步长锁定);固定比例时按较大的一边反算另一边 |
| width | INT | 102416–16384 | 图片宽度(步长按「空latent类型」的官方最小步长 8/16/32) |
| height | INT | 102416–16384 | 图片高度(步长按「空latent类型」的官方最小步长 8/16/32) |
| batch_size | INT | 11–4096 | 一次生成的图片数量(空 latent 的 batch 维度) |
| clip_name | COMBO | 文本编码器(models/text_encoders)。与官方 Load CLIP 同一个列表。 | |
| clip_type | COMBO | qwen_image | 类型:与官方 Load CLIP 一致;Qwen-Image 系文本编码器选 qwen_image |
| device | COMBO | default | 设备:与官方 Load CLIP 一致;default=自动,cpu=放内存 |
| max_length | INT | 5121–32768 | Generate Text · max_length:最大生成长度 |
| sampling_mode | COMBO | on | Generate Text · Sampling Mode:on=按下面的采样参数生成;off=贪心解码,下面的采样参数不生效 |
| temperature | FLOAT | 0.70000.01–2 | Generate Text · temperature |
| top_k | INT | 640–1000 | Generate Text · top_k |
| top_p | FLOAT | 0.950–1 | Generate Text · top_p |
| min_p | FLOAT | 0.050–1 | Generate Text · min_p |
| repetition_penalty | FLOAT | 1.050–5 | Generate Text · repetition_penalty |
| presence_penalty | FLOAT | 0.000–5 | Generate Text · presence_penalty |
| seed | INT | 00–18446744073709550000 | Generate Text · seed:生成种子(节点表面可点 🔁 切换生成后控制) |
| thinking | BOOLEAN | false | Generate Text · thinking:模型支持时以思考模式生成(思考内容不会进入输出) |
| use_default_template | BOOLEAN | true | Generate Text · use_default_template:使用模型自带的系统提示词 / 对话模板 |
| mtp | COMBO | auto | Generate Text · mtp:多 token 预测投机解码;没有 MTP 权重时无效 |
| internal_prompt | STRING | 节点内拼装/编辑的提示词正文(随工作流保存) | |
| manager_settings | STRING | 本节点配置 JSON:元素面板 elements + 追加设定 extra_settings + 生成后控制 run | |
| textopt | STRING | 外接提示词(优先于节点上的提示词框;接上线后提示词框锁定) | |
| imagesopt | IMAGE | 外接图像:一批图 = N 张参考图(官方图像编辑最多 10 张,提示词里可引用「第 N 张图」)。用「📦 批量图像」等多图节点直接接这里即可。 |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| 提示词 | STRING | — |
| 空latent | LATENT | — |