XB-BOX - 🖼️ 生图提示词预设Pro
Prompt presets, a correct empty latent, and an optional LLM — in one node
- images
- 提示词
- 空latent
- 提示词列表
- 提示词设定
The menu name is Chinese - XB-BOX - 🖼️ 生图提示词预设Pro - which is why you probably googled the class name instead. It's three jobs in one node: it assembles your prompt, builds a correct empty latent for whichever 2026 model you're running, and can optionally hand the whole thing to a local or remote LLM for enhancement or image reverse-description.
2026 prompting has two traps: every model has its own dialect, and every model has its own geometry. Z-Image, Flux 2 Klein, Anima and Krea 2 are LLM-encoded, so bracket weights are literal punctuation. And the latent your sampler wants is 4-channel on SDXL, 16 on most of the new crowd, 128 on Flux 2, 64 on Hunyuan.
The preset half (works with the LLM switched off)
latent_kind is the important one: pick the model you're actually generating with and the node uses that model's official latent format - channels, downsampling divisor, minimum step, batch ceiling - then builds a real empty latent. In the pack's own source the table is Z-image 16ch/8/step 16, Flux2 128ch/16/16, Qwen-image 16ch/8/32, Krea2 16ch/8/16, Anima 16ch/8/8, Boogu 16ch/8/8, SDXL 4ch/8/8, SD3 16ch/8/16, Hunyuan 64ch/32/32. Pick wrong and ComfyUI's fix_empty_latent_channels still rescues you by adjusting channels and scale - you get a picture, not an error, just with a pointless conversion in the middle.
aspect_ratio, width, height and batch_size are the usual geometry controls, but with the grid step taken from the latent type rather than a hard-coded 16.
preset_mode and three_view_text are the interesting pair. On 常规文生图 (plain text-to-image) your body text passes straight through. The three-, four- and five-view character modes and 背景纯透明 prepend a preset sentence describing a turnaround sheet - front/side/back panels plus a face close-up - or an RGBA transparent-background image. Edits to it are stored per mode and per language, so switching modes doesn't wipe them - the character-sheet pattern the consistency crowd has run for years, now a dropdown.
output_lang (中文/英文) governs the preset sentence, the element-panel separators and, with the LLM on, the language you get back. The panel itself is an 8-category chip picker in the frontend - style, camera, subject, pose, outfit, props, lighting, background.
The LLM half - and no, it isn't lying when it's off
use_llm defaults to false, and in the source both the model load and the API call sit inside that branch. Off means off - no GGUF, no request, no key. Turn it on and the node builds a system prompt from preset (38 model-specific enhancer presets in EN and ZH - Z-Image Turbo, Flux.2 Klein, Qwen-Image, Krea, Wan), your language, and task_preset (23 reverse/instruction presets, where * is a required placeholder and # is where your text lands), then demands only the final prompt back.
The task changes with what you feed it. Text only: enhance. Image only: caption. Image plus text: the text becomes an edit instruction, and the node forbids "originally X, now Y" comparison prose - it spots that phrasing and retries once. Anyone who has chained an edit model off a VLM knows why.
backend picks the engine: 本地模型 runs llama-cpp-python on your own GGUF; 在线 API uses whatever the config dialog saved (OpenAI-compatible - DeepSeek, Qwen, GLM, Kimi, Ollama, vLLM, LM Studio - plus Anthropic and Gemini). Keys live in ComfyUI's user directory, not the workflow: great for sharing, plaintext on disk, so treat them as secrets.
Outputs: 提示词 (the final prompt → CLIP Text Encode), 空latent (→ your sampler), 提示词列表 (one line per entry, for list consumers), 提示词设定 (the system prompt that ran - your debugging window). Optional inputs: text (external prompt; wired in, it wins and greys out the node's own box) and images for the vision path.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/WJLUOXIAO/XB_ToolBox
Or Manager → search XB_ToolBox. The preset half needs nothing beyond torch and ComfyUI itself.
For the local half you install llama-cpp-python yourself - it's deliberately commented out of the pack's requirements.txt. Nvidia users take a prebuilt CUDA wheel matching their Python version from the JamePeng releases; ROCm users compile. Put the GGUF (and, for VLM reverse-description, its matching mmproj) in ComfyUI/models/LLM/. The config's vram_limit then works out how many layers fit on your card.
Where people get burned
"未选择本地模型" on first run. The error itself tells you to open the 🤖 model-config dialog and pick a model. That dialog state lives in the node's internal settings JSON, so a half-configured node travels with your workflow.
Trusting the pack README over the pack. "No extra pip dependencies, plug and play" isn't true of the whole toolbox anymore, and neither prompt/LLM node appears in the README at all - everything above comes from the shipped source. It's also an LLM node, so it's expected to touch the network and load weights: skim a fresh pack before its first run.
Inputs (19)
| Name | Type | Default | Description |
|---|---|---|---|
| latent_kind | COMBO | Z-image | 空latent类型:选你正在用的模型即可(自动适配形状/下采样/步长) |
| output_lang | COMBO | 中文 [ZH] | 输出语言:影响预设句、元素拼装分隔符,同时作为 LLM 的输出语言要求 |
| preset_mode | COMBO | 无预设 | 预设模式:无预设=不前置任何设定词,只输出正文;其余档位把该档设定词置顶到最终提示词最顶端(本节点固定文生图,只有 4 档) |
| three_view_text | STRING | 生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。 | 预设模式的设定词(人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明 四个模式生效;无预设时自动隐藏) · 最终提示词输出时它会被【原封不动】加在最顶端; · 用户改过的设定词按「模式 + 语言」存进本节点(换模式 / 换语言都不会丢); · 启用 LLM 时它作为增强参考交给 LLM(LLM 只增强正文,不改写它) |
| skill_mode | COMBO | 自动 | SKILL 模式(面板上不再直接暴露三态):面板是「不使用 skill / 选文档」二选一;自动 = 用该档推荐文档(本节点固定 prompt_t2i),手动 = 用选中的那份,不用 = 不生效 |
| skill_name | COMBO | 不使用 | SKILL选择:support_llama/skills 里的 txt,整段作为 LLM 反推的 system prompt(放在最前面当角色与总规则);面板上是「不使用 skill / 选文档」二选一;技能文件自带语言与输出契约,选中后不再叠加节点的语言提示 |
| aspect_ratio | COMBO | Free | 画幅比例:Free=自由(仅按步长锁定);固定比例时按较大的一边反算另一边 |
| width | INT | 102416–16384 | 图片宽度(步长按「空latent类型」的官方最小步长 8/16/32) |
| height | INT | 102416–16384 | 图片高度(步长按「空latent类型」的官方最小步长 8/16/32) |
| batch_size | INT | 11–4096 | 一次生成的图片数量(空 latent 的 batch 维度) |
| use_llm | BOOLEAN | false | 关(默认)= 只输出拼装好的提示词 + 空latent(不加载模型、不调 API) 开 = 把(外接「提示词」或节点上的提示词)+ 外接「图像」交给 LLM 反推 |
| backend | COMBO | 本地模型 | 本地模型 = llama-cpp-python;在线 API = 在「🤖 大语言模型配置」弹窗里填服务商/模型/Key/地址 |
| preset | COMBO | Z-Image Turbo [ZH] | 增强预设(= 提示词设定 / system prompt),从「提示词设定」端口原样输出 |
| task_preset | COMBO | Normal - 描述 [ZH] | 反推预设(原「指令推理」的预设提示词);带 * 的预设里 * 是必填占位符 |
| open_api_settings | BOOLEAN | false | (兼容保留)API 设置已内嵌在「🤖 大语言模型配置」弹窗里,无需单独打开 |
| internal_prompt | STRING | 节点内拼装/编辑的提示词正文(随工作流保存) | |
| manager_settings | STRING | 本节点配置 JSON:元素面板(elements)+ LLM 反推(llm),随工作流保存 | |
| textopt | STRING | 外接提示词(优先于节点上的提示词框;接上线后提示词框锁定) | |
| imagesopt | IMAGE | 外接图像 / 视频帧(VLM 反推:图 → 提示词) |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| 提示词 | STRING | — |
| 空latent | LATENT | — |
| 提示词列表 | STRING | — |
| 提示词设定 | STRING | — |