Nodes/XB_ToolBox/XB-BOX - 🖼️ 生图提示词预设Pro
ComfyUI Node

XB-BOX - 🖼️ 生图提示词预设Pro

Prompt presets, a correct empty latent, and an optional LLM — in one node

By wjluoxiao·Created 6 months ago·Updated 5 days ago· 351
XB-BOX - 🖼️ 生图提示词预设Pro
  • images
  • 提示词
  • 空latent
  • 提示词列表
  • 提示词设定
◄latent_kindZ-image►
◄output_lang中文 [ZH]►
◄preset_mode无预设►
◄three_view_text生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。►
◄skill_mode自动►
◄skill_name不使用►
◄aspect_ratioFree►
◄width1024►
◄height1024►
◄batch_size1►
◄use_llmfalse►
◄backend本地模型►
◄presetZ-Image Turbo [ZH]►
◄task_presetNormal - 描述 [ZH]►
◄open_api_settingsfalse►
◄internal_prompt►
◄manager_settings►
◄text—►

The menu name is Chinese - XB-BOX - 🖼️ 生图提示词预设Pro - which is why you probably googled the class name instead. It's three jobs in one node: it assembles your prompt, builds a correct empty latent for whichever 2026 model you're running, and can optionally hand the whole thing to a local or remote LLM for enhancement or image reverse-description.

2026 prompting has two traps: every model has its own dialect, and every model has its own geometry. Z-Image, Flux 2 Klein, Anima and Krea 2 are LLM-encoded, so bracket weights are literal punctuation. And the latent your sampler wants is 4-channel on SDXL, 16 on most of the new crowd, 128 on Flux 2, 64 on Hunyuan.

The preset half (works with the LLM switched off)

latent_kind is the important one: pick the model you're actually generating with and the node uses that model's official latent format - channels, downsampling divisor, minimum step, batch ceiling - then builds a real empty latent. In the pack's own source the table is Z-image 16ch/8/step 16, Flux2 128ch/16/16, Qwen-image 16ch/8/32, Krea2 16ch/8/16, Anima 16ch/8/8, Boogu 16ch/8/8, SDXL 4ch/8/8, SD3 16ch/8/16, Hunyuan 64ch/32/32. Pick wrong and ComfyUI's fix_empty_latent_channels still rescues you by adjusting channels and scale - you get a picture, not an error, just with a pointless conversion in the middle.

aspect_ratio, width, height and batch_size are the usual geometry controls, but with the grid step taken from the latent type rather than a hard-coded 16.

preset_mode and three_view_text are the interesting pair. On 常规文生图 (plain text-to-image) your body text passes straight through. The three-, four- and five-view character modes and 背景纯透明 prepend a preset sentence describing a turnaround sheet - front/side/back panels plus a face close-up - or an RGBA transparent-background image. Edits to it are stored per mode and per language, so switching modes doesn't wipe them - the character-sheet pattern the consistency crowd has run for years, now a dropdown.

output_lang (中文/英文) governs the preset sentence, the element-panel separators and, with the LLM on, the language you get back. The panel itself is an 8-category chip picker in the frontend - style, camera, subject, pose, outfit, props, lighting, background.

The LLM half - and no, it isn't lying when it's off

use_llm defaults to false, and in the source both the model load and the API call sit inside that branch. Off means off - no GGUF, no request, no key. Turn it on and the node builds a system prompt from preset (38 model-specific enhancer presets in EN and ZH - Z-Image Turbo, Flux.2 Klein, Qwen-Image, Krea, Wan), your language, and task_preset (23 reverse/instruction presets, where * is a required placeholder and # is where your text lands), then demands only the final prompt back.

The task changes with what you feed it. Text only: enhance. Image only: caption. Image plus text: the text becomes an edit instruction, and the node forbids "originally X, now Y" comparison prose - it spots that phrasing and retries once. Anyone who has chained an edit model off a VLM knows why.

backend picks the engine: 本地模型 runs llama-cpp-python on your own GGUF; 在线 API uses whatever the config dialog saved (OpenAI-compatible - DeepSeek, Qwen, GLM, Kimi, Ollama, vLLM, LM Studio - plus Anthropic and Gemini). Keys live in ComfyUI's user directory, not the workflow: great for sharing, plaintext on disk, so treat them as secrets.

Outputs: 提示词 (the final prompt → CLIP Text Encode), 空latent (→ your sampler), 提示词列表 (one line per entry, for list consumers), 提示词设定 (the system prompt that ran - your debugging window). Optional inputs: text (external prompt; wired in, it wins and greys out the node's own box) and images for the vision path.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/WJLUOXIAO/XB_ToolBox

Or Manager → search XB_ToolBox. The preset half needs nothing beyond torch and ComfyUI itself.

For the local half you install llama-cpp-python yourself - it's deliberately commented out of the pack's requirements.txt. Nvidia users take a prebuilt CUDA wheel matching their Python version from the JamePeng releases; ROCm users compile. Put the GGUF (and, for VLM reverse-description, its matching mmproj) in ComfyUI/models/LLM/. The config's vram_limit then works out how many layers fit on your card.

Where people get burned

"未选择本地模型" on first run. The error itself tells you to open the 🤖 model-config dialog and pick a model. That dialog state lives in the node's internal settings JSON, so a half-configured node travels with your workflow.

Trusting the pack README over the pack. "No extra pip dependencies, plug and play" isn't true of the whole toolbox anymore, and neither prompt/LLM node appears in the README at all - everything above comes from the shipped source. It's also an LLM node, so it's expected to touch the network and load weights: skim a fresh pack before its first run.

CategoryXB_ToolBox/Image_Params

Inputs (19)

NameTypeDefaultDescription
latent_kindCOMBOZ-image空latent类型:选你正在用的模型即可(自动适配形状/下采样/步长)
output_langCOMBO中文 [ZH]输出语言:影响预设句、元素拼装分隔符,同时作为 LLM 的输出语言要求
preset_modeCOMBO无预设预设模式:无预设=不前置任何设定词,只输出正文;其余档位把该档设定词置顶到最终提示词最顶端(本节点固定文生图,只有 4 档)
three_view_textSTRING生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。预设模式的设定词(人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明 四个模式生效;无预设时自动隐藏) · 最终提示词输出时它会被【原封不动】加在最顶端; · 用户改过的设定词按「模式 + 语言」存进本节点(换模式 / 换语言都不会丢); · 启用 LLM 时它作为增强参考交给 LLM(LLM 只增强正文,不改写它)
skill_modeCOMBO自动SKILL 模式(面板上不再直接暴露三态):面板是「不使用 skill / 选文档」二选一;自动 = 用该档推荐文档(本节点固定 prompt_t2i),手动 = 用选中的那份,不用 = 不生效
skill_nameCOMBO不使用SKILL选择:support_llama/skills 里的 txt,整段作为 LLM 反推的 system prompt(放在最前面当角色与总规则);面板上是「不使用 skill / 选文档」二选一;技能文件自带语言与输出契约,选中后不再叠加节点的语言提示
aspect_ratioCOMBOFree画幅比例:Free=自由(仅按步长锁定);固定比例时按较大的一边反算另一边
widthINT102416–16384图片宽度(步长按「空latent类型」的官方最小步长 8/16/32)
heightINT102416–16384图片高度(步长按「空latent类型」的官方最小步长 8/16/32)
batch_sizeINT11–4096一次生成的图片数量(空 latent 的 batch 维度)
use_llmBOOLEANfalse关(默认)= 只输出拼装好的提示词 + 空latent(不加载模型、不调 API) 开 = 把(外接「提示词」或节点上的提示词)+ 外接「图像」交给 LLM 反推
backendCOMBO本地模型本地模型 = llama-cpp-python;在线 API = 在「🤖 大语言模型配置」弹窗里填服务商/模型/Key/地址
presetCOMBOZ-Image Turbo [ZH]增强预设(= 提示词设定 / system prompt),从「提示词设定」端口原样输出
task_presetCOMBONormal - 描述 [ZH]反推预设(原「指令推理」的预设提示词);带 * 的预设里 * 是必填占位符
open_api_settingsBOOLEANfalse(兼容保留)API 设置已内嵌在「🤖 大语言模型配置」弹窗里,无需单独打开
internal_promptSTRING节点内拼装/编辑的提示词正文(随工作流保存)
manager_settingsSTRING本节点配置 JSON:元素面板(elements)+ LLM 反推(llm),随工作流保存
textoptSTRING外接提示词(优先于节点上的提示词框;接上线后提示词框锁定)
imagesoptIMAGE外接图像 / 视频帧(VLM 反推:图 → 提示词)

Outputs (4)

NameTypeDescription
提示词STRING—
空latentLATENT—
提示词列表STRING—
提示词设定STRING—