XB-BOX - Image Prompt Preset
Prompt, Ratio and Empty Latent in XB-BOX
- 提示词
- 空latent
What it replaces
Every text-to-image workflow opens with three separate things: a prompt box, a resolution, and an EmptyLatentImage for the sampler. Change the ratio and forget the latent and you get a stretched or silently resized image. XB_ImagePromptPreset is all three in one, from XB_ToolBox ("XB-BOX") - a Chinese-authored beginner toolbox, young enough (v0.9.01) that there's essentially no community footprint to judge it by. So judge the code, which is short.
The two outputs
提示词 / "Prompt" (STRING) and 空latent / "Empty Latent" (LATENT). String into your positive CLIP Text Encode, latent into your sampler. There's no negative output - fine on CFG-1 distilled models like Z-Image Turbo or Flux 2 Klein, where the negative pass isn't computed at all, and a reason to encode one separately on SDXL.
The latent dropdown is the interesting part
latent_kind covers nine families - Z-image (the default), Flux2, Qwen-image, Krea2, Anima, Boogu, SDXL, SD3, Hunyuan - and each carries a real spec, not a label: channel count, downsample divisor, size step. SDXL is 4-channel/÷8, step 8. Z-Image, Qwen-Image, Krea2 and SD3 are 16-channel/÷8, step 16. Anima and Boogu are 16-channel, step 8. Flux 2 is 128-channel/÷16. Hunyuan is 64-channel/÷32.
The mechanism: an empty latent is a zeros tensor shaped [batch, channels, height÷divisor, width÷divisor], plus one piece of metadata - downscale_ratio_spacial - that ComfyUI's comfy/sample.py reads to decide whether your latent needs rescaling to fit the model you actually load. Channel count isn't cosmetic: a latent from the wrong VAE family gives noise or flat colour, not a subtly wrong picture.
Because it's all zeros, though, a wrong pick usually won't crash - ComfyUI pads the channels and rescales the size to your checkpoint. Match the dropdown to what you're actually loading.
Sizes snap, on purpose
width/height default to 1024 and step in 16s on the widget, but the effective step follows the latent kind: 8 for SDXL/Anima/Boogu, 32 for Hunyuan, 16 for the rest. aspect_ratio (Free, 1:1, 16:9, 9:16, 4:3, 3:4, 21:9) only locks the step when it's Free; a fixed ratio anchors on the larger of your two numbers, computes the other side, then snaps it. Type 1000, get 1008 - VAE divisibility, not a bug. batch_size (default 1) is the batch dimension: four images, one prompt, four seeds.
One caveat that still applies: SDXL's trained buckets are 1024², 1152×896, 1216×832, 1344×768, 1536×640, and a 16:9 snap at a 1024 anchor gives 1024×576, outside them. Generate on-bucket and crop. On Z-Image, Flux 2 and Anima the discrete-bucket constraint has loosened a lot.
The prompt side
output_lang (中文/英文) sets the language of the preset sentence and the joiner that assembles the prompt. preset_mode is either 常规文生图 - your body text only, fully under your control - or 人物三视图, which prepends a character-sheet sentence asking for four panels: face close-up, then full-body front, side and back. That sentence lives in three_view_text, is editable, and only applies in three-view mode. It's the turnaround-sheet practice, the usual way to lock a character's angles and build LoRA datasets.
The panel is a row of category buttons (style, angle, subject, pose, outfit, props, lighting, background), each opening a chip list that assembles into your prompt. Your body text lives in internal_prompt, the same string the panel's preview box shows - edit either, the other follows. manager_settings stores the chips per node, so two of these don't interfere. Both fields are advanced, so they sit in the node's properties view rather than on its face.
A craft note, because a picklist isn't a prompt: on SDXL-lineage and Anima models you want booru tags, which is exactly the shape the chip output gives you. On Z-Image, Flux 2 and Qwen-Image the encoder is an LLM reading your prompt as an instruction, so use the chips as a skeleton and write a sentence around them.
Install
ComfyUI Manager, search XB_ToolBox, restart. Manually:
cd ComfyUI/custom_nodes
git clone https://github.com/WJLUOXIAO/XB_ToolBox.git
The README's "NO extra pip dependencies required!" is only true of the core nodes. The pack's requirements.txt pulls PyAV, opencv-python, easyocr, onnxruntime, pythonnet and more, and __init__.py imports the whole node set inside one try/except - so a single missing dependency doesn't cost you one node, it prints 🚨 [XB-BOX] FATAL ERROR: XB_ToolBox Loading Failed! and registers nothing. Install Manager's way, which runs requirements for you. Nothing to download for this node - it builds the latent itself.
Common issues
- Prompt box looks empty. Your text is in the advanced
internal_promptfield. If the panel fails to render, restart and hard-refresh the browser tab; you can still edit that value from the node's properties. - Half the menu is Chinese. The pack ships
locales/zhandlocales/en, but only some nodes are translated. Search "XB" in the add-node menu. - My prompt starts with a paragraph.
preset_modeis still on 人物三视图. - The size changed after queuing. Wrong
latent_kind: ComfyUI rescaled the latent to fit the loaded model.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| latent_kind | COMBO | Z-image | 空latent类型:选你正在用的模型即可(自动适配形状/下采样/步长)。 Z-image 16ch·/8·16 Flux2 128ch·/16·16 Qwen-image 16ch·/8·32 Krea2 16ch·/8·16 Anima 16ch·/8·8 Boogu 16ch·/8·8 SDXL 4ch·/8·8 SD3 16ch·/8·16 Hunyuan 64ch·/32·32 (步长 = 各模型官方最小步长,仅 Qwen-image 按需求锁 32;通道/下采样同样按官方 latent_format) |
| output_lang | COMBO | 中文 [ZH] | 输出语言:影响预设句、元素拼装分隔符(面板里新加入的词条也按此语言) |
| preset_mode | COMBO | 无预设 | 预设模式:无预设=不前置任何设定词,只输出正文(完全可控);人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明 = 成句自动在正文前加对应预设句(下方预设句框仅这三个模式显示,切模式时未改过的默认句会自动跟随) |
| three_view_text | STRING | 生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。 | 预设句 / 设定词(人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明 四个模式生效;无预设时本框自动隐藏) · 用户改过的设定词按「模式 + 语言」存进本节点(换模式 / 换语言都不会丢); · 最终提示词输出时它会被原封不动加在正文最顶端 |
| aspect_ratio | COMBO | Free | 画幅比例:Free=自由(仅按步长锁定);固定比例时以输入中较大的一边为基准,自动反算另一边并锁定到步长倍数(与「图片参数大全」完全一致) |
| width | INT | 102416–16384 | 图片宽度(像素)。固定比例下改宽度会按比例重算高度;步长按「空latent类型」的官方最小步长(SDXL/Anima/Boogu=8,Flux2/Krea2/SD3/Z-image=16,Qwen-image/Hunyuan=32) |
| height | INT | 102416–16384 | 图片高度(像素)。固定比例下改高度会按比例重算宽度;步长按「空latent类型」的官方最小步长(8/16/32) |
| batch_size | INT | 11–4096 | 一次生成的图片数量(空 latent 的 batch 维度) |
| internal_prompt | STRING | 节点内编辑的提示词正文(与元素面板的预览框双向同步,随工作流保存)。人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明模式下,本字段前面会自动拼上对应预设句 | |
| manager_settings | STRING | 本节点独立保存的元素面板配置 JSON(已选词条 / 自建槽位 / 补充描述),随工作流保存,节点间互不影响 |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| 提示词 | STRING | — |
| 空latent | LATENT | — |