Nodes/XB_ToolBox/XB-BOX - Image Prompt Preset
ComfyUI Node

XB-BOX - Image Prompt Preset

Prompt, Ratio and Empty Latent in XB-BOX

By wjluoxiao·Created 6 months ago·Updated 5 days ago· 351
XB-BOX - Image Prompt Preset
    • 提示词
    • 空latent
    ◄latent_kindZ-image►
    ◄output_lang中文 [ZH]►
    ◄preset_mode无预设►
    ◄three_view_text生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。►
    ◄aspect_ratioFree►
    ◄width1024►
    ◄height1024►
    ◄batch_size1►
    ◄internal_prompt►
    ◄manager_settings►

    What it replaces

    Every text-to-image workflow opens with three separate things: a prompt box, a resolution, and an EmptyLatentImage for the sampler. Change the ratio and forget the latent and you get a stretched or silently resized image. XB_ImagePromptPreset is all three in one, from XB_ToolBox ("XB-BOX") - a Chinese-authored beginner toolbox, young enough (v0.9.01) that there's essentially no community footprint to judge it by. So judge the code, which is short.

    The two outputs

    提示词 / "Prompt" (STRING) and 空latent / "Empty Latent" (LATENT). String into your positive CLIP Text Encode, latent into your sampler. There's no negative output - fine on CFG-1 distilled models like Z-Image Turbo or Flux 2 Klein, where the negative pass isn't computed at all, and a reason to encode one separately on SDXL.

    The latent dropdown is the interesting part

    latent_kind covers nine families - Z-image (the default), Flux2, Qwen-image, Krea2, Anima, Boogu, SDXL, SD3, Hunyuan - and each carries a real spec, not a label: channel count, downsample divisor, size step. SDXL is 4-channel/÷8, step 8. Z-Image, Qwen-Image, Krea2 and SD3 are 16-channel/÷8, step 16. Anima and Boogu are 16-channel, step 8. Flux 2 is 128-channel/÷16. Hunyuan is 64-channel/÷32.

    The mechanism: an empty latent is a zeros tensor shaped [batch, channels, height÷divisor, width÷divisor], plus one piece of metadata - downscale_ratio_spacial - that ComfyUI's comfy/sample.py reads to decide whether your latent needs rescaling to fit the model you actually load. Channel count isn't cosmetic: a latent from the wrong VAE family gives noise or flat colour, not a subtly wrong picture.

    Because it's all zeros, though, a wrong pick usually won't crash - ComfyUI pads the channels and rescales the size to your checkpoint. Match the dropdown to what you're actually loading.

    Sizes snap, on purpose

    width/height default to 1024 and step in 16s on the widget, but the effective step follows the latent kind: 8 for SDXL/Anima/Boogu, 32 for Hunyuan, 16 for the rest. aspect_ratio (Free, 1:1, 16:9, 9:16, 4:3, 3:4, 21:9) only locks the step when it's Free; a fixed ratio anchors on the larger of your two numbers, computes the other side, then snaps it. Type 1000, get 1008 - VAE divisibility, not a bug. batch_size (default 1) is the batch dimension: four images, one prompt, four seeds.

    One caveat that still applies: SDXL's trained buckets are 1024², 1152×896, 1216×832, 1344×768, 1536×640, and a 16:9 snap at a 1024 anchor gives 1024×576, outside them. Generate on-bucket and crop. On Z-Image, Flux 2 and Anima the discrete-bucket constraint has loosened a lot.

    The prompt side

    output_lang (中文/英文) sets the language of the preset sentence and the joiner that assembles the prompt. preset_mode is either 常规文生图 - your body text only, fully under your control - or 人物三视图, which prepends a character-sheet sentence asking for four panels: face close-up, then full-body front, side and back. That sentence lives in three_view_text, is editable, and only applies in three-view mode. It's the turnaround-sheet practice, the usual way to lock a character's angles and build LoRA datasets.

    The panel is a row of category buttons (style, angle, subject, pose, outfit, props, lighting, background), each opening a chip list that assembles into your prompt. Your body text lives in internal_prompt, the same string the panel's preview box shows - edit either, the other follows. manager_settings stores the chips per node, so two of these don't interfere. Both fields are advanced, so they sit in the node's properties view rather than on its face.

    A craft note, because a picklist isn't a prompt: on SDXL-lineage and Anima models you want booru tags, which is exactly the shape the chip output gives you. On Z-Image, Flux 2 and Qwen-Image the encoder is an LLM reading your prompt as an instruction, so use the chips as a skeleton and write a sentence around them.

    Install

    ComfyUI Manager, search XB_ToolBox, restart. Manually:

    cd ComfyUI/custom_nodes
    git clone https://github.com/WJLUOXIAO/XB_ToolBox.git
    

    The README's "NO extra pip dependencies required!" is only true of the core nodes. The pack's requirements.txt pulls PyAV, opencv-python, easyocr, onnxruntime, pythonnet and more, and __init__.py imports the whole node set inside one try/except - so a single missing dependency doesn't cost you one node, it prints 🚨 [XB-BOX] FATAL ERROR: XB_ToolBox Loading Failed! and registers nothing. Install Manager's way, which runs requirements for you. Nothing to download for this node - it builds the latent itself.

    Common issues

    • Prompt box looks empty. Your text is in the advanced internal_prompt field. If the panel fails to render, restart and hard-refresh the browser tab; you can still edit that value from the node's properties.
    • Half the menu is Chinese. The pack ships locales/zh and locales/en, but only some nodes are translated. Search "XB" in the add-node menu.
    • My prompt starts with a paragraph. preset_mode is still on 人物三视图.
    • The size changed after queuing. Wrong latent_kind: ComfyUI rescaled the latent to fit the loaded model.
    CategoryXB_ToolBox/Image_Params

    Inputs (10)

    NameTypeDefaultDescription
    latent_kindCOMBOZ-image空latent类型:选你正在用的模型即可(自动适配形状/下采样/步长)。 Z-image 16ch·/8·16 Flux2 128ch·/16·16 Qwen-image 16ch·/8·32 Krea2 16ch·/8·16 Anima 16ch·/8·8 Boogu 16ch·/8·8 SDXL 4ch·/8·8 SD3 16ch·/8·16 Hunyuan 64ch·/32·32 (步长 = 各模型官方最小步长,仅 Qwen-image 按需求锁 32;通道/下采样同样按官方 latent_format)
    output_langCOMBO中文 [ZH]输出语言:影响预设句、元素拼装分隔符(面板里新加入的词条也按此语言)
    preset_modeCOMBO无预设预设模式:无预设=不前置任何设定词,只输出正文(完全可控);人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明 = 成句自动在正文前加对应预设句(下方预设句框仅这三个模式显示,切模式时未改过的默认句会自动跟随)
    three_view_textSTRING生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。预设句 / 设定词(人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明 四个模式生效;无预设时本框自动隐藏) · 用户改过的设定词按「模式 + 语言」存进本节点(换模式 / 换语言都不会丢); · 最终提示词输出时它会被原封不动加在正文最顶端
    aspect_ratioCOMBOFree画幅比例:Free=自由(仅按步长锁定);固定比例时以输入中较大的一边为基准,自动反算另一边并锁定到步长倍数(与「图片参数大全」完全一致)
    widthINT102416–16384图片宽度(像素)。固定比例下改宽度会按比例重算高度;步长按「空latent类型」的官方最小步长(SDXL/Anima/Boogu=8,Flux2/Krea2/SD3/Z-image=16,Qwen-image/Hunyuan=32)
    heightINT102416–16384图片高度(像素)。固定比例下改高度会按比例重算宽度;步长按「空latent类型」的官方最小步长(8/16/32)
    batch_sizeINT11–4096一次生成的图片数量(空 latent 的 batch 维度)
    internal_promptSTRING节点内编辑的提示词正文(与元素面板的预览框双向同步,随工作流保存)。人物三视图 / 人物四视图 / 人物五视图 / 背景纯透明模式下,本字段前面会自动拼上对应预设句
    manager_settingsSTRING本节点独立保存的元素面板配置 JSON(已选词条 / 自建槽位 / 补充描述),随工作流保存,节点间互不影响

    Outputs (2)

    NameTypeDescription
    提示词STRING—
    空latentLATENT—