Nodes/XB_ToolBox/XB-BOX - 🖼️ Qwen2.1提示词预设
ComfyUI Node

XB-BOX - 🖼️ Qwen2.1提示词预设

The text encoder you already loaded is a working LLM

By wjluoxiao·Created 6 months ago·Updated 5 days ago· 351
XB-BOX - 🖼️ Qwen2.1提示词预设
  • images
  • 提示词
  • 空latent
◄latent_kindQwen-image►
◄output_lang中文 [ZH]►
◄io_mode自动►
◄preset_mode无预设►
◄three_view_text生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。►
◄skill_mode自动►
◄skill_name不使用►
◄aspect_ratioFree►
◄width1024►
◄height1024►
◄batch_size1►
◄clip_name▾►
◄clip_typeqwen_image►
◄devicedefault►
◄max_length512►
◄sampling_modeon►
◄temperature0.7000►
◄top_k64►
◄top_p0.95►
◄min_p0.05►
◄repetition_penalty1.05►
◄presence_penalty0.00►
◄seed0►
◄thinkingfalse►
◄use_default_templatetrue►
◄mtpauto►
◄internal_prompt►
◄manager_settings►
◄text—►

The "text encoder" on modern image models isn't a little CLIP sidecar any more. On the Qwen-Image line it's a full vision-language model, and when ComfyUI loads it as CLIP it still holds the whole thing in VRAM. It can write.

This node exploits that. Click a few style/view/subject chips, type "a girl in a coat on a pier", hit run - and the encoder rewrites it into a long structured prompt, locally, offline, no API key, no Ollama, no second download. Mechanically it's ComfyUI's native Generate Text path, minus the assembly work.

What you're actually getting

Two halves in one node. The preset half is the same panel the pack's XB_ImagePromptPreset uses: 15 preset modes, split into a text-to-image group (三视图/四视图/五视图 character sheets, 背景纯透明, 图文版面, 信息图, 多格分镜, 广告分镜板) and an image-to-image group (保持主体换场景, 局部编辑, 老照片修复, 整图风格化, 360°全景, 多图指认合成) - plus an eight-category element palette (风格 / 视角 / 主体 / 姿态 / 装扮 / 道具 / 光影 / 背景).

The generator half is the official Load CLIP and Generate Text path, copied field-for-field.

The pairing is the point. prompt-engineering.md argues that on an LLM-encoded model your prompt is an instruction, not a token bag, so having a language model compose it is the right shape. The preset half keeps the composer honest: the output is the preset's 设定词 placed verbatim at the top, with the generated body underneath. That fixed header is your anchor against enhancer drift.

The mechanism, briefly

Your text goes in as 【正文】, the preset text as 【参考设定】; both are tokenized together with the encoder. Generation runs at your sampling settings, the result is decoded (a <think> block, if any, is stripped), and the node assembles 设定词 (untouched) + generated body with the right framing clauses. Two outputs:

  • 提示词 (STRING) - wire this into your text encode node, or into an edit encoder on the image side.
  • 空latent (LATENT) - wire this into the sampler. On the image-to-image path, don't: the official edit chain is input image → VAEEncode → sampler.

No negative prompt here, and that's correct: negatives are inert on LLM encoders.

The fields worth touching

Most inputs hide behind two popup buttons (🤖 LLM设置, ✨ 预设参数). The ones you'll actually set:

  • clip_name / clip_type / device - the Load CLIP trio. Drop your text encoder in ComfyUI/models/text_encoders, pick it, leave clip_type on qwen_image, device on default. cpu trades speed for card space. A mismatched type is the classic load failure.
  • preset_mode and io_mode - pick a mode; 自动 already checks whether an image is connected and picks the right framing (and the right SKILL file). Forget the reference image on a 图生图 mode and it logs a warning and falls back to plain text.
  • latent_kind / aspect_ratio / width / height / batch_size - nine latent kinds (Anima, Boogu, Flux2, Hunyuan, Krea2, Qwen-image, SD3, SDXL, Z-image), each with its own official size step, so width snaps to a legal value.
  • max_length - the 512 default is thin for 图文版面, 信息图, 多格分镜 and 广告分镜板; the author's own guidance is 1024–2048 or your layout rules get truncated mid-sentence.
  • sampling_mode / temperature / seed - off is greedy decoding and ignores the sampler settings entirely. Want the same prompt twice? Use off, or fix the seed.

The optional text input overrides and locks the on-node prompt box; images takes a batch as N reference images (up to 10, referenced by number inside the prompt).

Installing it

ComfyUI Manager, search XB_ToolBox, or:

cd ComfyUI/custom_nodes
git clone https://github.com/WJLUOXIAO/XB_ToolBox.git

Restart, then point it at your Qwen-Image text encoder. No models ship with this node, and no key is needed.

One caveat: the README claims "NO extra pip dependencies required! Plug and play." The repo's requirements.txt disagrees (PyAV, easyocr, opencv, onnxruntime, pythonnet), and nodes_audio_slicer.py imports av at module top, pulled in unguarded by __init__.py. If the pack never appears in your node list on a bare environment, that's usually why:

pip install -r ComfyUI/custom_nodes/XB_ToolBox/requirements.txt

Where people get burned

There's no bypass. Execute it and it generates, every queue, even if all you wanted was the latent. On a slow card that's real seconds - minutes with the long output modes.

The encoder cache holds one entry keyed on name/type/device, so alternating between two text encoders reloads each time. Selecting a SKILL file does nothing unless skill_mode is 手动, and SKILL does nothing at all when use_default_template is off - same as the official node ignoring system_prompt. JSON output with a rewritten_prompt field is the bundled skill files' contract, not a break. And on AMD/APU boxes where the iGPU enumerates first, HIP error: device kernel image is invalid means ComfyUI landed on the wrong device - relaunch with HIP_VISIBLE_DEVICES=1.

CategoryXB_ToolBox/Image_Params

Inputs (30)

NameTypeDefaultDescription
latent_kindCOMBOQwen-image空latent类型:选你正在用的模型即可(自动适配形状/下采样/步长)
output_langCOMBO中文 [ZH]输出语言:决定词表与设定词按哪种语言加载、生成文本的语言
io_modeCOMBO自动模版:自动 = 按有没有接参考图判断;文生图 / 图生图 = 手动指定。每个预设模式都有两套设定词,按它自动选用(模板切换会直接把设定词换成对应那套)
preset_modeCOMBO无预设预设模式:无预设=不前置任何设定词,只输出正文;其余档位把该档设定词置顶到最终提示词最顶端(设定词分文生图 / 图生图两套)
three_view_textSTRING生成平行排列的角色概念设计图,画面从左到右由四个独立面板组成:第一个面板是角色面部的精细特写肖像,第二个面板是人物正面全身站姿,第三个面板是人物侧面全身站姿,第四个面板是人物背面全身站姿。设定词:预设模式对应的设定文本;改过的按「模式 + 语言」记进节点,换模式 / 换语言都不会丢
skill_modeCOMBO自动SKILL 模式:自动 = 按预设模式适配(图生图档 → prompt_edit,其余 → prompt_t2i);手动 = 用下面选中的那一个;不用 = SKILL 完全不生效(哪怕选了也不生效)
skill_nameCOMBO不使用SKILL选择:support_llama/skills 里的 txt,整段作为模型自带的系统提示词(官方 Generate Text 的 system_prompt);仅在「SKILL 模式 = 手动」时生效;「使用内置模板」关闭时也不生效
aspect_ratioCOMBOFree画幅比例:Free=自由(仅按步长锁定);固定比例时按较大的一边反算另一边
widthINT102416–16384图片宽度(步长按「空latent类型」的官方最小步长 8/16/32)
heightINT102416–16384图片高度(步长按「空latent类型」的官方最小步长 8/16/32)
batch_sizeINT11–4096一次生成的图片数量(空 latent 的 batch 维度)
clip_nameCOMBO文本编码器(models/text_encoders)。与官方 Load CLIP 同一个列表。
clip_typeCOMBOqwen_image类型:与官方 Load CLIP 一致;Qwen-Image 系文本编码器选 qwen_image
deviceCOMBOdefault设备:与官方 Load CLIP 一致;default=自动,cpu=放内存
max_lengthINT5121–32768Generate Text · max_length:最大生成长度
sampling_modeCOMBOonGenerate Text · Sampling Mode:on=按下面的采样参数生成;off=贪心解码,下面的采样参数不生效
temperatureFLOAT0.70000.01–2Generate Text · temperature
top_kINT640–1000Generate Text · top_k
top_pFLOAT0.950–1Generate Text · top_p
min_pFLOAT0.050–1Generate Text · min_p
repetition_penaltyFLOAT1.050–5Generate Text · repetition_penalty
presence_penaltyFLOAT0.000–5Generate Text · presence_penalty
seedINT00–18446744073709550000Generate Text · seed:生成种子(节点表面可点 🔁 切换生成后控制)
thinkingBOOLEANfalseGenerate Text · thinking:模型支持时以思考模式生成(思考内容不会进入输出)
use_default_templateBOOLEANtrueGenerate Text · use_default_template:使用模型自带的系统提示词 / 对话模板
mtpCOMBOautoGenerate Text · mtp:多 token 预测投机解码;没有 MTP 权重时无效
internal_promptSTRING节点内拼装/编辑的提示词正文(随工作流保存)
manager_settingsSTRING本节点配置 JSON:元素面板 elements + 追加设定 extra_settings + 生成后控制 run
textoptSTRING外接提示词(优先于节点上的提示词框;接上线后提示词框锁定)
imagesoptIMAGE外接图像:一批图 = N 张参考图(官方图像编辑最多 10 张,提示词里可引用「第 N 张图」)。用「📦 批量图像」等多图节点直接接这里即可。

Outputs (2)

NameTypeDescription
提示词STRING—
空latentLATENT—