Nodes/XB_ToolBox/XB-llama - ✨ 提示词增强反推
ComfyUI Node

XB-llama - ✨ 提示词增强反推

The prompt writer without the rest — local or API, text or image

By wjluoxiao·Created 6 months ago·Updated 5 days ago· 351
XB-llama - ✨ 提示词增强反推
  • images
  • 提示词全量
  • 提示词列表
  • 提示词设定
◄backend本地模型►
◄presetZ-Image Turbo [ZH]►
◄task_presetNormal - 描述 [ZH]►
◄open_api_settingsfalse►
◄custom_prompt►
◄manager_settings{}►
◄text—►

XB-llama - ✨ 提示词增强反推 (that's the display name - search the class name, not the menu label) is the LLM step on its own. No latent, no element panel, no geometry: text, an image, or both in, a prompt out. If the pack's all-in-one preset node is too much node for what you're doing, this is the piece you wanted.

You reach for it when the rest of the graph is already built and the only thing missing is the blank-page problem. Two jobs, and the community converged on both: enhance a rough idea into a model-appropriate prompt, and caption an existing image so you can iterate on it. The reason to run it locally rather than in a browser tab isn't quality - an 8B model doesn't out-write a frontier API - it's that local is uncensored, offline and free per call.

How it works

The node is a thin shell over the pack's inference engine, so the interesting part is what it assembles. It builds a system prompt from preset - 38 enhancer presets, one per model family and language: Z-Image Turbo, the Flux.2 line including both Klein variants, Qwen-Image 2512 and Edit 2511, Krea 2, Anima, Animagine, Wan and LTX video, MiniMax, Ideogram, Pony v6, storyboard. These aren't vague vibe instructions, they're dialect packs. The Z-Image one spells out that bracket weight syntax like (word:1.5) is forbidden, because DiT renders brackets as literal characters, and tells the model to write flowing prose instead. Same rule the prompt-engineering consensus states from the other side: on an LLM-encoded model, bracket weights are never applied.

task_preset is the second half: 23 instruction presets - description, tags, detailed, extreme-detailed, cinematic, creative analysis and story modes, refine-and-expand, and vision presets like bounding-box detection. Learn the author's convention: a * is a placeholder you must fill, a # is where your text gets dropped.

backend chooses the engine. 本地模型 runs your own GGUF through llama-cpp-python; 在线 API uses the provider/model/key/URL in the config dialog, covering the OpenAI-compatible field (OpenAI, DeepSeek, Qwen, GLM, Kimi, Ollama, vLLM, LM Studio) plus Anthropic and Gemini. Those settings persist in ComfyUI's user directory rather than the workflow - so sharing a graph doesn't share your key, but they are plaintext on disk.

Local config is VRAM-aware: give it a vram_limit and it computes how many GGUF layers fit, from the model's layer count. force_offload frees that memory after the run - the thing that makes an LLM node viable next to a diffusion model on one card. save_states keeps conversation state on disk so you can iterate without re-feeding context.

Inputs and outputs that matter

Required fields: backend, preset, task_preset, open_api_settings (a compatibility toggle - the API window lives in the 🤖 model-config dialog now), custom_prompt (your hand-typed text, or the placeholder for a * preset) and manager_settings (a JSON blob the frontend writes; don't touch it).

Optional: text and images. text is where automation lives - wire a string in and the node's own prompt box locks, because the port wins. images is the vision port: a captioner feeding this node is the assistant pattern image-to-video and remix workflows all run on.

Three outputs:

  • 提示词全量 - the text, straight into CLIP Text Encode
  • 提示词列表 - same content as a list, for consumers that want lines
  • 提示词设定 - the system prompt actually used; the fastest way to find out what a preset really says while you tune it

Install

cd ComfyUI/custom_nodes
git clone https://github.com/WJLUOXIAO/XB_ToolBox

Manager → search XB_ToolBox works too. Restart afterwards.

Models go in ComfyUI/models/LLM/, which the pack registers for .gguf, .safetensors and .bin files, so the loader sees whatever you drop there. A vision-language model also needs its mmproj and a matching chat handler.

The one dependency you must install by hand is llama-cpp-python, because the pack deliberately leaves it commented out of requirements.txt with a note pointing at the prebuilt Nvidia CUDA wheels (match your Python version) or a ROCm compile. On the API path you need nothing at all.

Where people get burned

"未选择本地模型." You flipped to a local backend without choosing a model in the config dialog. The error message says so explicitly.

"chat_handler 不能为 None! (加载了 mmproj 视觉模块)". Real error string, real cause: you loaded a vision projector but no matching handler. VLM reverse is a two-part download, not one.

A llama-cpp-python version complaint. Your wheel is older than the handlers the node expects; the error points you at the updated releases page. Reinstall it.

Chat scaffolding in your output. Preamble, assistant delimiters, markdown fences and reasoning traces are the classic failure of every LLM-in-graph node. This one adds hard output rules and strips thinking, but a chatty model still leaks sometimes. Pick small and obedient over big and clever - rewriting is the job, and it's why an 8B model is usually the right pick.

CategoryXB-llama

Inputs (8)

NameTypeDefaultDescription
backendCOMBO本地模型本地模型 = 用下面的模型/mmproj 跑 llama-cpp-python 在线 API = 用「⚙️ 打开 API 设置…」里配置的服务(存本地,不进工作流)
presetCOMBOZ-Image Turbo [ZH]提示词增强预设(原「✨ 提示词增强预设」) 其内容作为 system_prompt 使用,并原样从 system_prompt 端口输出
task_presetCOMBONormal - 描述 [ZH]反推任务预设(原「💬 指令推理」的预设提示词) 带 * 的预设里,* 表示必填占位符;带 # 的预设会把「文本」填入该位置
open_api_settingsBOOLEANfalse点击后弹出 API 设置窗口(provider / 模型名 / API Key / 地址 / 温度 / Token) 配置保存在 ComfyUI user 目录,不会写入工作流
custom_promptSTRING—
manager_settingsSTRING{}—
textoptSTRING外接提示词(优先于节点上的手写提示词编辑框) 接上线后节点上的编辑框会锁定不可编辑
imagesoptIMAGE外接图像 / 视频帧(VLM 反推:图 → 提示词)

Outputs (3)

NameTypeDescription
提示词全量STRING—
提示词列表STRING—
提示词设定STRING—