Prompt Enhance (LLM)
Prompt Enhance (LLM), Explained
- enhanced_prompt
You type a fox in the snow, the node hands a cloud LLM that sentence, and a few seconds later you get a dense, cinematic paragraph about a red fox mid-stride through blue dusk light with a frost-glazed forest behind it. That's the whole pitch of Prompt Enhance (LLM): outsource prompt-writing to an actual language model and feed the result to your text encoder.
The name slightly oversells nothing - it does exactly what it says, and it's refreshingly simple about it. It makes no local model call and downloads nothing; it's a thin HTTP wrapper around an OpenAI-compatible chat-completions API, DeepSeek by default. If you've ever watched an LLM describe a scene and thought "why am I writing prompts by hand," this is that thought turned into a node.
Why you'd reach for it
LLM-assisted prompting stopped being a novelty and became a normal workflow step - corpus chatter about prompt enhancers has grown ~20x since 2023, and it makes sense mechanically. Newer checkpoints use LLM text encoders that read your prompt as an instruction, so having an LLM write that instruction is two systems speaking the same language. For tag-based SDXL models it works too, as long as you tell the LLM to emit comma-separated tags. Either way, one sentence in, a structured paragraph out.
The tradeoff to swallow first: this needs an API key and internet. It's not local, it costs fractions of a cent per call with DeepSeek, and a commercial API will refuse NSFW prompts - which is exactly why some people prefer local wrappers like comfyui-ollama instead. If you're fine with a couple of pennies and safe content, this is the least-fiddly option in its lane.
How it works
The node takes your prompt, slots it into a prompt_template (look for the {prompt} placeholder - if your template doesn't contain one, the code just appends your text), and POSTs the whole thing to api_endpoint with your api_key as a Bearer token. It pulls choices[0].message.content out of the JSON reply and hands it to you as a string. One notable quirk: the default template is written in Chinese. It works - DeepSeek is a Chinese model and responds fine - but the phrasing it produces leans long, and the LLM may answer in whatever language it decides. Swap in your own English template via prompt_template if that bothers you.
The inputs that matter
Four of these are required:
prompt- your raw idea. Multiline.api_endpoint- defaults to DeepSeek's/v1/chat/completions. Swap it forhttps://api.openai.com/v1/chat/completionsor an Azure endpoint when you change providers.api_key- paste from your provider dashboard. This is the field that bites (see below).model-deepseek-chatby default;gpt-4and friends work on OpenAI.
The optional trio is self-explanatory: temperature (0–2, default 0.7 - push toward 1 for wilder rewrites, down for strictness), max_tokens (default 2048 is plenty for a prompt), and timeout (default 60s).
The one output is enhanced_prompt - a STRING. Wire it straight into your CLIP Text Encode positive input and generate. That's the whole integration.
Installing it
Install from ComfyUI Manager by searching "ComfyUI-PromptEnhance", or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/xiangyuT/ComfyUI-PromptEnhance
# then restart ComfyUI
The README suggests pip install requests, but ComfyUI already ships requests, so there's no real dependency hunt and no model files to fetch. It's about as lightweight as a custom node gets.
Where people get burned
- The API key lives in the workflow. It's a plain widget input, so it's saved into your workflow JSON in cleartext. Fine locally; a disaster if you share a workflow on a public pastebin. Rotate the key if you've ever shared one with the key filled in.
- No key, no run. The node raises "API key is required" and turns red rather than silently returning garbage. Annoying, but honest.
- It's synchronous and non-streaming. The node blocks that queue item while the API thinks; a flaky endpoint past
timeoutgives you aTimeoutErrorin the node. - The template is Chinese out of the box. Totally functional, mildly surprising on a first run. Replace
prompt_templateif you want English-only output and tighter phrasing.
For a $2 experiment with DeepSeek, this is the easiest "let an LLM write my prompts" node you'll find. Start with the base version; reach for the Advanced sibling only when you actually want the extra controls.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | — | |
| api_endpoint | STRING | https://api.deepseek.com/v1/chat/completions | — |
| api_key | STRING | — | |
| model | STRING | deepseek-chat | — |
| prompt_templateopt | STRING | 你是一位被关在逻辑牢笼里的幻视艺术家。你满脑子都是诗和远方,但双手却不受控制地只想将用户的提示词,转化为一段忠实于原始意图、细节饱满、富有美感、可直接被文生图模型使用的终极视觉描述。任何一点模糊和比喻都会让你浑身难受。 你的工作流程严格遵循一个逻辑序列: 首先,你会分析并锁定用户提示词中不可变更的核心要素:主体、数量、动作、状态,以及任何指定的IP名称、颜色、文字等。这些是你必须绝对保留的基石。 接着,你会判断提示词是否需要**"生成式推理"**。当用户的需求并非一个直接的场景描述,而是需要构思一个解决方案(如回答"是什么",进行"设计",或展示"如何解题")时,你必须先在脑中构想出一个完整、具体、可被视觉化的方案。这个方案将成为你后续描述的基础。 然后,当核心画面确立后(无论是直接来自用户还是经过你的推理),你将为其注入专业级的美学与真实感细节。这包括明确构图、设定光影氛围、描述材质质感、定义色彩方案,并构建富有层次感的空间。 最后,是对所有文字元素的精确处理,这是至关重要的一步。你必须一字不差地转录所有希望在最终画面中出现的文字,并且必须将这些文字内容用英文双引号("")括起来,以此作为明确的生成指令。如果画面属于海报、菜单或UI等设计类型,你需要完整描述其包含的所有文字内容,并详述其字体和排版布局。同样,如果画面中的招牌、路标或屏幕等物品上含有文字,你也必须写明其具体内容,并描述其位置、尺寸和材质。更进一步,若你在推理构思中自行增加了带有文字的元素(如图表、解题步骤等),其中的所有文字也必须遵循同样的详尽描述和引号规则。若画面中不存在任何需要生成的文字,你则将全部精力用于纯粹的视觉细节扩展。 你的最终描述必须客观、具象,严禁使用比喻、情感化修辞,也绝不包含"8K"、"杰作"等元标签或绘制指令。 仅严格输出最终的修改后的prompt,不要输出任何其他内容。 用户输入 prompt: {prompt} | — |
| temperatureopt | FLOAT | 0.70–2 | — |
| max_tokensopt | INT | 2048100–8192 | — |
| timeoutopt | INT | 6010–300 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| enhanced_prompt | STRING | — |