ComfyUI Node

Qwen Image 2.1 Prompt Enhancer

It writes the picture, and hands you the ratio separately

By T8mars·Created 2 months ago·Updated 2 days ago· 347
Qwen Image 2.1 Prompt Enhancer
  • reference_images
  • provider_config
  • rewritten_prompt
  • wh_ratio
  • qwen_image_request_json
  • enhancement_report_json
◄prompt►
◄input_mode文生图 / Text-to-image►
◄max_output_chars0►
◄wh_ratioauto►
◄transparent_alphafalse►
◄api_mode贞贞平价小屋(推荐)►
◄ai_workshop_modelgemini-3.5-flash►
◄api_key►
◄custom_model►
◄openai_base_url►
◄seed0►
◄local_modelQwen3.8-27B-Q4_K_M.gguf►
◄local_mmprojAUTO(自动匹配)►
◄local_context_size32768►
◄local_max_tokens16384►
◄local_think_mode关闭(推荐,速度优先)►
◄local_reasoning_effortmedium►
◄local_video_sample_fps2.00►
◄local_unload_policy执行后卸载(推荐)►
◄local_comfy_memory_policyAUTO(显存不足时释放)►
◄recovery_slot►
◄recovery_actionnormal►
◄rewrite_profile经典兼容 / Classic►

What this node actually is

It's a text node. It doesn't generate an image, doesn't download Qwen weights, and has nothing to do with the Qwen Image diffusion model on your drive. You give it a rough brief - plus up to ten reference images if you're editing - and it hands that to a chat LLM pinned to a frozen "Image Prompt Rewriting Expert" skill, which returns one long English paragraph describing the finished frame, plus the aspect ratio as a separate field.

Yes, the name is confusing: "Qwen Image 2.1" is the model you're writing for, not the model doing the writing. The default cloud route runs qwen/qwen3.8-flash-next.

Why you'd reach for it: prompt-enhancer nodes went from novelty to routine, and the category has two well-documented failure modes - chat scaffolding leaking into your prompt, and the enhancer inventing detail you never asked for. This node's answer to both is the skill: describe the finished image as if you were looking at it, and pin the parts of your brief that must survive verbatim.

How it works, mechanically

The skill lives in the repo at official_skills/qwen-image-2.1/SKILL.md with a checked hash, so your output contract can't drift under you. The node ships it alongside four per-run rules: which mode you're in, how many ordered reference images there are, how to treat the ratio, and whether transparency is required.

Anything you fixed - visible text strings, named objects, counts, stated colours, stated positions - is copied character for character. Everything you left open, the model decides and states. Notes like "4K, no noise" are obeyed silently and never echoed as picture content.

Two choices worth understanding. The ratio lives only in the wh_ratio field and is deliberately never repeated in the prose, so a downstream width/height picker gets one clean value. And cloud requests get an 8192-token output budget by default, because the model's hidden reasoning shares it - a long skill can otherwise eat the whole allowance before a word of description is written.

The inputs you'll actually set

prompt is required, and it can be converted to an input socket to take an upstream text node. input_mode is either 文生图 / Text-to-image or 图像编辑 / Image edit - the node rejects text mode with images connected, and edit mode with none. reference_images is an autogrow socket taking 1–10 ordered IMAGE inputs; it never silently drops one.

max_output_chars defaults to 0, meaning the model picks a Skill-compliant length. A non-zero value triggers at most one bounded correction, and the node never cuts a sentence locally - if it's still over, the full result is kept and the report flags it. wh_ratio is auto or one of 1:1 / 3:4 / 4:3 / 16:9 / 9:16 / 2:1 / 1:2; auto lets the skill pick its 3:2 or 2:3 default and report the choice back. transparent_alpha is off unless you need RGBA and a transparent background. For the channel, api_mode covers ZhenZhen Affordable AI Shop (the default), AI Workshop (gemini-3.5-flash), an OpenAI-compatible endpoint, or local GGUF, which needs no key at all.

Outputs and where they go

rewritten_prompt is the one that matters - wire it into whatever the downstream image or edit node calls its prompt. wh_ratio is a STRING that belongs in your ratio or resolution picker, not in the prompt. The other two are diagnostics: what was asked, and whether the response was structured, how many corrections ran, and whether it blew your cap.

Installing it

Search the Manager for MiniMax H3 / Seedance 2.0 / Music 3 Prompt Enhancer (T8), or install from Git:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-prompt-enhancer-T8.git

Restart ComfyUI, then Ctrl+F5 the browser - the front end ships with the package and a stale tab shows you the old node. Cloud mode adds no dependencies beyond what ComfyUI ships. Local GGUF is the real lift: a GGUF in ComfyUI/models/LLM, a matching mmproj for image editing, and a llama runtime.

If the Registry copy is pending or flagged, ComfyUI-Manager can quietly reinstall an older "active" version and still report success. Update from GitHub instead, and never keep two copies of the pack in custom_nodes.

Where people get burned

Pasting the ratio into the brief. The fixed ratio is intentionally not repeated in the prose - but if you write "16:9" in your text, it becomes a fixed fact and it will surface in the description. Use the dropdown.

Truncated output is the other one. The generation budget is shared with the model's reasoning, so the fix is a bigger budget, not a shorter brief. And if the model returns non-JSON even after one correction, the text is kept in rewritten_prompt, structured_response goes false and wh_ratio comes back empty - check the report before assuming nothing ran.

Last thing: on cloud routes your brief and up to ten images leave your machine. Keys belong in SEEDANCE_API_KEY / T8STAR_API_KEY / OPENAI_API_KEY, since saving one into the node writes it into the workflow JSON.

CategoryT8/Qwen Image

Inputs (25)

NameTypeDefaultDescription
promptSTRING—
input_modeCOMBO文生图 / Text-to-image2 options: 文生图 / Text-to-image, 图像编辑 / Image edit
max_output_charsINT00–120000 让模型自行决定完整长度;非零会执行一次有界长度纠正,超出时保留完整结果并在报告标记。
wh_ratioCOMBOauto8 options: auto, 1:1, 3:4, 4:3, 16:9, 9:16, +2
transparent_alphaBOOLEANfalse—
api_modeCOMBO贞贞平价小屋(推荐)4 options: 贞贞平价小屋(推荐), 贞贞的AI工坊(图片/视频), OpenAI兼容接口(备用), 本地 GGUF(llama.cpp / Qwen,离线)
ai_workshop_modelCOMBOgemini-3.5-flash2 options: gemini-3.5-flash, Custom(自定义)
reference_imagesoptCOMFY_AUTOGROW_V3—
api_keyoptSTRING—
custom_modeloptSTRING—
openai_base_urloptSTRING—
seedoptINT00–18446744073709550000—
local_modeloptCOMBOQwen3.8-27B-Q4_K_M.gguf1 options: Qwen3.8-27B-Q4_K_M.gguf
local_mmprojoptCOMBOAUTO(自动匹配)2 options: AUTO(自动匹配), mmproj-F16.gguf
local_context_sizeoptINT327688192–65536—
local_max_tokensoptINT16384256–61440—
local_think_modeoptCOMBO关闭(推荐,速度优先)2 options: 关闭(推荐,速度优先), 开启(质量优先)
local_reasoning_effortoptCOMBOmedium3 options: low, medium, xhigh
local_video_sample_fpsoptFLOAT2.000.25–8—
local_unload_policyoptCOMBO执行后卸载(推荐)3 options: 执行后卸载(推荐), 保持驻留, 空闲10分钟后卸载
local_comfy_memory_policyoptCOMBOAUTO(显存不足时释放)2 options: AUTO(显存不足时释放), 不主动释放 ComfyUI 模型
recovery_slotoptSTRING—
recovery_actionoptSTRINGnormal—
provider_configoptT8_LLM_PROVIDER_CONFIG—
rewrite_profileoptCOMBO经典兼容 / Classic默认经典兼容,旧工作流不变。编辑专用只改指定部分,多图说明来源,中文需求默认中文;仅图像编辑生效。

Outputs (4)

NameTypeDescription
rewritten_promptSTRING—
wh_ratioSTRING—
qwen_image_request_jsonSTRING—
enhancement_report_jsonSTRING—