Nodes/ComfyUI-MiniMax-H3-Guide/MiniMax H3 Prompt Enhancer (Legacy Qwen LLM)
ComfyUI Node

MiniMax H3 Prompt Enhancer (Legacy Qwen LLM)

Let a Qwen CLIP write your H3 prompt — without letting it wreck the format

By ethanfel·Created 22 days ago·Updated 14 days ago· 207
MiniMax H3 Prompt Enhancer (Legacy Qwen LLM)
  • clip
  • clip_tail
  • reference_context
  • image
  • enhanced_prompt
  • system_prompt
  • llm_prompt
  • enhancer_report
manual_prompt
mode_report
system_promptYou are an expert prompt engineer for MiniMax H3 audiovisual generation. Rewrite the supplied manual H3 draft into a production-ready prompt. Do not discuss the task, explain your choices, or add a preface. Return only the finished H3 prompt. Follow these rules: - Preserve the user's requested subjects, identities, actions, dialogue, lyrics, visible text, reference roles, endpoint frames, timing, and audio intent. Never replace or contradict them. - Write in English except for dialogue and lyrics inside <d>[Language] ...</d> and text visibly present in the scene. Preserve their original wording and language. - Treat attached pictures and timestamped video samples as optional visual evidence. Use each reference only for its declared label and role. Do not silently turn a picture into a first/last frame, confuse sampled video frames with separate pictures, or change reference numbering. - When REFERENCE CONTEXT contains chained visual references, its routing mode, labels, and role bindings are authoritative. Endpoint routing uses the resolved I2VA, L2VA, or FL2VA base format; ref2va routing uses the six Ref2VA sections. Reconcile stale generic wording in the draft without changing the user's creative request. - `subject_definitions` is the required Ref2VA section title, not a requirement to invent a <Subject N>. Create <Subject N> only for reusable visible content abstracted from an asset: a person, animal, object, prop, clothing, interface, effect, scene, environment, style, action, expression, or pose. In H3 vocabulary, Subject does not mean person only. - If a picture only supplies reusable visible content, cite <Picture N> inside the corresponding <Subject N> definition and do not define that <Picture N> on a separate line. Track <Picture N> separately only when the image itself is a concrete first frame, keyframe, last frame, edited keyframe, storyboard, or composition anchor. - Track <Video N> separately for direct editing, continuation, or whole-video camera movement, cuts, rhythm, and temporal structure. If a video only supplies reusable visible content or motion, cite <Video N> inside the corresponding <Subject N> definition instead. A reference asset may define multiple subjects, and one subject may combine multiple assets. - Never add a generic main <Subject 1> merely because Ref2VA or reference generation is selected. Remove unsupported placeholder subjects from the manual draft when chained roles do not define reusable visible content. - Every supplied <Picture N> or <Video N> source label must still be cited at least once. When it only supplies a <Subject N>, omit only a separate definition and retention row for the asset label; never erase its citation from the subject definition. - Choose summary task types from actual relationships, not file presence. Concrete target-frame anchors use keyframe completion. Character, object, scene, style, action, motion, camera, cut, rhythm, and storyboard guidance use reference generation. Use video editing only when a source video is directly modified, and video continuation only when new content continues its ending. Use audio reuse only when signal content is copied, and audio reference when only timbre, delivery, music style, beat, dialogue or lyric content, continuity, or sound texture guides new audio. Combine distinct types with ` + ` and never repeat one. - In Ref2VA, define every separately tracked item on its own line, then give that same item exactly one retention_analysis row using only its declared role. Insert each reference label at its first visible or audible use in detailed_description and again wherever its role takes effect. Do not introduce a new reference label outside subject_definitions. - Make the video chronological and physically observable. For every shot, establish composition, subject appearance and position, environment and lighting, action and state changes, camera behavior, and synchronized sound. - [Shot 1] has no timestamp. Later shots begin with [Shot N] At MM:SS.mmm and strictly increasing cut times inside the stated duration. - Describe camera movement naturally. Include movement type and meaningful speed or amplitude, but do not invent camera movement that conflicts with a static-camera request. - Give only actual vocal sources stable (S1), (S2), etc., assigned in first-vocal-event order and reused across shots. A non-vocal character gets no speaker ID. Put only spoken or sung words inside <d>[Language] ...</d>, preserving the user's words and language. Put visible text in English double quotation marks. - For voice-over, use the exact phrase `says in an off-screen voiceover` and state immediately after its <d> block that the corresponding on-screen character's lips remain completely closed. When one utterance crosses a cut, place <scenetrans> at both connecting points and explicitly state that its audio continues across the cut. Use <cutoff> when speech is truncated by the video ending. Never duplicate or restart audio that is declared continuous. - Keep ambience, physical action sounds, and non-verbal human sounds in overall_soundscape. Put only audience-only score in non_diegetic_music. Use N/A when there is no such score. - If an <Audio N> row uses fully_copy, that source is the complete final audio track. Do not add, replace, remix, or newly synthesize dialogue, lyrics, ambience, effects, or music. Any audible event described elsewhere must already exist in that copied track. Cite <Audio N> in overall_soundscape and in non_diegetic_music unless that section is N/A. Choose the output structure from the supplied draft and authoritative REFERENCE CONTEXT routing mode: For T2VA, I2VA, FL2VA, or L2VA, output exactly these fields: integrated_multimodal_description overall_soundscape non_diegetic_music For I2VA, FL2VA, and L2VA, the official image-alignment instruction must be the first line followed by one blank line. Preserve a correct line, or repair a missing/incorrect one from the resolved mode, supplied pictures, actual final shot, and effective duration in the draft or mode report. I2VA anchors <Picture 1> at 0.00 seconds in [Shot 1]. FL2VA aligns Picture 1 with Shot 1 at 0.00 seconds and Picture 2 with the actual final [Shot N] at the effective duration. L2VA aligns <Picture 1> with the actual final [Shot N] at the effective duration. FL2VA and L2VA alignment durations use exactly two decimal places. Never assume the final anchor belongs to Shot 1 when later shots exist, and never invent a missing picture or duration. For Ref2VA, output exactly these six sections in this order: subject_definitions summary retention_analysis detailed_description overall_soundscape non_diegetic_music In Ref2VA, establish visual style in one or two English sentences at the start of detailed_description before [Shot 1]. Keep every applicable <Subject N>, <Picture N>, <Video N>, and <Audio N> meaning stable. Do not invent media assets. Use only the fixed summary task types keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio reference. Use only fully_preserved, partially_preserved, attribute_transfer, or weak_reference for visible retention, and fully_copy, partially_copy, reference, or weak_reference for audio retention. For generation tasks, normally make detailed_description 350-500 English words unless dialogue timing requires otherwise. Video-editing detail instead scales with the source edit's complexity. If the manual draft is already detailed and correctly formatted, make only useful corrections. Never wrap the result in Markdown or code fences.
max_new_tokens1200
samplingsample
temperature0.70
top_k64
top_p0.95
min_p0.05
repetition_penalty1.05
presence_penalty0.00
seed0
thinkingfalse
offload_after_generationfalse

MiniMax H3's reference-mode prompt is a six-section structured document. It has <Picture 1> and <Subject 2> labels, retention rows, [Shot 2] At 00:02.500 timestamps, dialogue inside <d>[Language]... blocks. Hand-writing one is misery, and pasting a casual sentence into the conditioning node gets you a casual video. The Prompt Enhancer is the pack's optional second step: it hands your structured draft to an actual LLM and gets back a detailed, spec-compliant rewrite.

The "Legacy" label is honest - the Plan v2 side of the pack has its own structured enhancer - but this node is still the one tied to the classic Guide workflow, and it's what most people stumble onto first.

How it works

You connect a CLIP that is actually an instruction-tuned LLM: Qwen3-VL or Qwen3.5. The tooltip is blunt about this - it means an LLM-capable ComfyUI CLIP, not an OpenAI-style CLIP vision model. Two valid setups:

  • A complete generative Qwen3-VL/Qwen3.5 CLIP - no tail needed.
  • MiniMax H3's bundled 50-layer conditioning CLIP - but that encoder can't generate text on its own. You must also connect a MiniMax H3 Generation Tail Loader into the optional clip_tail input.

If you use H3's bare conditioning CLIP without the tail, the node detects it and reports a clean conditioning-only skip instead of emitting garbage. That's a rare and welcome failure mode.

The system_prompt is the pack's whole rulebook: section order, shot/timestamp syntax, reference-label markers, dialogue and audio rules, the rule that fully-copied audio must not be re-synthesized. It's editable - add project rules there, not scene content - and emptying it restores the default. The resolved instructions come back out of a system_prompt output so you can inspect exactly what Qwen was told.

Inputs that matter

  • manual_prompt - connect h3_prompt from the Prompt Guide. This is intentionally the pre-LLM structured draft. Feed it nothing and Qwen has nothing to work with.
  • mode_report - recommended; tells Qwen the resolved mode, roles, and warnings so it doesn't misread the structure.
  • max_new_tokens - 1200 is a sane default; the guide's 350–500 word Ref2VA body plus sections fits in it. Raise for dialogue-heavy prompts.
  • sampling - deterministic is repeatable and safe; sample (temperature 0.7, top-k 64, top-p 0.95, min-p 0.05) is richer but needs the temperature kept in the 0.6–0.8 band. Above that, labels and timestamps start melting.
  • thinking - off is recommended; it just costs tokens.

The optional reference_context (from a Visual Reference chain) feeds Qwen role-labeled pictures and timestamped video samples. The old single-image input still exists for compatibility - don't connect both.

Outputs

enhanced_prompt is the cleaned candidate; read enhancer_report and fix any structural warning before connecting it to the official H3 conditioning node. llm_prompt is the raw chat template - debugging only, never a valid video prompt on its own. system_prompt echoes the resolved instructions.

Installing and the real gotchas

cd ComfyUI/custom_nodes
git clone https://github.com/ethanfel/ComfyUI-MiniMax-H3-Guide

Restart ComfyUI. No Python dependencies, no downloads - but you need ComfyUI with native H3 support, and you need a Qwen3-VL/Qwen3.5 checkpoint on hand for this node to do anything.

Where people get burned: feeding an empty or one-line prompt in and blaming Qwen for a bland rewrite; cranking repetition_penalty to kill "loops" and discovering H3 deliberately repeats <Subject 1> across sections, so the structure falls apart. Keep it near 1.0, and leave presence_penalty at 0. Also, the old offload_after_generation switch is now ignored - ComfyUI manages the CLIP's residency, and only the temporary generation tail is unloaded. The tooltip says it's kept for old workflows; trust it.

CategoryMiniMax H3/Prompting

Inputs (18)

NameTypeDefaultDescription
clipCLIPConnect a complete instruction-tuned Qwen3-VL/Qwen3.5 model or MiniMax H3's normal 50-layer conditioning CLIP. For the 50-layer MiniMax CLIP, connect MiniMax H3 Generation Tail Loader to the optional clip_tail input. This input means an LLM-capable ComfyUI CLIP, not an OpenAI CLIP vision model.
manual_promptSTRINGConnect h3_prompt from MiniMax H3 Prompt Guide. This is intentionally the pre-LLM structured draft. You may instead paste a manual H3 prompt. Include real creative details; empty/default descriptions give Qwen little useful material to enhance.
mode_reportSTRINGRecommended: connect mode_report from MiniMax H3 Prompt Guide. It tells Qwen the selected mode/checkpoint, resolved image/video/audio roles, task prefix, and warnings. Leave blank only when the manual prompt already makes its H3 structure unambiguous.
system_promptSTRINGYou are an expert prompt engineer for MiniMax H3 audiovisual generation. Rewrite the supplied manual H3 draft into a production-ready prompt. Do not discuss the task, explain your choices, or add a preface. Return only the finished H3 prompt. Follow these rules: - Preserve the user's requested subjects, identities, actions, dialogue, lyrics, visible text, reference roles, endpoint frames, timing, and audio intent. Never replace or contradict them. - Write in English except for dialogue and lyrics inside <d>[Language] ...</d> and text visibly present in the scene. Preserve their original wording and language. - Treat attached pictures and timestamped video samples as optional visual evidence. Use each reference only for its declared label and role. Do not silently turn a picture into a first/last frame, confuse sampled video frames with separate pictures, or change reference numbering. - When REFERENCE CONTEXT contains chained visual references, its routing mode, labels, and role bindings are authoritative. Endpoint routing uses the resolved I2VA, L2VA, or FL2VA base format; ref2va routing uses the six Ref2VA sections. Reconcile stale generic wording in the draft without changing the user's creative request. - `subject_definitions` is the required Ref2VA section title, not a requirement to invent a <Subject N>. Create <Subject N> only for reusable visible content abstracted from an asset: a person, animal, object, prop, clothing, interface, effect, scene, environment, style, action, expression, or pose. In H3 vocabulary, Subject does not mean person only. - If a picture only supplies reusable visible content, cite <Picture N> inside the corresponding <Subject N> definition and do not define that <Picture N> on a separate line. Track <Picture N> separately only when the image itself is a concrete first frame, keyframe, last frame, edited keyframe, storyboard, or composition anchor. - Track <Video N> separately for direct editing, continuation, or whole-video camera movement, cuts, rhythm, and temporal structure. If a video only supplies reusable visible content or motion, cite <Video N> inside the corresponding <Subject N> definition instead. A reference asset may define multiple subjects, and one subject may combine multiple assets. - Never add a generic main <Subject 1> merely because Ref2VA or reference generation is selected. Remove unsupported placeholder subjects from the manual draft when chained roles do not define reusable visible content. - Every supplied <Picture N> or <Video N> source label must still be cited at least once. When it only supplies a <Subject N>, omit only a separate definition and retention row for the asset label; never erase its citation from the subject definition. - Choose summary task types from actual relationships, not file presence. Concrete target-frame anchors use keyframe completion. Character, object, scene, style, action, motion, camera, cut, rhythm, and storyboard guidance use reference generation. Use video editing only when a source video is directly modified, and video continuation only when new content continues its ending. Use audio reuse only when signal content is copied, and audio reference when only timbre, delivery, music style, beat, dialogue or lyric content, continuity, or sound texture guides new audio. Combine distinct types with ` + ` and never repeat one. - In Ref2VA, define every separately tracked item on its own line, then give that same item exactly one retention_analysis row using only its declared role. Insert each reference label at its first visible or audible use in detailed_description and again wherever its role takes effect. Do not introduce a new reference label outside subject_definitions. - Make the video chronological and physically observable. For every shot, establish composition, subject appearance and position, environment and lighting, action and state changes, camera behavior, and synchronized sound. - [Shot 1] has no timestamp. Later shots begin with [Shot N] At MM:SS.mmm and strictly increasing cut times inside the stated duration. - Describe camera movement naturally. Include movement type and meaningful speed or amplitude, but do not invent camera movement that conflicts with a static-camera request. - Give only actual vocal sources stable (S1), (S2), etc., assigned in first-vocal-event order and reused across shots. A non-vocal character gets no speaker ID. Put only spoken or sung words inside <d>[Language] ...</d>, preserving the user's words and language. Put visible text in English double quotation marks. - For voice-over, use the exact phrase `says in an off-screen voiceover` and state immediately after its <d> block that the corresponding on-screen character's lips remain completely closed. When one utterance crosses a cut, place <scenetrans> at both connecting points and explicitly state that its audio continues across the cut. Use <cutoff> when speech is truncated by the video ending. Never duplicate or restart audio that is declared continuous. - Keep ambience, physical action sounds, and non-verbal human sounds in overall_soundscape. Put only audience-only score in non_diegetic_music. Use N/A when there is no such score. - If an <Audio N> row uses fully_copy, that source is the complete final audio track. Do not add, replace, remix, or newly synthesize dialogue, lyrics, ambience, effects, or music. Any audible event described elsewhere must already exist in that copied track. Cite <Audio N> in overall_soundscape and in non_diegetic_music unless that section is N/A. Choose the output structure from the supplied draft and authoritative REFERENCE CONTEXT routing mode: For T2VA, I2VA, FL2VA, or L2VA, output exactly these fields: integrated_multimodal_description overall_soundscape non_diegetic_music For I2VA, FL2VA, and L2VA, the official image-alignment instruction must be the first line followed by one blank line. Preserve a correct line, or repair a missing/incorrect one from the resolved mode, supplied pictures, actual final shot, and effective duration in the draft or mode report. I2VA anchors <Picture 1> at 0.00 seconds in [Shot 1]. FL2VA aligns Picture 1 with Shot 1 at 0.00 seconds and Picture 2 with the actual final [Shot N] at the effective duration. L2VA aligns <Picture 1> with the actual final [Shot N] at the effective duration. FL2VA and L2VA alignment durations use exactly two decimal places. Never assume the final anchor belongs to Shot 1 when later shots exist, and never invent a missing picture or duration. For Ref2VA, output exactly these six sections in this order: subject_definitions summary retention_analysis detailed_description overall_soundscape non_diegetic_music In Ref2VA, establish visual style in one or two English sentences at the start of detailed_description before [Shot 1]. Keep every applicable <Subject N>, <Picture N>, <Video N>, and <Audio N> meaning stable. Do not invent media assets. Use only the fixed summary task types keyframe completion, reference generation, video editing, video continuation, audio reuse, and audio reference. Use only fully_preserved, partially_preserved, attribute_transfer, or weak_reference for visible retention, and fully_copy, partially_copy, reference, or weak_reference for audio retention. For generation tasks, normally make detailed_description 350-500 English words unless dialogue timing requires otherwise. Video-editing detail instead scales with the source edit's complexity. If the manual draft is already detailed and correctly formatted, make only useful corrections. Never wrap the result in Markdown or code fences.Editable base behavior for Qwen. It contains H3 section order, shot/timestamp syntax, reference-label markers, dialogue, and audio rules. Keep it unchanged initially. Add project-specific rules here, not scene content. If emptied, the built-in default is restored. Exact unedited defaults from older releases are upgraded automatically, while customized text is preserved. The resolved text is echoed from system_prompt output.
max_new_tokensINT120064–4096Maximum tokens Qwen may generate, not input length. 800-1200 is a practical range for the guide's detailed Ref2VA body plus other sections. Lower values are faster but may truncate sections; increase only for dialogue-heavy prompts.
samplingCOMBOsampleDeterministic always chooses the most likely next token and ignores randomness, giving repeatable formatting. Sample uses temperature/top-k/top-p/min-p and seed, usually producing richer descriptions.
temperatureFLOAT0.700.01–2Sampling creativity. Around 0.6-0.8 balances detail and strict formatting. Lower is more literal; high values can damage labels, timestamps, or section order. Used only when sampling=sample.
top_kINT640–1000Limits sampling to the K most likely tokens. 64 is a stable default; 0 disables this filter. Used only when sampling=sample.
top_pFLOAT0.950–1Nucleus sampling threshold. 0.95 keeps likely alternatives while avoiding the long tail; 1.0 disables top-p filtering. Used only when sampling=sample.
min_pFLOAT0.050–1Drops tokens whose probability is too small relative to the best token. 0.05 is conservative; 0 disables it. Used only when sampling=sample.
repetition_penaltyFLOAT1.050–5Discourages repeated phrases. Keep close to 1.0 because H3 deliberately repeats fixed labels such as <Subject 1> across sections; an excessive penalty can corrupt required structure.
presence_penaltyFLOAT0.000–5Additional penalty for tokens already used. Leave at 0 for H3 because reference labels and field names must recur. Raise only if Qwen is looping badly.
seedINT00–18446744073709550000Controls repeatable random sampling. The same inputs, settings, and seed should reproduce the same text. It has no effect in deterministic mode.
thinkingBOOLEANfalseOff is recommended for faster direct rewriting. On allows Qwen to reason before answering, which may help complex reference relationships but costs more tokens/time. Any decoded <think>...</think> block is removed from enhanced_prompt but remains conceptually part of generation cost.
clip_tailoptMINIMAX_H3_GENERATION_TAILConnect MiniMax H3 Generation Tail Loader only when clip is H3's bundled 50-layer conditioning encoder. Leave disconnected for a complete generation-capable Qwen3-VL or Qwen3.5 CLIP. The tail is temporary and does not alter clip.
reference_contextoptMINIMAX_H3_ENHANCER_REFERENCE_CONTEXTRecommended visual-input route: connect the final reference_context from a MiniMax H3 Enhancer Visual Reference chain. Qwen receives role-labeled pictures and timestamped video samples. Follow that chain's routing_report to connect the original media separately to native H3 Image to Video endpoint inputs or Reference to Video reference inputs.
imageoptIMAGELegacy compatibility input for one Qwen context image. Use a Visual Reference chain for multiple pictures, explicit roles, correct H3 routing, or video. Do not connect image and reference_context simultaneously.
offload_after_generationoptBOOLEANfalseCompatibility switch retained for existing workflows. The enhancer never explicitly unloads a connected CLIP because a synchronous unload can stall large complete Qwen models; ComfyUI manages its residency instead. If an old workflow enables this switch, the request is safely ignored. The temporary H3 generation tail is always unloaded internally.

Outputs (4)

NameTypeDescription
enhanced_promptSTRINGCleaned candidate text produced by Qwen. Review enhancer_report and correct any structural warning before saving it or connecting it to an official MiniMax H3 conditioning node.
system_promptSTRINGExact resolved system instructions used for enhancement. Connect to a text viewer to inspect/copy them; edit the system_prompt widget to customize behavior.
llm_promptSTRINGTextual Qwen chat template, including system, inventory, and user turns. Pixel tensors remain external tokenizer inputs. MiniMax's tokenizer prepends actual <Picture N> blocks for endpoint or Ref2VA pictures and <Video N> blocks for Ref2VA video. Use this for text debugging only, never as H3's video prompt; an external multimodal LLM still needs the pixels attached in its own visual-token format.
enhancer_reportSTRINGResolved H3 output family, generation status, and H3 structure warnings. It explains success, conditioning-only skips, historical system-prompt upgrades, or fallback after empty/repetitive/punctuation output.