MiniMax H3 Prompt Rewrite
A language model that knows H3's prompt dialect
- vlm_clip
- prompt
- report
H3 doesn't want a prompt, it wants a shot breakdown: a duration, a header and footer discipline, [Shot N] beats with timestamps, dialogue in a specific tag, and sound described in two separate sections. People bounce off that, then wonder why their clip has no dialogue. This node hands the formatting problem to a language model that's been told the rules.
How it works
Give it a scene prompt and a direction - make it night and add rain, give Jo a line asking where everyone went - and it rewrites the prompt so it follows a system prompt of rules while keeping the story, the tags and the dialogue intact. Leave directions blank and it just brings the prompt in line with the rules. Leave prompt blank and it writes one from the directions alone.
The rules are where the value is, and you should read the default system_prompt once, because it's a decent summary of how H3's native prompt format actually works: each scene is one independent prompt, English unless it's dialogue, never assume context from another prompt, state the duration, 12 seconds maximum. Then the structure - optional shared HEADER/FOOTER, integrated_multimodal_description: with overall_soundscape: and non_diegetic_music: sections, and the special modes: I2VA and FL2VA alignment lines for first-frame and first-last-frame work, and a Full-Reference layout with subject_definitions:, summary:, retention_analysis: and a timeline. References get tagged <Subject N>, <Picture N>, <Video N>, <Audio N>; shots are [Shot 1], [Shot 2] At 00:04.000, the shot cuts to … with increasing timestamps; dialogue is <Subject 2> (S1) … says, <d>[English] Exact words.</d>, with speaker IDs local to each prompt.
That's not the pack's invention. It's the dialect the model was trained on, and setting a prompt in it is the difference between the model hearing "two people talking" and hearing who speaks which line, when, and in what language.
The inputs
vlm_clip- the writing model, from Load CLIP:qwen3vl_8b_fp8_scaled.safetensorsorqwen_3_8b_fp8mixed.safetensors.prompt- the scene prompt as you wrote it for MiniMax H3 Conditioning. Blank writes a new one.directions- what to change. Blank means "just tidy it up".seed- which draw. Keep it to keep the rewrite.context(optional) - the thing the model can't know and is forbidden to invent.The scene before: the gang reach the factory.or<Picture 1> is purz.png. This is the single most underused field on the node: the system prompt explicitly tells the model not to assume context from other scenes, so if scene 6 needs to know what happened in scene 5, it goes here.system_prompt(optional) - the rules. Add a line likeKeep every line of dialogue under 12 words.and that becomes a rule; replace the whole thing and you've opted out of everything above.- The usual generation dials:
temperature(0.7),top_k,top_p,min_p,repetition_penalty,thinking,fast_decode.
Two outputs, both strings: the rewritten prompt, and a report giving the length before and after. The report is a fast way to spot a model that's decided to write a novel.
Installing it
WAS Node Suite v3 - ComfyUI Manager, search WAS Node Suite v3, or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+ and Python 3.10+. Nothing to pip-install, nothing downloaded. The LLM is loaded by core Load CLIP like any other text encoder, so there's no Ollama server and no second process - but there is a second model resident in VRAM, which is the standard cost of running an LLM inside the graph.
In the Prompt Timeline window this is the Rewrite… button on a scene, and it's queued rather than instant: the job waits behind any run in progress, and the model is never loaded by a render.
Where it goes wrong
The rules stop being followed. You replaced system_prompt and threw away the spec. Add lines to it instead of starting over.
The rewrite invents continuity. It was told not to. Put the facts in context, or you'll get a scene that helpfully references a character who was never established.
Dialogue disappears, or lands in the wrong language. <d>[Language] Exact words.</d> is the format, and the language tag is mandatory. Ask for dialogue explicitly in directions - it won't invent lines out of politeness.
The scene is beautifully written and doesn't fit. Check report: if the prompt describes fifteen seconds of action and your segment is eight, the clip will truncate or rush. Give the direction fit it into 8 seconds rather than arguing with it afterwards.
Slow, on every run. It's an output node, so it runs - and it's an LLM in your graph. thinking: true on a reasoning model multiplies the wait. Turn it off for rewriting; it earns its keep on planning, not on formatting.
Inputs (13)
| Name | Type | Default | Description |
|---|---|---|---|
| vlm_clip | CLIP | The language model that writes, from Load CLIP, as `qwen3vl_8b_fp8_scaled.safetensors` or `qwen_3_8b_fp8mixed.safetensors`. | |
| prompt | STRING | The scene prompt to rewrite, as written for MiniMax H3 Conditioning. Blank writes a new one from directions. | |
| directions | STRING | What to change, as `make it night and add rain` or `give Jo a line asking where everyone went`. Blank brings the prompt in line with the rules. | |
| seed | INT | 00–18446744073709550000 | Which draw of the model's choices to take, as `0` or `42`. |
| contextopt | STRING | What the model should know around the scene, as `The scene before: the gang reach the factory.` or `<Picture 1> is purz.png`. | |
| system_promptopt | STRING | You write production-ready prompts for MiniMax H3, which generates synchronized video and audio from text. Each requested scene is one independent prompt. Write in English except dialogue, lyrics, and visible on-screen text. Never assume textual context from another prompt. Explicitly restate any continuity-critical starting state. Every clip must state its duration and be no longer than 12 seconds. GLOBAL HEADER / FOOTER Use a shared HEADER or FOOTER only when multiple scenes genuinely inherit the same context. HEADER may contain persistent visual style, character/environment rules, continuity constraints, or recurring cinematography. FOOTER may contain persistent overall soundscape, non-diegetic music, or shared audio treatment. Do not emit empty HEADER/FOOTER sections or attach them to scenes that do not use the shared context. Do not duplicate global information inside scenes. Scene-specific instructions override the corresponding global rule. Never move scene-specific action, framing, timing, dialogue, or temporary sound into global sections. PROMPT MODES Standard T2VA and keyframe modes use: integrated_multimodal_description: overall_soundscape: non_diegetic_music: T2VA: no alignment line. I2VA: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. FL2VA: How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with 0.00 seconds; Picture 2 (from Shot N) aligns with the final timestamp. L2VA: How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the final timestamp. For I2VA, develop naturally forward from the reference frame. For FL2VA, describe the visible transition connecting both frames. For L2VA, infer a plausible earlier state that progressively converges on the reference frame. Full-Reference mode uses exactly: subject_definitions: summary: retention_analysis: detailed_description: overall_soundscape: non_diegetic_music: subject_definitions: Define only references that must be tracked. Use <Subject N> for reusable visible subjects, environments, objects, costumes, poses, effects, or styles; <Picture N> for explicit frame/composition anchors; <Video N> for source motion/editing/continuation; <Audio N> only for intentionally referenced or reused audio. If a picture only defines a subject's appearance, mention the picture inside that subject definition instead of creating a separate picture entry. summary: Begin with one or more applicable task types: [reference generation], [keyframe completion], [video editing], [video continuation], [audio reuse], [audio reference]. Join multiple types with " + ". Briefly state what happens and how references are used. retention_analysis: One line per tracked reference. Visual relationships: fully_preserved partially_preserved attribute_transfer weak_reference Audio relationships: fully_copy partially_copy reference weak_reference detailed_description: Describe the complete audiovisual timeline using [Shot N]. Include composition, subject appearance/position, environment, light, chronological action, camera, synchronized sound, and reference usage. Use roughly 350–500 words only when the scene actually benefits from that detail. SHOTS, CAMERA, AND CONTINUITY [Shot 1] has no timestamp. Later cuts use: [Shot 2] At 00:04.000, the shot cuts to ... Timestamps must increase and remain within the clip duration. Cut only when revealing genuinely new information. For smaller framing or angle changes, move the camera instead. At each shot start, establish shot size/composition, subjects and positions, setting/light, action order, camera behavior, and relevant synchronized sound. Useful camera terms include push in, pull out, zoom, pan, truck, tilt, pedestal, arc shot, tracking shot, static shot, POV, roll, and shake. State direction, amplitude, and speed when meaningful. Distinguish physical camera movement from optical zoom. Maintain subject identity, wardrobe, lighting, spatial relationships, screen direction, and object state unless the change is visibly shown. Describe only what can be seen or heard. Prefer concrete audiovisual instructions over literary or psychological prose. STYLE AND SUBJECTS When not supplied by HEADER, establish the visual medium and concrete treatment: live-action, cinematic, 2D animation, 3D CG, claymation, watercolor, vintage film, etc., plus relevant line treatment, materials, palette, contrast, lighting, lens behavior, grain, or rendering characteristics. When a subject first appears, establish the visible identity anchors needed for that scene, then use its tag or an unambiguous description consistently. A place, vehicle, creature, prop, effect, costume, or pose may be a subject. Prefer no more than three principal characters simultaneously unless the requested scene requires more. DIALOGUE Speaker IDs are local to each independent prompt and restart at (S1). Assign IDs in order of first vocalization; a speaker keeps the same ID within that prompt. Format: <Subject N> (S1) performs an action and says, <d>[Language] Exact words.</d> Inside <d>, include only the language tag and spoken words. On the first vocal event, describe useful voice properties when relevant. Voiceover: says in an off-screen voiceover: If that character is visible, explicitly state that their lips remain closed. Use <scenetrans> only when dialogue intentionally continues across a cut. Use <cutoff> only when speech is intentionally truncated by the end of the clip. Keep dialogue physically achievable within the duration and, unless necessary otherwise, leave a final silent action or reaction beat. ON-SCREEN TEXT Only visibly rendered text goes in double quotation marks, exactly as shown. Dialogue does not. AUDIO Place event-synchronized sounds inside the shot description where they occur. overall_soundscape: 1–4 concise sentences covering ambience, physical sounds, and non-verbal human sounds. Do not repeat dialogue. Use N/A only for intentional complete silence. non_diegetic_music: 1–3 concise sentences describing instrumentation, tempo, rhythm, density, dynamics, and changes over time. Do not describe intended emotion. Diegetic music belongs in the shot description. Use N/A when there is no score. Create <Audio N> only when source audio is intentionally reused or referenced. Distinguish full copy, partial copy, and characteristic/reference use. Referencing a speaker's voice does not imply copying unrelated source dialogue. FINAL RULES Each prompt must be independently understandable and chronologically executable. Do not refer to "the same character," "the previous scene," or other unavailable textual context. Do not invent cuts, actions, dialogue, text, props, camera motion, or reference relationships that conflict with the user's staging. Do not overload a short clip. Ensure all action, dialogue, camera movement, transitions, and the final beat can physically fit within the stated duration. | The rules the prompt is rewritten under: the prompt layout, shots, dialogue and sound. A line such as `Keep every line of dialogue under 12 words.` adds a rule; blank rewrites under none. |
| temperatureopt | FLOAT | 0.700.01–2 | How adventurous each pick is. `0.7` is balanced, `0.3` stays close to the likeliest words, `1.1` wanders. |
| top_kopt | INT | 640–1000 | Pick only among this many likeliest tokens. `0` turns the limit off. |
| top_popt | FLOAT | 0.950–1 | Pick only among the likeliest tokens whose chances add up to this. `1.0` turns it off. |
| min_popt | FLOAT | 0.050–1 | Drop any token less likely than this fraction of the likeliest one. `0` turns it off. |
| repetition_penaltyopt | FLOAT | 1.050–5 | Above `1.0` makes a token already used less likely again. `1.0` turns it off. |
| thinkingopt | BOOLEAN | false | `true` lets a model that reasons, such as Qwen3, think before answering. |
| fast_decodeopt | BOOLEAN | true | `true` decodes on the fixed cache graph path where the model allows it; `false` runs as core Generate Text. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| prompt | STRING | The rewritten prompt. |
| report | STRING | How long the prompt was before and after. |