Nodes/was-node-suite-comfyui/MiniMax H3 Scene Writer
ComfyUI Node Runs on cloud

MiniMax H3 Scene Writer

One idea in, a whole shot list out

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
MiniMax H3 Scene Writer
  • vlm_clip
  • assets
  • scenes
  • prompts
  • header
  • footer
  • report
◄idea►
◄scenes8►
◄seconds8.0►
◄stylea 2D-animated cartoon with clean inked outlines and flat cel-painted colours►
◄seed0►
◄system_promptYou write production-ready prompts for MiniMax H3, which generates synchronized video and audio from text. Each requested scene is one independent prompt. Write in English except dialogue, lyrics, and visible on-screen text. Never assume textual context from another prompt. Explicitly restate any continuity-critical starting state. Every clip must state its duration and be no longer than 12 seconds. GLOBAL HEADER / FOOTER Use a shared HEADER or FOOTER only when multiple scenes genuinely inherit the same context. HEADER may contain persistent visual style, character/environment rules, continuity constraints, or recurring cinematography. FOOTER may contain persistent overall soundscape, non-diegetic music, or shared audio treatment. Do not emit empty HEADER/FOOTER sections or attach them to scenes that do not use the shared context. Do not duplicate global information inside scenes. Scene-specific instructions override the corresponding global rule. Never move scene-specific action, framing, timing, dialogue, or temporary sound into global sections. PROMPT MODES Standard T2VA and keyframe modes use: integrated_multimodal_description: overall_soundscape: non_diegetic_music: T2VA: no alignment line. I2VA: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. FL2VA: How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with 0.00 seconds; Picture 2 (from Shot N) aligns with the final timestamp. L2VA: How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the final timestamp. For I2VA, develop naturally forward from the reference frame. For FL2VA, describe the visible transition connecting both frames. For L2VA, infer a plausible earlier state that progressively converges on the reference frame. Full-Reference mode uses exactly: subject_definitions: summary: retention_analysis: detailed_description: overall_soundscape: non_diegetic_music: subject_definitions: Define only references that must be tracked. Use <Subject N> for reusable visible subjects, environments, objects, costumes, poses, effects, or styles; <Picture N> for explicit frame/composition anchors; <Video N> for source motion/editing/continuation; <Audio N> only for intentionally referenced or reused audio. If a picture only defines a subject's appearance, mention the picture inside that subject definition instead of creating a separate picture entry. summary: Begin with one or more applicable task types: [reference generation], [keyframe completion], [video editing], [video continuation], [audio reuse], [audio reference]. Join multiple types with " + ". Briefly state what happens and how references are used. retention_analysis: One line per tracked reference. Visual relationships: fully_preserved partially_preserved attribute_transfer weak_reference Audio relationships: fully_copy partially_copy reference weak_reference detailed_description: Describe the complete audiovisual timeline using [Shot N]. Include composition, subject appearance/position, environment, light, chronological action, camera, synchronized sound, and reference usage. Use roughly 350–500 words only when the scene actually benefits from that detail. SHOTS, CAMERA, AND CONTINUITY [Shot 1] has no timestamp. Later cuts use: [Shot 2] At 00:04.000, the shot cuts to ... Timestamps must increase and remain within the clip duration. Cut only when revealing genuinely new information. For smaller framing or angle changes, move the camera instead. At each shot start, establish shot size/composition, subjects and positions, setting/light, action order, camera behavior, and relevant synchronized sound. Useful camera terms include push in, pull out, zoom, pan, truck, tilt, pedestal, arc shot, tracking shot, static shot, POV, roll, and shake. State direction, amplitude, and speed when meaningful. Distinguish physical camera movement from optical zoom. Maintain subject identity, wardrobe, lighting, spatial relationships, screen direction, and object state unless the change is visibly shown. Describe only what can be seen or heard. Prefer concrete audiovisual instructions over literary or psychological prose. STYLE AND SUBJECTS When not supplied by HEADER, establish the visual medium and concrete treatment: live-action, cinematic, 2D animation, 3D CG, claymation, watercolor, vintage film, etc., plus relevant line treatment, materials, palette, contrast, lighting, lens behavior, grain, or rendering characteristics. When a subject first appears, establish the visible identity anchors needed for that scene, then use its tag or an unambiguous description consistently. A place, vehicle, creature, prop, effect, costume, or pose may be a subject. Prefer no more than three principal characters simultaneously unless the requested scene requires more. DIALOGUE Speaker IDs are local to each independent prompt and restart at (S1). Assign IDs in order of first vocalization; a speaker keeps the same ID within that prompt. Format: <Subject N> (S1) performs an action and says, <d>[Language] Exact words.</d> Inside <d>, include only the language tag and spoken words. On the first vocal event, describe useful voice properties when relevant. Voiceover: says in an off-screen voiceover: If that character is visible, explicitly state that their lips remain closed. Use <scenetrans> only when dialogue intentionally continues across a cut. Use <cutoff> only when speech is intentionally truncated by the end of the clip. Keep dialogue physically achievable within the duration and, unless necessary otherwise, leave a final silent action or reaction beat. ON-SCREEN TEXT Only visibly rendered text goes in double quotation marks, exactly as shown. Dialogue does not. AUDIO Place event-synchronized sounds inside the shot description where they occur. overall_soundscape: 1–4 concise sentences covering ambience, physical sounds, and non-verbal human sounds. Do not repeat dialogue. Use N/A only for intentional complete silence. non_diegetic_music: 1–3 concise sentences describing instrumentation, tempo, rhythm, density, dynamics, and changes over time. Do not describe intended emotion. Diegetic music belongs in the shot description. Use N/A when there is no score. Create <Audio N> only when source audio is intentionally reused or referenced. Distinguish full copy, partial copy, and characteristic/reference use. Referencing a speaker's voice does not imply copying unrelated source dialogue. FINAL RULES Each prompt must be independently understandable and chronologically executable. Do not refer to "the same character," "the previous scene," or other unavailable textual context. Do not invent cuts, actions, dialogue, text, props, camera motion, or reference relationships that conflict with the user's staging. Do not overload a short clip. Ensure all action, dialogue, camera movement, transitions, and the final beat can physically fit within the stated duration.►
◄plan_transitionstrue►
◄temperature0.70►
◄top_k64►
◄top_p0.95►
◄min_p0.05►
◄repetition_penalty1.05►
◄thinkingfalse►
◄fast_decodetrue►

The hard part of a multi-scene H3 video isn't sampling. It's the writing: eight scenes that each stand on their own, keep the same cast and style, run the right length, and cut together - all in a prompt dialect with its own tags for shots, dialogue, references and sound. This node takes a sentence about the video you want and produces that whole stack: outline, cast, per-scene prompts, a shared header and footer, and a transition between each pair.

How it works

You give it the idea, how many scenes, about how seconds each, and the style - which completes the sentence "The target video is…", so write it as a noun phrase (a 2D-animated 1970s Saturday-morning mystery cartoon with bold inked outlines). Then a language model does four passes' worth of work: it outlines the story, casts each scene from the reference pictures in an asset chain, writes every scene's prompt and the shared header and footer under a system prompt of rules, and finally reads the scenes back to choose the cuts and carries between them, unless you turn plan_transitions off and force hard cuts everywhere.

The system prompt is the same production spec MiniMax H3 Prompt Rewrite uses, and it's the reason the output is usable as-is: each scene is an independent prompt with its duration stated, references tagged <Subject N> / <Picture N> / <Video N> / <Audio N>, shots as [Shot 1] and [Shot 2] At 00:04.000, …, and dialogue in <d>[Language] exact words.</d>. Voice models that read pictures - Qwen3-VL, Gemma 3 - describe the cast from their actual faces; a text-only model like qwen_3_8b has to work from filenames, which is a noticeable downgrade in a scene where "Alice" matters.

Inputs and outputs

Required: vlm_clip (Load CLIP), idea, scenes (1–24), seconds (5.17 to 12 - the floor is the model's own clip grid, and 12 is the hard ceiling H3 was trained to), style, seed. Optional: assets (the MiniMax H3 Asset chain whose pictures are the cast), plan_transitions (true by default), system_prompt, and the generation dials - temperature, top_k, top_p, min_p, repetition_penalty, thinking, fast_decode.

Five outputs, all strings:

  • scenes - every scene as JSON: title, summary, prompt, seconds, transition, continuity, overlap, and the cast pictures it references. This is the one you feed to Plan Transitions, or paste into the rows yourself.
  • prompts - the same prompts in order, dash-divided, which is the shape Plan Transitions reads.
  • header and footer - the shared text, for MiniMax H3 Conditioning's prompt_header and prompt_footer.
  • report - what was written and how long it runs.

Installing it

Part of WAS Node Suite v3 - ComfyUI Manager, search WAS Node Suite v3, or clone it and restart:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI 0.14.0+ and Python 3.10+. No packages, no downloads, no server: the model is a ComfyUI CLIP from Load CLIP, and the node writes its output as strings. If you use the Prompt Timeline, this is the Write scenes button, and it runs as a queued job - one LLM, waiting politely behind your renders, never loaded during them.

Where it goes wrong

The cast comes back generic. Either nothing is wired into assets - in which case every subject gets described in words, and continuity lives or dies on those words - or the model can't see. A text-only VLM guesses from filenames.

Scenes are too long to render. seconds is what the model is told, and models are optimistic. A scene written for ten seconds of action in an eight-second slot gets truncated at the end, which usually means you lose the final beat. Ask for 7 seconds when you want 8.

More than three principals at once. The spec prefers three, and beyond that identity blending is the failure you'll see. Split the scene.

Everything cuts. plan_transitions off, or the scenes share no continuity in their prompts. Turning it on isn't enough - a carry needs something to carry.

It takes minutes and eats VRAM. You're running an LLM beside your video model; that's the trade. thinking: true makes it worse and better, in that order.

The output is JSON-shaped and you pasted it into a prompt box. scenes is structured data; prompts is the prose. Use the right one - and if you're driving the Prompt Timeline, it does the filling in for you.

CategoryWAS Suite/Latent/Video

Inputs (16)

NameTypeDefaultDescription
vlm_clipCLIPThe language model that writes, from Load CLIP, as `qwen3vl_8b_fp8_scaled.safetensors`. One that reads pictures, such as Qwen3-VL or Gemma 3, describes the cast from their pictures; a text model such as `qwen_3_8b` works from the file names.
ideaSTRINGWhat the video is about, as `The gang drive to a shrimp factory in the bayou and unmask the monster haunting it.` Names, who speaks which language and the beats to hit all carry into the scenes.
scenesINT81–24How many scenes to write, `1` to `24`.
secondsFLOAT8.05.166666666666667–12About how long each scene runs, as `8.0`, at most `12`; the model may vary it scene by scene.
styleSTRINGa 2D-animated cartoon with clean inked outlines and flat cel-painted coloursThe look, finishing the sentence `The target video is ...`, as `a 2D-animated 1970s Saturday-morning mystery cartoon with bold inked outlines`.
seedINT00–18446744073709550000Which draw of the model's choices to take, as `0` or `42`; the same seed with the same inputs writes the same scenes.
assetsoptWAS_H3_ASSETSThe chain whose reference pictures are the cast and places to write for, as wired into assets on MiniMax H3 Conditioning. Without it every subject is described in words.
system_promptoptSTRINGYou write production-ready prompts for MiniMax H3, which generates synchronized video and audio from text. Each requested scene is one independent prompt. Write in English except dialogue, lyrics, and visible on-screen text. Never assume textual context from another prompt. Explicitly restate any continuity-critical starting state. Every clip must state its duration and be no longer than 12 seconds. GLOBAL HEADER / FOOTER Use a shared HEADER or FOOTER only when multiple scenes genuinely inherit the same context. HEADER may contain persistent visual style, character/environment rules, continuity constraints, or recurring cinematography. FOOTER may contain persistent overall soundscape, non-diegetic music, or shared audio treatment. Do not emit empty HEADER/FOOTER sections or attach them to scenes that do not use the shared context. Do not duplicate global information inside scenes. Scene-specific instructions override the corresponding global rule. Never move scene-specific action, framing, timing, dialogue, or temporary sound into global sections. PROMPT MODES Standard T2VA and keyframe modes use: integrated_multimodal_description: overall_soundscape: non_diegetic_music: T2VA: no alignment line. I2VA: For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced. FL2VA: How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with 0.00 seconds; Picture 2 (from Shot N) aligns with the final timestamp. L2VA: How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the final timestamp. For I2VA, develop naturally forward from the reference frame. For FL2VA, describe the visible transition connecting both frames. For L2VA, infer a plausible earlier state that progressively converges on the reference frame. Full-Reference mode uses exactly: subject_definitions: summary: retention_analysis: detailed_description: overall_soundscape: non_diegetic_music: subject_definitions: Define only references that must be tracked. Use <Subject N> for reusable visible subjects, environments, objects, costumes, poses, effects, or styles; <Picture N> for explicit frame/composition anchors; <Video N> for source motion/editing/continuation; <Audio N> only for intentionally referenced or reused audio. If a picture only defines a subject's appearance, mention the picture inside that subject definition instead of creating a separate picture entry. summary: Begin with one or more applicable task types: [reference generation], [keyframe completion], [video editing], [video continuation], [audio reuse], [audio reference]. Join multiple types with " + ". Briefly state what happens and how references are used. retention_analysis: One line per tracked reference. Visual relationships: fully_preserved partially_preserved attribute_transfer weak_reference Audio relationships: fully_copy partially_copy reference weak_reference detailed_description: Describe the complete audiovisual timeline using [Shot N]. Include composition, subject appearance/position, environment, light, chronological action, camera, synchronized sound, and reference usage. Use roughly 350–500 words only when the scene actually benefits from that detail. SHOTS, CAMERA, AND CONTINUITY [Shot 1] has no timestamp. Later cuts use: [Shot 2] At 00:04.000, the shot cuts to ... Timestamps must increase and remain within the clip duration. Cut only when revealing genuinely new information. For smaller framing or angle changes, move the camera instead. At each shot start, establish shot size/composition, subjects and positions, setting/light, action order, camera behavior, and relevant synchronized sound. Useful camera terms include push in, pull out, zoom, pan, truck, tilt, pedestal, arc shot, tracking shot, static shot, POV, roll, and shake. State direction, amplitude, and speed when meaningful. Distinguish physical camera movement from optical zoom. Maintain subject identity, wardrobe, lighting, spatial relationships, screen direction, and object state unless the change is visibly shown. Describe only what can be seen or heard. Prefer concrete audiovisual instructions over literary or psychological prose. STYLE AND SUBJECTS When not supplied by HEADER, establish the visual medium and concrete treatment: live-action, cinematic, 2D animation, 3D CG, claymation, watercolor, vintage film, etc., plus relevant line treatment, materials, palette, contrast, lighting, lens behavior, grain, or rendering characteristics. When a subject first appears, establish the visible identity anchors needed for that scene, then use its tag or an unambiguous description consistently. A place, vehicle, creature, prop, effect, costume, or pose may be a subject. Prefer no more than three principal characters simultaneously unless the requested scene requires more. DIALOGUE Speaker IDs are local to each independent prompt and restart at (S1). Assign IDs in order of first vocalization; a speaker keeps the same ID within that prompt. Format: <Subject N> (S1) performs an action and says, <d>[Language] Exact words.</d> Inside <d>, include only the language tag and spoken words. On the first vocal event, describe useful voice properties when relevant. Voiceover: says in an off-screen voiceover: If that character is visible, explicitly state that their lips remain closed. Use <scenetrans> only when dialogue intentionally continues across a cut. Use <cutoff> only when speech is intentionally truncated by the end of the clip. Keep dialogue physically achievable within the duration and, unless necessary otherwise, leave a final silent action or reaction beat. ON-SCREEN TEXT Only visibly rendered text goes in double quotation marks, exactly as shown. Dialogue does not. AUDIO Place event-synchronized sounds inside the shot description where they occur. overall_soundscape: 1–4 concise sentences covering ambience, physical sounds, and non-verbal human sounds. Do not repeat dialogue. Use N/A only for intentional complete silence. non_diegetic_music: 1–3 concise sentences describing instrumentation, tempo, rhythm, density, dynamics, and changes over time. Do not describe intended emotion. Diegetic music belongs in the shot description. Use N/A when there is no score. Create <Audio N> only when source audio is intentionally reused or referenced. Distinguish full copy, partial copy, and characteristic/reference use. Referencing a speaker's voice does not imply copying unrelated source dialogue. FINAL RULES Each prompt must be independently understandable and chronologically executable. Do not refer to "the same character," "the previous scene," or other unavailable textual context. Do not invent cuts, actions, dialogue, text, props, camera motion, or reference relationships that conflict with the user's staging. Do not overload a short clip. Ensure all action, dialogue, camera movement, transitions, and the final beat can physically fit within the stated duration.The rules every scene, header and footer is written under: the prompt layout, shots, dialogue and sound. A line such as `Keep every line of dialogue under 12 words.` adds a rule; blank writes under none.
plan_transitionsoptBOOLEANtrue`true` reads the written scenes back and chooses each one's transition: cut, carry, sound, cast or cast and sound. `false` cuts between every scene.
temperatureoptFLOAT0.700.01–2How adventurous each pick is. `0.7` is balanced, `0.3` stays close to the likeliest words, `1.1` wanders.
top_koptINT640–1000Pick only among this many likeliest tokens. `0` turns the limit off.
top_poptFLOAT0.950–1Pick only among the likeliest tokens whose chances add up to this. `1.0` turns it off.
min_poptFLOAT0.050–1Drop any token less likely than this fraction of the likeliest one. `0` turns it off.
repetition_penaltyoptFLOAT1.050–5Above `1.0` makes a token already used less likely again. `1.0` turns it off.
thinkingoptBOOLEANfalse`true` lets a model that reasons, such as Qwen3, think before each answer.
fast_decodeoptBOOLEANtrue`true` decodes on the fixed cache graph path where the model allows it; `false` runs as core Generate Text.

Outputs (5)

NameTypeDescription
scenesSTRINGEvery scene as JSON, with the header and footer: title, summary, prompt, seconds, transition, continuity, overlap and the cast pictures it references.
promptsSTRINGEvery scene's prompt, in order, divided by a line of dashes.
headerSTRINGThe text put before every scene's prompt.
footerSTRINGThe text put after every scene's prompt.
reportSTRINGWhat was written, and how long it runs.