Nodes/MiniMax H3 Planner/H3 Story Planner
ComfyUI Node

H3 Story Planner

One brief in, a whole planned and written timeline out

By AIJigyasa·Created 23 days ago·Updated 20 days ago· 4
H3 Story Planner
  • project
  • cast
  • timeline
  • beat_sheet
  • report
◄idea►
◄total_seconds30.0►
◄min_segment_seconds5.0►
◄max_segment_seconds10.0►
◄formatcinematic ad►
◄dialogueauto►
◄audio_roleperformed on camera►
◄window4►
◄single_call_max8►
◄chain_first_framesfalse►
◄providerOllama (Local)►
◄ollama_urlhttp://127.0.0.1:11434►
◄ollama_modelqwen3-vl:8b►
◄temperature0.35►
◄reuse_existingtrue►
◄seed0►
◄style_prefix_override►
◄api_key►
◄api_model►
◄max_output_tokens8192►
◄num_ctx16384►
◄keep_alive10m►
◄request_timeout900►

For an advert, a short film, a UGC piece or a product video, the per-segment approach has a structural flaw: writing each segment in its own model call, told to make each one differ from the last, means nothing carries wardrobe, location, time of day, or where anyone was standing. Every shot reinvents the world and the stitched result reads as a slideshow.

H3 Story Planner inverts that. Continuity stops being something plumbed between calls and becomes a property of one generation: a single brief goes in, a whole planned and written timeline comes out.

Two kinds of call, then code does the rest

First call: the bible and the beat sheet. Look, style prefix, the frozen world block, the canonical cast wording, the scenes, and for every segment its duration, its action, and the physical state it hands to the next.

Then the prose, expanded in windows (window, default 4 segments) each seeing the previous window's closing state, so a long piece never needs one enormous reply. At or below single_call_max segments (8 by default) the whole video is written in one prose call, which is what makes it consistent. A 30-second ad is one call instead of the twenty-one the old path needed.

Then assembly is deterministic - durations snap to the frame ladder, the style prefix and cast wording are copied rather than requested, and every enforcement in the pack gets reapplied:

  • every segment's summary states the same location, time of day and light in identical words
  • a shot that ignores where the previous clip ended is given that state
  • a segment that runs past its own ending into the next has the extra shots removed and reported, so an action never renders twice
  • references are bracketed and bound in the section H3 actually reads, and a tag the cast doesn't have is stripped
  • dialogue is budgeted for the runtime, written inside <d>[English] …</d> with a stable speaker ID, and a segment with no line is told nobody speaks

Inputs worth setting

project and idea are the wires. idea is the whole brief - what happens, who's in it, what it's for. Then total_seconds, and the clip length window: min_segment_seconds (longer clips cost less per second of video, because per-run overhead is amortised) and max_segment_seconds (a ceiling, not a target).

  • format - auto, short film, cinematic ad, ugc, product, explainer, music video.
  • dialogue - auto decides from the brief; spoken lines makes the planner write the actual script and render it as H3 dialogue with speaker IDs; none keeps it picture only.
  • audio_role - what the connected audio is. The first three treat it as the soundtrack, reused exactly: performed on camera when a visible subject raps or sings it. Voice sample (clone the timbre) is the opposite and it's the one for an ad - the audio is only a timbre, the dialogue is written here and spoken in that voice. For two characters, connect one voice sample each and say who is who in the brief: "use <audio 1> for the man and <audio 2> for the woman". The report shows the casting and warns when the role contradicts the brief.
  • chain_first_frames - hands each clip's last frame to the next within a scene. Hard visual continuity at the joins, but artifacts compound across a long chain, so it's off by default and the tooltip says leave it off unless you need it.
  • reuse_existing / seed - re-queueing with nothing changed costs no model calls. Change seed to replan everything.

Optional: cast, style_prefix_override, api_key (leave it empty - a key typed here is saved into the workflow), api_model, max_output_tokens, num_ctx, keep_alive, request_timeout. Outputs: timeline, beat_sheet, report.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/AIJigyasa/ComfyUI-H3-Planner
git clone https://github.com/AIJigyasa/ComfyUI-H3-Prompt-Creator

Restart. The Story Planner needs a language model - local Ollama (default qwen3-vl:8b) or OpenAI / Anthropic / OpenRouter / Gemini through the Prompt Creator's provider stack - and it needs that sibling pack installed, since it imports its engine rather than copying it. ffmpeg on PATH is needed for the pack as a whole.

Where people get burned

Drop your references onto the Cast Board first. The planner sends reference images to the vision model so it writes what the subject actually looks like. Give it one person's face and a written description of someone else and H3 resolves the contradiction by blending them.

Don't leave this node in your render graph. With reuse_existing off it rewrites the prompts on every queue, which resets the segments that changed, which discards their clips. H3 Vault Write's auto-advance actually refuses to loop when it detects a segment stored twice in a row, and names this exact cause. Plan in one workflow, render in another.

A local 8B is not a director. The community ceiling with small local models on structured writing is real: they remove the blank page. If your beat sheet reads thin, raise temperature slightly or run the planning pass through an API provider.

CategoryH3 Planner

Inputs (25)

NameTypeDefaultDescription
projectH3_PROJECT—
ideaSTRINGthe whole brief: what happens, who is in it, what it is for
total_secondsFLOAT30.02–600—
min_segment_secondsFLOAT5.00.5–15shortest clip the planner may ask for. Longer clips cost less per second of video, because the fixed per-run overhead is amortised.
max_segment_secondsFLOAT10.01–15longest clip your card can render. A ceiling, not a target — the planner varies length to suit each beat.
formatCOMBOcinematic ad7 options: auto, short film, cinematic ad, ugc, product, explainer, +1
dialogueCOMBOautowhether anyone speaks on camera. 'spoken lines' makes the planner write the actual script and render it as H3 <d>[English] ...</d> dialogue with speaker IDs; 'auto' decides from the brief; 'none' keeps it picture only.
audio_roleCOMBOperformed on camerawhat the connected audio IS. The first three treat it as the soundtrack, reused exactly: 'performed on camera' when a visible subject raps or sings it. 'voice sample' is the opposite and is the one for an ad: the audio is only a timbre, the dialogue is written here and generated in that voice. Ignored with no audio in the cast.
windowINT41–12segments written per prose call once the piece is too long for one reply. Each window sees the previous window's closing state.
single_call_maxINT81–40at or below this many segments the whole video is written in ONE prose call, which is what makes it consistent. Above it, windows are used.
chain_first_framesBOOLEANfalsestart each segment from the previous one's last frame, within a scene. Hard visual continuity at the joins, but artifacts compound across a long chain — leave off unless you need it.
providerCOMBOOllama (Local)1 options: Ollama (Local)
ollama_urlSTRINGhttp://127.0.0.1:11434—
ollama_modelSTRINGqwen3-vl:8b—
temperatureFLOAT0.350–1.2—
reuse_existingBOOLEANtruekeep segments already written from the same brief
seedINT00–4294967295change to replan the whole video
castoptH3_CAST—
style_prefix_overrideoptSTRING—
api_keyoptSTRING—
api_modeloptSTRING—
max_output_tokensoptINT8192256–32768—
num_ctxoptINT163842048–131072—
keep_aliveoptSTRING10m—
request_timeoutoptINT90030–3600—

Outputs (3)

NameTypeDescription
timelineH3_TIMELINE—
beat_sheetSTRING—
reportSTRING—