Nodes/MiniMax H3 Planner/H3 Segment Prompter
ComfyUI Node

H3 Segment Prompter

A slice of a long prompt is not a prompt — this rewrites each one properly

By AIJigyasa·Created 24 days ago·Updated 21 days ago· 4
H3 Segment Prompter
  • project
  • cast
  • timeline
  • report
◄h3_prompt►
◄total_seconds60.0►
◄segment_seconds10.0►
◄shots_per_segment1►
◄formatauto►
◄audio_roleperformed on camera►
◄providerOllama (Local)►
◄ollama_urlhttp://127.0.0.1:11434►
◄ollama_modelqwen3-vl:8b►
◄temperature0.25►
◄reuse_existingtrue►
◄seed0►
◄idea►
◄style_prefix_override►
◄api_key►
◄api_model►
◄max_output_tokens3072►
◄num_ctx8192►
◄keep_alive10m►
◄request_timeout600►

Here's the trap in multi-segment H3 work. You write one beautiful six-section prompt for the whole video. You cut it into twelve pieces. Every piece is now broken: it has no style prefix, no subject_definitions of its own, no soundscape, and its timestamps start wherever the cut happened to fall. H3 renders each segment standalone, so each segment needs a complete prompt - not a fragment of one.

H3 Segment Prompter writes those. One long prompt in, per-segment prompts out.

What it enforces in code rather than by asking

This is the part that makes it worth the model call. Everything a language model is unreliable about is checked and fixed after the fact, in Python:

  • the style prefix appears verbatim at the head of [Shot 1], so segment six can't drift to a different look
  • [Shot 1] carries no timestamp, and later shots are respaced strictly inside the segment's rendered duration
  • all essential action completes by the planned duration - the ladder overshoot becomes an explicit hold, not a new action
  • subject_definitions covers only the tags that segment actually cites, de-duplicated
  • (S1)/(S2) speaker IDs stay stable across segments
  • tags above the cast size are flagged

That last set is why this node and the no-model H3 Segment Slicer exist side by side: the slicer can't drift because it never rewrites anything, and this one rewrites so it needs the guardrails. The slicer's own notes are blunt about the failure mode - a model asked to preserve a style prefix will paraphrase, and a paraphrase can contradict the very prefix it was told to keep.

Inputs that matter

Wire project. Then h3_prompt takes the h3_prompt output of your H3 prompt creator, run at the FULL video duration - not one clip. total_seconds is the same number you gave the prompt creator, and getting these two out of step is the classic way to produce nonsense.

  • shots_per_segment - 1 gives one segment per shot in the treatment and the most detail per clip. Raise it to pack several shots into one longer clip while they still fit under the cap.
  • segment_seconds - a rough length per segment, only used when there's no treatment to cut on.
  • format - auto, short film, cinematic ad, ugc, product, explainer, music video. It steers emphasis; auto adds nothing.
  • audio_role - who makes the sound in the connected audio. Set performed on camera when the subject raps or sings it and their mouth has to match the words; without that, H3 plays the track over someone sitting silently. The role is ignored when the cast has no audio asset.
  • reuse_existing and seed - segments already written from the same inputs are skipped, and locked segments are never touched. Re-running after one edit costs one call, not twelve. Change seed to rewrite everything.

Optional: cast, idea (extra direction on every segment, or the whole plan when h3_prompt is empty), style_prefix_override, and the provider block - provider, ollama_url, ollama_model, temperature, num_ctx, max_output_tokens, request_timeout.

Outputs are timeline and report.

It needs a model, and it needs the sibling pack

This node imports ComfyUI-H3-Prompt-Creator rather than copying it, reusing that pack's system guide, providers, JSON repair and shot-timing enforcement. Without it installed you get a plain message saying so, and every other node in the pack still works.

cd ComfyUI/custom_nodes
git clone https://github.com/AIJigyasa/ComfyUI-H3-Planner
git clone https://github.com/AIJigyasa/ComfyUI-H3-Prompt-Creator

Restart both times. Default provider is local Ollama (qwen3-vl:8b at http://127.0.0.1:11434); OpenAI, Anthropic, OpenRouter and Gemini come through the Prompt Creator's provider stack.

Where people get burned

Leave api_key empty. ComfyUI saves every widget value inside the workflow file, so a key typed into that field travels with anything you share or export. Set the provider's environment variable instead.

One call per segment, and a local 8B is not a writer. The community's honest ceiling with small local models on structured prompt work is "removes the blank page", not "writes like a director". If a segment comes back limp, raise temperature a little, or use an API provider for the planning pass and keep everything else local.

A planner in the render graph will destroy your renders. With reuse_existing off, every queue rewrites the prompts, the timeline discards clips that no longer match their spec, and segments go outstanding again right after they finish. Plan in one workflow, render in another.

CategoryH3 Planner

Inputs (22)

NameTypeDefaultDescription
projectH3_PROJECT—
h3_promptSTRINGthe h3_prompt output of your H3 prompt creator, run at the FULL video duration. Leave empty to plan from `idea` instead.
total_secondsFLOAT60.01–600the whole video's length — the same number you gave the prompt creator
segment_secondsFLOAT10.01–15rough length per segment; only used when there is no treatment to cut on
shots_per_segmentINT11–121 = one segment per shot in the treatment, and one prompt written for it — the most detail per clip. Raise it to pack several shots into one longer clip while they still fit under the segment cap. With only an idea (no treatment) it is the most shots each clip may have, and each shot is at least 1.5s: 1 = one continuous shot per clip.
formatCOMBOautowhat kind of video this is. Steers emphasis per segment; auto adds nothing.
audio_roleCOMBOperformed on camerawho makes the sound in the connected audio. Set 'performed on camera' when the subject raps, sings or speaks it and their mouth must match the words - without that H3 plays the track over someone sitting silently. Ignored when no audio asset is in the cast.
providerCOMBOOllama (Local)1 options: Ollama (Local)
ollama_urlSTRINGhttp://127.0.0.1:11434—
ollama_modelSTRINGqwen3-vl:8b—
temperatureFLOAT0.250–1.2—
reuse_existingBOOLEANtrueskip segments already written from the same inputs
seedINT00–4294967295change to rewrite every segment
castoptH3_CAST—
ideaoptSTRINGextra direction applied to every segment; the whole plan when h3_prompt is empty
style_prefix_overrideoptSTRINGblank = lift it off [Shot 1] of the treatment
api_keyoptSTRINGprefer the environment variable — a key typed here is saved into the workflow
api_modeloptSTRING—
max_output_tokensoptINT3072256–32768—
num_ctxoptINT81922048–131072—
keep_aliveoptSTRING10m—
request_timeoutoptINT60030–3600—

Outputs (2)

NameTypeDescription
timelineH3_TIMELINE—
reportSTRING—