H3 Intent Graph Producer
Write down what the shot is about before you say it prettily
- request
- reference_registry
- intent_graph
- producer_report
Most prompt problems are actually intent problems wearing a costume. You wrote a paragraph; the model read it as instructions; you meant something slightly different by three of the phrases; and the video is off in a way you can't name. This node forces the naming.
What it does
It builds a deterministic intent graph from your request and reference registry: the subjects, the action, the anchors the subject owns, and how the segment develops. It's the join between "what I asked for" and "what the prompt will say."
The distinguishing feature is what it won't do. Its contract is explicit that it builds a manual graph - no free-text inference, no model reading your prose and guessing at structure. If you describe a courier and a bridge and never say which one is the subject, the graph doesn't decide for you. It leaves the detail unspecified, visibly.
That matters because it's also the node that resolves the silence question. H3 generates audio jointly with the picture, so an empty soundscape in a plan means the model is told nothing about sound - which is not the same as being told the video is silent. The pack refuses to render overall_soundscape: N/A unless silence was explicitly authored. This node is where you author it.
Inputs and outputs
Required:
- request - the typed request from H3 Context Request.
- reference_registry - the canonical registry from H3 Reference Registry.
- subject_label - the short identity of your primary subject, e.g.
the courier. - action_description - what they're doing.
Optional, and each one is a decision you're choosing to make explicit:
- secondary_subject_label and secondary_action_description - a second subject with its own action. Two subjects in one clip is the case that most often falls apart, so naming both rather than implying one is worth the ten seconds.
- complete_silence - boolean, default off. Turn it on only when silence is genuinely the intent. This is the switch the README points at when validation passes but guide readiness reports Incomplete.
- keyframe_binding -
unbound(default),first_frame,last_frame, orboth. Which source role your primary subject owns. Anchoring the subject to a frame is a different request from merely showing them. - segment_development -
unspecified(default),anchor_development, oranchor_static_hold. In plain terms: does the shot evolve away from its anchor frame, or hold it? These are caller declarations about content, never observations inferred from your media.
Outputs: intent_graph (H3_INTENT_GRAPH) and producer_report (H3_DOWNSTREAM_PRODUCER_REPORT). The graph and report pair feeds H3 Directive Authority Producer, H3 Full Reference Timeline Producer and H3 Local Reconstruction Acceptance - all of which want both halves.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/rookiestar28/ComfyUI-MiniMaxH3-Studio.git
# restart ComfyUI
Not in the registry yet, so it's a manual clone. No Python dependencies declared, no weights downloaded at install, prebuilt browser extension, Python 3.10+. If you want to see a working graph, workflows/m15_02_downstream_minimal.json wires this node into the producer chain with the exact inputs above.
Why you'd reach for it on the canvas
The sidebar builds an intent graph for you under the hood - the Context page's numbered stages are the same pipeline, and its promise is that the canvas and the sidebar never disagree about your prompt. So if you're happy clicking through the UI, you may never touch this node.
Reach for it when you want that structure visible and reusable. An intent graph is the part of a video workflow worth templating: same subject, same relationship between subject and anchor, different action. It's also the honest place to admit what you haven't decided. unspecified isn't an error state - it's a hole you can see, which is much better than a hole that gets filled in by a language model and never mentioned (prompt-engineering.md).
And if you're shotgunning generations: check this node before you re-roll the seed. Nine times out of ten, the reason the model keeps drifting from your idea is that the intent you're asking for is ambiguous, and no sampler setting fixes an ambiguous brief.
Gotchas
complete_silence is the one people get wrong in both directions. Left off with no audio intent, you get an incomplete guide-readiness report - not a failure, a nudge. Turned on by habit, you've commanded a silent video, and the model will honour you.
One more: this node doesn't need the reference registry to be interesting, but it does require one on the socket. A registry with no assets in it is fine for a text-to-video job; what isn't fine is expecting the graph to invent subject identities from your prose, because it won't.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| request | H3_CONTEXT_REQUEST | — | |
| reference_registry | H3_REFERENCE_REGISTRY | — | |
| subject_label | STRING | — | |
| action_description | STRING | — | |
| secondary_subject_labelopt | STRING | — | |
| secondary_action_descriptionopt | STRING | — | |
| complete_silenceopt | BOOLEAN | false | — |
| keyframe_bindingopt | COMBO | unbound | 4 options: unbound, first_frame, last_frame, both |
| segment_developmentopt | COMBO | unspecified | 3 options: unspecified, anchor_development, anchor_static_hold |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| intent_graph | H3_INTENT_GRAPH | — |
| producer_report | H3_DOWNSTREAM_PRODUCER_REPORT | — |