H3 Hard Constraint Producer
Dialogue the model isn't allowed to reword
- hard_constraints
- producer_report
Everything else in your prompt is a suggestion. This node is where you write down the things that are not.
If your clip has a spoken line, a sign in the background, or a brand name that must appear exactly as typed, an LLM-encoded video model will happily paraphrase it. H3 reads your prompt as an instruction with structure, not as a bag of tokens (prompt-engineering.md), which means phrase-level fidelity is something you have to assert rather than hope for.
What it does
It assembles exact, caller-authored constraints - and then refuses to touch them. The module's contract is "exact caller-authored hard constraints without inference or rewriting." Nothing is normalised, spell-checked, summarised or tightened. What you type is what the constraint says, and downstream stages are required to preserve it character for character.
The node also emits its own report alongside the constraint set, which matters more than it sounds: it's a typed claim that this constraint bundle exists and is unmodified, and the nodes downstream - the directive authority producer, the acceptance gate - require that report to be paired with this exact set.
Inputs and outputs
Nine required widgets and three optional ones. The required set:
- dialogue - the spoken line, exactly.
- visible_text - on-screen text, signs, titles, captions. Anything that has to be legible in frame.
- required_content - things that must appear.
- forbidden_content - things that must not. This is where "no watermark," "no extra characters," "no logos" goes, and it's the constraint people most often forget to write down and then complain about.
- timing_start and timing_end - when a constrained thing happens. Leave blank for "anywhere in the clip."
- keep_target, keep_value and change_replacement - the keep/change pair: what entity you're pinning, what value to preserve, and what to substitute.
keep_targetis a closed list:subject,scene,action,camera,style,audio,dialogue,lyrics,visible_text,asset.
The optional three are all about speech delivery:
- dialogue_language -
autoby default, or one of the eleven supported languages (Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish).autorecognises Korean, Japanese and Chinese from their scripts; other scripts stay untagged, with a warning. Crucially, the spoken words never change - this is a hint about pronunciation and script, not a translation. - dialogue_speaker - a short identity phrase, like
the guide. Not a person; a role. - dialogue_delivery -
on_screenorvoiceover. Voiceover means the speaker's lips stay closed, which is a real difference in generated footage and the sort of instruction a model only obeys if you said it.
Outputs: hard_constraints (H3_HARD_CONSTRAINTS) and producer_report. The constraint set feeds the request path, the directive authority producer, and the acceptance gate.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/rookiestar28/ComfyUI-MiniMaxH3-Studio.git
# restart ComfyUI
Not in the Comfy Registry yet, so the clone is the route. No Python dependencies (dependencies = []), no downloads at install, prebuilt extension. Python 3.10+. The node needs no model weights - you can build and inspect constraint sets on a machine that has never run H3, and examples/minimal_core_pipeline.py exercises this path in pure Python.
Where people get burned
Timing hints are strings, not numbers, and the pipeline expects exact timing points. If you feed it something creative, expect a validation error rather than a reinterpretation - and expect the diagnostic to name the contact point. The same author's earlier pack, ComfyUI-OpenClaw, is built on the same instinct: strict boundaries and fail-closed defaults rather than convenient guesses.
The second trap is empty dialogue plus an intended-silent video. If the clip is meant to be silent, an empty constraint set reads as "unspecified," which is different from "silent," and you'll see guide readiness come back Incomplete even though validation passed. The README's remedy is specific: turn on complete_silence in H3 Intent Graph Producer. Silence you asserted is a fact; silence you omitted is a gap.
Third: keep the words short. A constraint is not a script. If you need fifteen seconds of dialogue, you need more than one clip's worth of duration, and H3's clips cap at 15 seconds (MiniMax H3 panel). Write what fits, then plan a longer video as segments.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| dialogue | STRING | — | |
| visible_text | STRING | — | |
| required_content | STRING | — | |
| forbidden_content | STRING | — | |
| timing_start | STRING | — | |
| timing_end | STRING | — | |
| keep_target | COMBO | subject | 10 options: subject, scene, action, camera, style, audio, +4 |
| keep_value | STRING | — | |
| change_replacement | STRING | — | |
| dialogue_languageopt | COMBO | auto | 12 options: auto, Arabic, Chinese, English, French, German, +6 |
| dialogue_speakeropt | STRING | — | |
| dialogue_deliveryopt | COMBO | on_screen | 2 options: on_screen, voiceover |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| hard_constraints | H3_HARD_CONSTRAINTS | — |
| producer_report | H3_DOWNSTREAM_PRODUCER_REPORT | — |