Nodes/TrentNodes/Ultimate H3 Cowboy Promptor
ComfyUI Node

Ultimate H3 Cowboy Promptor

H3 prompts for any shot, not just the one trick

By TrentHunter82·Created 9 months ago·Updated 4 days ago· 36
Ultimate H3 Cowboy Promptor
  • video
  • frames
  • audio
  • subject_1_image
  • subject_2_image
  • subject_3_image
  • subject_4_image
  • subject_5_image
  • subject_6_image
  • h3_prompt
  • duration_seconds
  • fps
  • analysis_json
  • h3_checkpoint_hint
  • ref_image_1
  • ref_image_2
  • ref_image_3
  • ref_image_4
  • ref_image_5
  • ref_image_6
  • ref_video
  • ref_video_audio
  • ref_audio
  • width
  • height
  • length
  • label_map
h3_moderef
subjects
target_descriptionthe courier ducks under a roller shutter
vlm_provideranthropic
modelauto
fps24.00
api_key
video_rolesubject_source
audio_rolenone
cut_times
dialogue
constraint_notes
duration_override0.00
max_frames_to_analyze8
seed0
base_picture_rolefirst_frame
fl2va_normalize_picture_tagsfalse
snap_duration_to_h3_gridtrue
subject_rows2
subject_1_kindcharacter
subject_1_name
subject_1_description
subject_2_kindenvironment
subject_2_name
subject_2_description
subject_3_kindcharacter
subject_3_name
subject_3_description
subject_4_kindcharacter
subject_4_name
subject_4_description
subject_5_kindcharacter
subject_5_name
subject_5_description
subject_6_kindcharacter
subject_6_name
subject_6_description
music_videofalse
music_sourceauto
lyrics
music_description

The H3 Auto Prompt Generator does one job brilliantly: character replacement from a reference image. The Ultimate H3 Cowboy Promptor is the node for everything else H3 can do. Where the other node is a focused tool, this one writes a MiniMax H3 prompt for any kind of shot - people, animals, objects, scenes, styles, actions, several subjects at once - and then hands the assets straight to the sampler. The README is clear that both are installed on purpose: one job done well, and one job done for the whole spec.

The promise is validated, not vibed: all six of MiniMax's own worked examples pass this node with zero errors, per the README.

The big idea: subjects are rows, not syntax

Instead of typing a monster prompt, you fill in rows. Each row is a kind (character, environment, animal, object, wardrobe, interface, effect, style, action, expression, pose), an optional name, and a description. character and environment sit at the top of the kind list and become the guide's own person and scene; the rest are the guide's other reusable kinds.

One rule runs the whole way through - row N is <Subject N> is subject_N_image is <Picture N>. So a skipped row leaves a hole rather than closing it up, and the node says so. A row needs no image (that's how you ask for a style or an action), and the typed subjects field behind the advanced panel is the escape hatch for a seventh subject or one that cites two pictures at once.

Two formats, one node

h3_mode picks the skeleton:

  • ref - the six-section Ref2VA format from reference assets. This is the multi-subject mode.
  • base_T2VA / base_I2VA / base_FL2VA / base_L2VA - the three-field base format: text alone, first frame, first+last frames, last frame. These need a different H3 checkpoint, which the h3_checkpoint_hint output names for you.

vlm_provider (same ten-option list as the Auto node) chooses the brain, target_description says what the target video should show, and cut_times wires Cut Detective in to pin the shot timeline.

The part that keeps everything honest: nothing is wired twice

The killer feature is the outputs. The images, video, and audio you plug in come back out as ref_image_1..6, ref_video (as IMAGE frames - what the sampler actually takes), ref_video_audio, and ref_audio, plus the width, height, and length that MiniMaxH3ReferenceToVideo wants and a label_map you can read against the prompt.

The point: the prompt's <Picture N> tags cannot drift from what the sampler receives, because the node hands you the exact assets it referenced. No "prompt says Picture 2 but I wired the wrong image" debugging session. length sits on H3's 17k+5 frame grid, and snap_duration_to_h3_grid (on by default) makes the prompt state the duration that grid really produces - ask for 2.00 seconds and H3 renders 2.33, so the prompt says 2.33.

Gotchas

  • Backends are the same seven hosted/local options as the Auto node - this is a VLM-driven node, so it needs a provider (or a local model) to write the prose.
  • Rows and images must stay in order. Slot 1 is <Picture 1>; if you wire subject images out of order, the prompt follows the slots, not your intent.
  • snap_duration_to_h3_grid off means the prompt claims a shorter video than H3 produces. On by default, and there's a reason.

Install

One of ~69 nodes in TrentNodes:

# ComfyUI Manager: search "Trent Nodes"

# or:
cd ComfyUI/custom_nodes
git clone https://github.com/TrentHunter82/TrentNodes.git
cd TrentNodes && pip install -r requirements.txt

Pack install covers it; provider clients (anthropic/openai/google-genai) and keys are on you if you use hosted VLMs. If H3 character replacement is your whole world, the Auto node is leaner. The moment you want H3 to do more than swap a performer, this is the one you reach for.

CategoryTrent/VLM

Inputs (50)

NameTypeDefaultDescription
h3_modeCOMBOrefref writes the six-section Ref2VA format from reference assets. The base_* modes write the three-field base format and need a different H3 checkpoint (see h3_checkpoint_hint): T2VA is text only, I2VA anchors the first frame, FL2VA the first and last, L2VA the last. In base mode the subjects field is ignored - there are no reference labels at all.
subjectsSTRINGEXTRA subjects, typed. The rows below are the normal way in; this is the escape hatch for what a row cannot say - a seventh subject, or one that cites two pictures at once. One line each: <kind> [name] [@Picture N ...] -- <features> Kinds: character, environment, animal, object, wardrobe, interface, effect, style, action, expression, pose. Typed lines are numbered AFTER the filled rows. Ignored in the base_* modes, which have no reference labels at all.
target_descriptionSTRINGthe courier ducks under a roller shutterWhat the target video should show.
vlm_providerCOMBOanthropicWhich VLM writes the prompt.
modelSTRINGauto'auto' uses the provider's default model.
videooptVIDEOSource clip. Preferred input; carries its fps.
framesoptIMAGEAlternative to video: an IMAGE batch. Set fps to match.
fpsoptFLOAT24.001–240Frame rate of the frames input.
audiooptAUDIOThe clip's track, for a provider that can hear it (gemini). Becomes <Audio 1>.
api_keyoptSTRINGBlank uses the provider's environment variable.
video_roleoptCOMBOsubject_sourceWhat the wired video is FOR - this is what decides the task type, and wiring alone cannot imply it. subject_source: it just shows a subject, so it is cited inside that subject's line and gets no entry. structure_reference: its camera, cuts and rhythm are followed. edit_source: the target IS this video, edited. continuation_source: the target continues from where it ends.
audio_roleoptCOMBOnoneWhat the wired audio is FOR. reuse: the signal is copied into the target (adds 'audio reuse'). reference: only its timbre, style or beat is followed, not the signal (adds 'audio reference').
cut_timesoptSTRINGWire any Cut Detective output here. The shot list becomes ground truth for the [Shot N] times. A count mismatch is only an error when a video is being edited - otherwise the target's structure is not bound to the reference's. In base mode there is no source clip, so this reads as the shot structure you are asking for: more than one entry is what lets FL2VA write more than one shot.
dialogueoptSTRINGExact spoken words. They must reach a <d> block verbatim, or the run retries once.
constraint_notesoptSTRINGThings that must not change. Folded into the prompt as positive assertions, because H3 has no negative field and the official format ends at non_diegetic_music - nothing may follow it.
duration_overrideoptFLOAT0.000–6000 = use the clip's own duration. In a base mode with no clip this IS the target length: it reaches the instruction line as S.SS, and leaving it at 0 makes the node guess 5 seconds and warn.
max_frames_to_analyzeoptINT82–16Keyframes sampled from the clip for the VLM.
seedoptINT00–2147483647Passed to providers that support seeding.
subject_1_imageoptIMAGEReference image for line 1 of the subjects field. This slot IS <Picture 1> - the numbers always match, so wire them in order from slot 1. In a base_* mode there are no subjects: slot 1 is the anchor frame, and slot 2 is FL2VA's last frame.
subject_2_imageoptIMAGEReference image for line 2 of the subjects field. This slot IS <Picture 2> - the numbers always match, so wire them in order from slot 1. In a base_* mode there are no subjects: slot 1 is the anchor frame, and slot 2 is FL2VA's last frame.
subject_3_imageoptIMAGEReference image for line 3 of the subjects field. This slot IS <Picture 3> - the numbers always match, so wire them in order from slot 1. In a base_* mode there are no subjects: slot 1 is the anchor frame, and slot 2 is FL2VA's last frame.
subject_4_imageoptIMAGEReference image for line 4 of the subjects field. This slot IS <Picture 4> - the numbers always match, so wire them in order from slot 1. In a base_* mode there are no subjects: slot 1 is the anchor frame, and slot 2 is FL2VA's last frame.
subject_5_imageoptIMAGEReference image for line 5 of the subjects field. This slot IS <Picture 5> - the numbers always match, so wire them in order from slot 1. In a base_* mode there are no subjects: slot 1 is the anchor frame, and slot 2 is FL2VA's last frame.
subject_6_imageoptIMAGEReference image for line 6 of the subjects field. This slot IS <Picture 6> - the numbers always match, so wire them in order from slot 1. In a base_* mode there are no subjects: slot 1 is the anchor frame, and slot 2 is FL2VA's last frame.
base_picture_roleoptCOMBOfirst_frameBase mode only: whether a single wired picture is the FIRST frame or the LAST one. Nothing in the pixels says which, so it has to be declared. h3_mode already declares it too (base_I2VA is first, base_L2VA is last); this is the cross-check, and a disagreement warns and follows h3_mode.
fl2va_normalize_picture_tagsoptBOOLEANfalsebase_FL2VA only. MiniMax's guide writes bare 'Picture 1' and 'Shot 1' for FL2VA, with no brackets, while I2VA and L2VA bracket both - in the instruction line AND in the body of its own worked example. Off reproduces that. On rewrites them to <Picture 1> and [Shot 1]. No validator can tell which generates a better video, so this exists to be A/B'd; the setting is recorded in analysis_json.
snap_duration_to_h3_gridoptBOOLEANtrueH3 renders whole frames on a 17k+5 grid at 24 fps, so it rounds a length UP: ask for 2.00 seconds and you get 2.33. On, the prompt states the length H3 really produces, and the length output matches it. Off keeps the requested number, so the prompt claims a shorter video than the one H3 makes. Both numbers are recorded in analysis_json.
subject_rowsoptINT20–6How many subject rows to show. It only controls the node face: a row with anything typed in it is always used, and the count grows on its own when you wire an image or fill the last row. Up to 6.
subject_1_kindoptCOMBOcharacterWhat subject_1_image / row 1 IS. character and environment are the two everyone needs; the rest are the guide's other reusable kinds. The kind picks the features the prompt asks for - a character gets face, hair and garments, an environment gets surfaces and light, a style gets grain and grade.
subject_1_nameoptSTRINGOptional name for row 1, e.g. 'Aria Voss' or 'the loading bay'. Blank is fine - MiniMax's own example names nobody.
subject_1_descriptionoptSTRINGWhat row 1 looks like: the features the target video must keep. Type here and this row becomes <Subject 1>, and subject_1_image becomes <Picture 1>. An empty row is not a subject.
subject_2_kindoptCOMBOenvironmentWhat subject_2_image / row 2 IS. character and environment are the two everyone needs; the rest are the guide's other reusable kinds. The kind picks the features the prompt asks for - a character gets face, hair and garments, an environment gets surfaces and light, a style gets grain and grade.
subject_2_nameoptSTRINGOptional name for row 2, e.g. 'Aria Voss' or 'the loading bay'. Blank is fine - MiniMax's own example names nobody.
subject_2_descriptionoptSTRINGWhat row 2 looks like: the features the target video must keep. Type here and this row becomes <Subject 2>, and subject_2_image becomes <Picture 2>. An empty row is not a subject.
subject_3_kindoptCOMBOcharacterWhat subject_3_image / row 3 IS. character and environment are the two everyone needs; the rest are the guide's other reusable kinds. The kind picks the features the prompt asks for - a character gets face, hair and garments, an environment gets surfaces and light, a style gets grain and grade.
subject_3_nameoptSTRINGOptional name for row 3, e.g. 'Aria Voss' or 'the loading bay'. Blank is fine - MiniMax's own example names nobody.
subject_3_descriptionoptSTRINGWhat row 3 looks like: the features the target video must keep. Type here and this row becomes <Subject 3>, and subject_3_image becomes <Picture 3>. An empty row is not a subject.
subject_4_kindoptCOMBOcharacterWhat subject_4_image / row 4 IS. character and environment are the two everyone needs; the rest are the guide's other reusable kinds. The kind picks the features the prompt asks for - a character gets face, hair and garments, an environment gets surfaces and light, a style gets grain and grade.
subject_4_nameoptSTRINGOptional name for row 4, e.g. 'Aria Voss' or 'the loading bay'. Blank is fine - MiniMax's own example names nobody.
subject_4_descriptionoptSTRINGWhat row 4 looks like: the features the target video must keep. Type here and this row becomes <Subject 4>, and subject_4_image becomes <Picture 4>. An empty row is not a subject.
subject_5_kindoptCOMBOcharacterWhat subject_5_image / row 5 IS. character and environment are the two everyone needs; the rest are the guide's other reusable kinds. The kind picks the features the prompt asks for - a character gets face, hair and garments, an environment gets surfaces and light, a style gets grain and grade.
subject_5_nameoptSTRINGOptional name for row 5, e.g. 'Aria Voss' or 'the loading bay'. Blank is fine - MiniMax's own example names nobody.
subject_5_descriptionoptSTRINGWhat row 5 looks like: the features the target video must keep. Type here and this row becomes <Subject 5>, and subject_5_image becomes <Picture 5>. An empty row is not a subject.
subject_6_kindoptCOMBOcharacterWhat subject_6_image / row 6 IS. character and environment are the two everyone needs; the rest are the guide's other reusable kinds. The kind picks the features the prompt asks for - a character gets face, hair and garments, an environment gets surfaces and light, a style gets grain and grade.
subject_6_nameoptSTRINGOptional name for row 6, e.g. 'Aria Voss' or 'the loading bay'. Blank is fine - MiniMax's own example names nobody.
subject_6_descriptionoptSTRINGWhat row 6 looks like: the features the target video must keep. Type here and this row becomes <Subject 6>, and subject_6_image becomes <Picture 6>. An empty row is not a subject.
music_videooptBOOLEANfalseWrite the prompt as a music video. non_diegetic_music becomes the lead audio section instead of 'N/A', overall_soundscape thins to what is audible under the track, cuts are described as landing on the beat, and performance to camera becomes the action. Put the sung words in lyrics and the track in music_description.
music_sourceoptCOMBOautoWhere the music comes from. auto: the song is declared as <Audio 1> when audio is connected, on the assumption that the same file reaches H3. generate_score: H3 invents the track. reuse_audio_1: the track reaches H3 as <Audio 1> even with nothing wired here. Declaring the reuse is what adds 'audio reuse' to the task type.
lyricsoptSTRINGExact sung words, in their original language. They must reach a <d>[Language] ...</d> block in the shot where they are heard, or the run retries once. Blank means the mouth moves to the music with no intelligible words - H3 invents nonsense syllables if you ask for singing without giving it any.
music_descriptionoptSTRINGThe track, for non_diegetic_music: genre, instrumentation, tempo or BPM, and how it develops. Example: 'downtempo synthwave, ~92 BPM, analog pad and gated drums, the filter opens into the chorus at the second cut'. Blank lets the model infer it from the attached audio, or from the visuals if there is none.

Outputs (18)

NameTypeDescription
h3_promptSTRINGThe H3 prompt.
duration_secondsFLOATTarget duration. With snap_duration_to_h3_grid on this is the length H3 really produces, not the length asked for.
fpsINTFrame rate of the source clip, rounded.
analysis_jsonSTRINGEverything the run decided, including both durations.
h3_checkpoint_hintSTRINGWhich H3 checkpoint this prompt is written for.
ref_image_1IMAGEWhatever you plugged into subject_1_image, untouched. Wire it to the sampler's ref_image_1 socket. A gap stays a gap: the prompt says <Picture 1> for this slot, so compacting it here is exactly the mismatch this node exists to stop.
ref_image_2IMAGEWhatever you plugged into subject_2_image, untouched. Wire it to the sampler's ref_image_2 socket. A gap stays a gap: the prompt says <Picture 2> for this slot, so compacting it here is exactly the mismatch this node exists to stop.
ref_image_3IMAGEWhatever you plugged into subject_3_image, untouched. Wire it to the sampler's ref_image_3 socket. A gap stays a gap: the prompt says <Picture 3> for this slot, so compacting it here is exactly the mismatch this node exists to stop.
ref_image_4IMAGEWhatever you plugged into subject_4_image, untouched. Wire it to the sampler's ref_image_4 socket. A gap stays a gap: the prompt says <Picture 4> for this slot, so compacting it here is exactly the mismatch this node exists to stop.
ref_image_5IMAGEWhatever you plugged into subject_5_image, untouched. Wire it to the sampler's ref_image_5 socket. A gap stays a gap: the prompt says <Picture 5> for this slot, so compacting it here is exactly the mismatch this node exists to stop.
ref_image_6IMAGEWhatever you plugged into subject_6_image, untouched. Wire it to the sampler's ref_image_6 socket. A gap stays a gap: the prompt says <Picture 6> for this slot, so compacting it here is exactly the mismatch this node exists to stop.
ref_videoIMAGEThe clip as IMAGE frames, which is what the sampler's ref_video_ socket takes. Empty in base mode, which has no <Video 1>.
ref_video_audioAUDIOThe wired audio, for when the clip's own track is reused. Connect this OR ref_audio, not both. Empty in base mode.
ref_audioAUDIOThe same audio again, for when only its timbre or beat is referenced. Connect this OR ref_video_audio, not both.
widthINTSampler width: the video's framing if there is one, else the first wired picture's, on H3's canvas grid.
heightINTSampler height, from the same source as width.
lengthINTFrame count on H3's 17k+5 grid. Wire it to the sampler's length.
label_mapSTRINGWhat each <Picture i> / <Video k> / <Audio j> tag will refer to once the sampler numbers them. Read it against the prompt.