Nodes/Trent Nodes/H3 Codex Promptor (Trent)
ComfyUI Node

H3 Codex Promptor (Trent)

Your ChatGPT subscription writes the video prompt

By TrentHunter82·Created 10 months ago·Updated a day ago· 42
H3 Codex Promptor (Trent)
  • first_frame
  • last_frame
  • reference_images
  • video_frames
  • audio
  • h3_prompt
  • checkpoint_hint
  • validation_report
  • overall_soundscape
  • non_diegetic_music
◄creative_brief►
◄duration_seconds6.0►
◄modeauto►
◄actiongenerate►
◄prompt_text►
◄refinement►
◄reference_roles►
◄context►
◄dialogue►
◄source_soundscape►
◄source_music►
◄sound_log►
◄fps24.00►
◄max_frames8►
◄model►
◄reasoning_effortdefault►
◄revision0►
◄timeout_seconds600►
◄cut_times—►

MiniMax H3 doesn't take a paragraph. It takes a document - either the three-field base format (integrated_multimodal_description, overall_soundscape, non_diegetic_music) or the six-section Ref2VA one, which adds subject definitions, a retention analysis and <Picture 1> / <Video 1> / <Audio 1> labels. Get the structure wrong and you're debugging a 33B video model when the real problem was your formatting.

MiniMax publishes an official prompt-writing guide, so Trent's answer is: don't paraphrase it in a system prompt, load the real thing. H3 Codex Promptor sends your brief and reference frames to Codex CLI (the one your ChatGPT subscription covers), gets a prompt written against MiniMax's own skill document, checks it locally, and drops it into an editable box wired to your sampler. No API key, despite the name.

How it works

The node spawns codex app-server --listen stdio:// as a private subprocess in a throwaway temp dir, with OPENAI_API_KEY and CODEX_API_KEY stripped out of the environment. The session is pinned to ChatGPT sign-in, so it bills your subscription and can't quietly fall back to API billing. Shell, web search and multi-agent are disabled - the model gets the skill, your images and a schema, nothing else.

The system prompt is MiniMax's official h3-prompt-writing skill plus your mode's reference guide - downloaded at a pinned commit on first use, SHA-256 verified, cached under ComfyUI/user/default/trentnodes/h3_codex/skills/. Codex returns the body plus an assumptions list; deterministic local checks then run on formatting, shot timing and reference labels. Fail them and the model gets exactly one corrective turn. Whatever comes back is returned as-is with the diagnostics attached - nothing is silently rewritten.

The inputs that matter

creative_brief is the box you fill in. duration_seconds (default 6) must match the clip you'll generate; the official guide is written around 4–15s.

mode on auto reads your wiring: text only → T2VA, first_frame → I2VA, last_frame → L2VA, both → FL2VA, and any of reference_images, video_frames or audio → Ref2VA. Note that last one: wiring audio alone flips you to the six-section Ref2VA format - correct H3 behaviour, but it surprises people who only wanted to declare <Audio 1>. Base modes want exact picture counts (1 for I2VA/L2VA, 2 for FL2VA), and labelling follows first frame, last frame, then reference_images in batch order. Wire H3 in that same order or the prompt describes the wrong picture.

action is generate / refine / locked, and prompt_text is the editable result, saved inside the workflow. refinement is the change request for an existing prompt. video_frames takes a VHS IMAGE batch and sends sampled keyframes (max_frames, default 8, timestamps from fps) - never the video, never its audio. reference_roles, context and dialogue explain what each reference supplies and carry the exact words H3 should speak; source_soundscape, source_music and sound_log are built for H3 Audio Soundscaper's outputs, because this node cannot hear anything.

Outputs

h3_prompt goes into your H3 text input. checkpoint_hint names the shape you need - MiniMax-H3-Base-FL2VA for base modes, MiniMax-H3-Base-Ref2VA for Ref2VA. validation_report holds the diagnostics, frame times, assumptions and skill revision; right-click → Show Codex validation report shows the same. overall_soundscape and non_diegetic_music are split out of the body for workflows that want those as separate text inputs.

Installing it

cd ComfyUI/custom_nodes
git clone https://github.com/TrentHunter82/TrentNodes.git
cd TrentNodes && pip install -r requirements.txt

Or search "Trent Nodes" in ComfyUI Manager. The requirements file is the whole pack's (opencv, matplotlib, psd-tools, timm, transnetv2-pytorch, the API SDKs) - a heavy install for a node gallery.

The real prerequisite isn't pip:

# same environment/user ComfyUI runs as (inside WSL, if that's where ComfyUI lives)
codex login

The CLI must be on ComfyUI's PATH or at ~/.local/bin/codex, otherwise set TRENT_CODEX_BIN. Restart ComfyUI, refresh the browser, and start from example_workflows/H3_Codex_Promptor.json.

Where people get burned

Nothing changed on Run. Locked mode is the default after a successful generate, and locked prompts ignore edits to your brief and references by design. Hit Generate, or use New variation.

A seed variation returned the identical prompt. That's the cache: results are keyed on brief, images, mode, duration, model, effort, skill version and revision, so a new H3 seed spends no Codex usage. Changed your Codex model, or want a fresh draft? Bump revision.

Login expired. Run codex login as the ComfyUI user; a cached codex login status isn't proof.

Generate/Refine buttons do nothing inside a subgraph. They're main-canvas only; set action and press Run. And the first unlocked run needs GitHub once for the skill - after that it's local.

Two things before you invest: your brief, context and sampled frames go to Codex while H3 renders locally, and the H3 Community License excludes the US, EU, UK and South Korea - a fair number of readers aren't licensed to run the local weights at all.

CategoryTrent/VLM

Inputs (24)

NameTypeDefaultDescription
creative_briefSTRINGDescribe what should happen. Codex uses your ChatGPT subscription to write an H3 prompt.
duration_secondsFLOAT6.00.1–600Match the actual H3 clip duration. The official guide targets 4–15 seconds.
modeCOMBOautoAuto: first/last frame inputs choose base modes; general references choose Ref2VA. Override to match your checkpoint.
actionCOMBOgenerateGenerate from your brief, refine the editor using your request, or use the locked editor with no Codex call.
prompt_textSTRINGEditable result. Lock uses this exact text. Saved inside the workflow.
refinementSTRINGWhat to change in the current prompt, e.g. slower camera movement; preserve dialogue exactly.
first_frameoptIMAGEOpening keyframe; each batch image receives a Picture label.
last_frameoptIMAGEEnding keyframe. With first_frame, auto chooses FL2VA.
reference_imagesoptIMAGEReference batch in the same order you wire into H3. Auto selects Ref2VA.
video_framesoptIMAGEVideo as a frame batch (VHS IMAGE output). Codex sees sampled frames, not the full video or its audio.
audiooptAUDIOMarks Audio 1 as available to your H3 workflow. Metadata only; connect Soundscaper analysis below to describe its content.
reference_rolesoptSTRINGExplain what each reference supplies, e.g. Picture 1 identity; Video 1 motion; Audio 1 voice.
contextoptSTRINGExtra scene, character, story, style, or continuity context.
dialogueoptSTRINGExact spoken words, lyrics, and speaker information to preserve.
source_soundscapeoptSTRINGConnect H3 Audio Soundscaper overall_soundscape here.
source_musicoptSTRINGConnect H3 Audio Soundscaper non_diegetic_music here.
sound_logoptSTRINGConnect H3 Audio Soundscaper sound_log or a transcript here.
fpsoptFLOAT24.000.1–240Source video fps, used for sampled-frame timestamps.
max_framesoptINT82–32Target sampled video frames. With cut_times, raised as needed to inspect opening/middle/ending frames in every shot.
modeloptSTRINGEmpty uses your Codex configured model. Otherwise enter a model available to your subscription.
reasoning_effortoptCOMBOdefault5 options: default, low, medium, high, xhigh
revisionoptINT00–2147483647Increase to request another draft with the same inputs. Keep fixed when rendering H3 seed variations.
timeout_secondsoptINT60030–3600—
cut_timesoptSTRINGConnect Cut Detective cut_times, shot_table or cuts_json. Uses exact shot starts; duration_seconds ends the final shot. Use the same source video and fps.

Outputs (5)

NameTypeDescription
h3_promptSTRING—
checkpoint_hintSTRING—
validation_reportSTRING—
overall_soundscapeSTRING—
non_diegetic_musicSTRING—