Nodes/IAMCCS-nodes/IAMCCS Prompter — MiniMax H3 Screenplay
ComfyUI Node

IAMCCS Prompter — MiniMax H3 Screenplay

Write the scene in sections, not one giant prompt

By IAMCCS·Created 11 months ago·Updated 4 days ago· 113
IAMCCS Prompter — MiniMax H3 Screenplay
  • cine_linx
  • cine_linx
  • final_prompt
  • project_json
  • report
project_data{"schema": "iamccs.minimax_h3.prompter_project", "schema_version": 3, "project_name": "Platform at Blue Hour", "task_mode": "t2va", "injection_target": "global", "writing_mode": "guided", "merge_policy": "replace", "ai_direction": "", "ai_scope": "active_field", "ai_visual_roles": {}, "sections": {"scene": "A rain-polished railway platform before sunrise. <Subject 1>, a tired courier in a charcoal coat, waits beside a silver case while an empty train approaches through blue mist.", "shot_list": "0.00-2.00s: hold a medium-wide profile. 2.00-4.50s: the train enters and throws moving reflections across the platform. 4.50s-end: <Subject 1> turns toward camera and grips the case.", "acting": "Restrained performance: shoulders tense first, then the eyes react, then one deliberate turn. Preserve natural blinking and breathing.", "dialogue": "<Subject 1> (S1): <d>[English] Not this train.</d>", "light_and_image": "Cool dawn ambience, practical sodium lamps, wet reflections, restrained contrast, realistic skin texture, cinematic depth without artificial glow.", "camera": "One slow lateral tracking move at chest height with mild foreground parallax; no cut and no change of lens language.", "production_sound": "Distant rail vibration, light rain on metal roofing, one approaching brake squeal, coat movement, clear close dialogue with matching platform reverb.", "non_diegetic_music": "A sparse low cello pulse enters only after the train becomes visible; keep it separate from the physical scene sound.", "negatives": "No identity drift, no duplicate people, no wardrobe change, no warped hands, no sudden zoom, no jump cut, no subtitles, no logo.", "reference_use": "Use <Picture 1> as the complete opening-frame authority for identity, wardrobe, composition, lens perspective, lighting direction and visible environment. Animate from it rather than redesigning it.", "identity_continuity_locks": "Keep <Subject 1>'s face, hairline, coat, silver case, body proportions and screen side unchanged. Preserve the platform geometry and time of day.", "boundary_frames": "Open exactly on <Picture 1> and arrive naturally at <Picture 2> as the final composition. Treat both pictures as full-frame boundaries, not loose style references.", "action": "The character crosses the connected space in one continuous action. Movement should develop physically toward the final pose with stable identity and coherent screen direction.", "subject_definitions": "<Subject 1>: the principal performer shown in <Picture 1>; preserve face, body proportions, wardrobe and signature accessories.\n<Subject 2>: the compact silver case; preserve its shape, scale, surface marks and position relative to <Subject 1>.", "summary": "A tense cinematic beat in which <Subject 1> notices an approaching threat while protecting <Subject 2>. The result should feel observational, grounded and continuous.", "retention_analysis": "Retain identity and wardrobe from <Picture 1>. Retain the physical timing and camera rhythm from <Video 1> only where supplied. Use <Audio 1> for voice character or cadence only when it is connected; do not invent an unseen speaker.", "detailed_description": "Begin with the supplied reference composition. <Subject 1> hears the approaching train, tightens one hand around <Subject 2>, then turns with a controlled breath. Use a single lateral camera move and preserve spatial geography. If dialogue is desired: <Subject 1> (S1): <d>[English] Not this train.</d>", "overall_soundscape": "Layer the location ambience, contact sounds, movement and dialogue in chronological order. Keep perspective and reverberation consistent with camera distance; avoid wall-to-wall effects.", "v2va_subject_definitions": "<Picture 1> defines the replacement <Subject 1>: [write only visible identity, body, wardrobe, material or object facts that must be preserved].\n<Video 1> contains source <Subject 2>: [identify exactly what is being replaced]. Each additional <Picture N> may define only the named <Subject N> or continuity attribute.", "v2va_source_video_authority": "<Video 1> is the temporal source authority for duration, action timing, body or object motion, camera path, framing, occlusion order, environment and edit rhythm. Preserve those source relationships unless an interval instruction below explicitly changes one.", "v2va_replacement_retention": "Replace source <Subject 2> from <Video 1> with replacement <Subject 1> from <Picture 1>. Retain [list source environment, secondary subjects, interactions, contact points, lighting response and camera behavior]. Change only [list the requested identity, object, clothing or appearance attributes].", "v2va_interval_edits": "[Define source-time intervals from <Video 1>, for example 00:00.00-00:02.50, and state the visible replacement action or retained event in each interval. Leave this field untimed when the edit applies uniformly to the complete source video.]", "v2va_sound_policy": "[State whether connected source-video audio is retained, replaced, muted or supplemented. Name <Audio 1> only when an audio reference is actually connected. Keep dialogue wording, lip timing, contact sounds and ambience consistent with the chosen policy.]", "v2va_exclusions": "Do not change unselected subjects, source environment, camera trajectory, duration, occlusion order or interactions. No identity blending between <Subject 1> and <Subject 2>, duplicate replacement, geometry drift, temporal jump, subtitle, logo or invented reference.", "audio_drive_contract": "Treat the connected custom audio as the timing authority. Preserve its order, pauses, breaths and duration; do not invent, remove or reorder speech.", "audio_subject_map": "<Subject 1> (S1): [describe the visible speaker and the identity/reference facts that must remain stable].", "audio_scene_intent": "[Describe location, time, dramatic purpose and the visible starting situation.]", "audio_timed_performance": "[Map audible phrases, pauses and breaths to chronological facial expression, gaze, gesture and body action.]", "audio_dialogue_map": "<Subject 1> (S1): <d>[Language] ...</d>", "audio_visual_sync": "[Describe visible mouth articulation, breath, contact or musical actions that must synchronize with the connected audio.]", "audio_camera_sync": "[Describe one coherent framing and camera move that supports the timed performance without hiding the speaker.]", "audio_environment": "[Describe only environmental ambience and contact sounds not already fixed by the custom audio.]", "audio_continuity_locks": "[List identity, wardrobe, anatomy, prop, geography, eyeline and lip-visibility facts that cannot drift.]"}}
task_modet2va
injection_targetglobal
writing_modeguided
merge_policyreplace
character_budget6800
assistant_draft

The single biggest practical limit of MiniMax H3 is that its prompt has a hard character cap, and the people getting the best results treat it like a screenplay with separate fields - scene, shot list, acting, dialogue, lighting, camera, sound - rather than one blob of prose. IAMCCS_Prompter is a structured writing desk that makes you work that way: you fill the sections, it assembles them in the order the H3 conditioning actually wants, and hands the result downstream as a clean final prompt plus a cine_linx contract.

The default project_data field is a full JSON template (schema v3) with every section pre-labeled: scene, shot_list, acting, dialogue, light_and_image, camera, production_sound, non_diegetic_music, negatives, reference_use, boundary_frames, and the V2V-specific blocks (v2va_subject_definitions, v2va_source_video_authority, v2va_interval_edits, v2va_sound_policy). You don't write JSON by hand - the node ships a UI that edits these fields, and project_json on the output carries the portable version.

The mode switch that matters

  • task_mode - t2va, i2va, fl2va, ref2va, v2va_object_swap, audio_driven. This selects which sections actually get composed, so a v2va_object_swap prompt emphasizes the replacement/source-video contract and an audio_driven one leans on the audio_drive_contract and audio_timed_performance blocks.
  • injection_target - global, local_auto, or local_1..3. When a Shotboard is downstream, this decides whether your composed prompt replaces/appends the global prompt or lands in a specific timeline slot.
  • writing_mode - manual (you write), guided (the desk enforces the structure), assistant_fill (an optional assistant_draft text fills only the empty boxes).
  • merge_policy - replace or append.
  • character_budget - it watches your character count and refuses to emit a prompt over H3's absolute limit (the error tells you exactly how many characters to cut - genuinely helpful, not cryptic).

Optional cine_linx in is the interesting one: it lets an upstream IAMCCS Cine H3 Vision Info node route analyzed visual context (what the reference frame actually shows) into the declared global/local target without embedding image tensors into the prompt node itself.

Outputs

final_prompt (the assembled plain-text prompt - the thing you wire into a standalone H3 conditioning node), cine_linx (for the Shotboard path, which also gets a visible INJECT action in the UI), project_json, and report.

The honest take

For the Shotboard workflow this is the right tool and it removes a real class of error - "I wrote the prompt in the wrong mode and the references got treated as style instead of identity." Used standalone it's just a well-structured prompt assembler, and there are cheaper ones if all you need is a text box. The value is the contract: the same structured project can drive a Shotboard injection or a direct native conditioning node, and it keeps your scene data as a reusable file rather than a widget you'll misplace.

Install via ComfyUI Manager (search "IAMCCS") or cd ComfyUI/custom_nodes && git clone https://github.com/IAMCCS/IAMCCS-nodes.git. The pack's own user guide (docs/IAMCCS_PROMPTER_MINIMAX_H3_USER_GUIDE_EN.txt) walks the two connection methods - Shotboard injection vs. direct final_prompt - and it's worth a skim before your first serious run.

CategoryIAMCCS/MiniMax H3/Prompting

Inputs (8)

NameTypeDefaultDescription
project_dataSTRING{"schema": "iamccs.minimax_h3.prompter_project", "schema_version": 3, "project_name": "Platform at Blue Hour", "task_mode": "t2va", "injection_target": "global", "writing_mode": "guided", "merge_policy": "replace", "ai_direction": "", "ai_scope": "active_field", "ai_visual_roles": {}, "sections": {"scene": "A rain-polished railway platform before sunrise. <Subject 1>, a tired courier in a charcoal coat, waits beside a silver case while an empty train approaches through blue mist.", "shot_list": "0.00-2.00s: hold a medium-wide profile. 2.00-4.50s: the train enters and throws moving reflections across the platform. 4.50s-end: <Subject 1> turns toward camera and grips the case.", "acting": "Restrained performance: shoulders tense first, then the eyes react, then one deliberate turn. Preserve natural blinking and breathing.", "dialogue": "<Subject 1> (S1): <d>[English] Not this train.</d>", "light_and_image": "Cool dawn ambience, practical sodium lamps, wet reflections, restrained contrast, realistic skin texture, cinematic depth without artificial glow.", "camera": "One slow lateral tracking move at chest height with mild foreground parallax; no cut and no change of lens language.", "production_sound": "Distant rail vibration, light rain on metal roofing, one approaching brake squeal, coat movement, clear close dialogue with matching platform reverb.", "non_diegetic_music": "A sparse low cello pulse enters only after the train becomes visible; keep it separate from the physical scene sound.", "negatives": "No identity drift, no duplicate people, no wardrobe change, no warped hands, no sudden zoom, no jump cut, no subtitles, no logo.", "reference_use": "Use <Picture 1> as the complete opening-frame authority for identity, wardrobe, composition, lens perspective, lighting direction and visible environment. Animate from it rather than redesigning it.", "identity_continuity_locks": "Keep <Subject 1>'s face, hairline, coat, silver case, body proportions and screen side unchanged. Preserve the platform geometry and time of day.", "boundary_frames": "Open exactly on <Picture 1> and arrive naturally at <Picture 2> as the final composition. Treat both pictures as full-frame boundaries, not loose style references.", "action": "The character crosses the connected space in one continuous action. Movement should develop physically toward the final pose with stable identity and coherent screen direction.", "subject_definitions": "<Subject 1>: the principal performer shown in <Picture 1>; preserve face, body proportions, wardrobe and signature accessories.\n<Subject 2>: the compact silver case; preserve its shape, scale, surface marks and position relative to <Subject 1>.", "summary": "A tense cinematic beat in which <Subject 1> notices an approaching threat while protecting <Subject 2>. The result should feel observational, grounded and continuous.", "retention_analysis": "Retain identity and wardrobe from <Picture 1>. Retain the physical timing and camera rhythm from <Video 1> only where supplied. Use <Audio 1> for voice character or cadence only when it is connected; do not invent an unseen speaker.", "detailed_description": "Begin with the supplied reference composition. <Subject 1> hears the approaching train, tightens one hand around <Subject 2>, then turns with a controlled breath. Use a single lateral camera move and preserve spatial geography. If dialogue is desired: <Subject 1> (S1): <d>[English] Not this train.</d>", "overall_soundscape": "Layer the location ambience, contact sounds, movement and dialogue in chronological order. Keep perspective and reverberation consistent with camera distance; avoid wall-to-wall effects.", "v2va_subject_definitions": "<Picture 1> defines the replacement <Subject 1>: [write only visible identity, body, wardrobe, material or object facts that must be preserved].\n<Video 1> contains source <Subject 2>: [identify exactly what is being replaced]. Each additional <Picture N> may define only the named <Subject N> or continuity attribute.", "v2va_source_video_authority": "<Video 1> is the temporal source authority for duration, action timing, body or object motion, camera path, framing, occlusion order, environment and edit rhythm. Preserve those source relationships unless an interval instruction below explicitly changes one.", "v2va_replacement_retention": "Replace source <Subject 2> from <Video 1> with replacement <Subject 1> from <Picture 1>. Retain [list source environment, secondary subjects, interactions, contact points, lighting response and camera behavior]. Change only [list the requested identity, object, clothing or appearance attributes].", "v2va_interval_edits": "[Define source-time intervals from <Video 1>, for example 00:00.00-00:02.50, and state the visible replacement action or retained event in each interval. Leave this field untimed when the edit applies uniformly to the complete source video.]", "v2va_sound_policy": "[State whether connected source-video audio is retained, replaced, muted or supplemented. Name <Audio 1> only when an audio reference is actually connected. Keep dialogue wording, lip timing, contact sounds and ambience consistent with the chosen policy.]", "v2va_exclusions": "Do not change unselected subjects, source environment, camera trajectory, duration, occlusion order or interactions. No identity blending between <Subject 1> and <Subject 2>, duplicate replacement, geometry drift, temporal jump, subtitle, logo or invented reference.", "audio_drive_contract": "Treat the connected custom audio as the timing authority. Preserve its order, pauses, breaths and duration; do not invent, remove or reorder speech.", "audio_subject_map": "<Subject 1> (S1): [describe the visible speaker and the identity/reference facts that must remain stable].", "audio_scene_intent": "[Describe location, time, dramatic purpose and the visible starting situation.]", "audio_timed_performance": "[Map audible phrases, pauses and breaths to chronological facial expression, gaze, gesture and body action.]", "audio_dialogue_map": "<Subject 1> (S1): <d>[Language] ...</d>", "audio_visual_sync": "[Describe visible mouth articulation, breath, contact or musical actions that must synchronize with the connected audio.]", "audio_camera_sync": "[Describe one coherent framing and camera move that supports the timed performance without hiding the speaker.]", "audio_environment": "[Describe only environmental ambience and contact sounds not already fixed by the custom audio.]", "audio_continuity_locks": "[List identity, wardrobe, anatomy, prop, geography, eyeline and lip-visibility facts that cannot drift.]"}}
task_modeCOMBOt2va6 options: t2va, i2va, fl2va, ref2va, v2va_object_swap, audio_driven
injection_targetCOMBOglobal5 options: global, local_auto, local_1, local_2, local_3
writing_modeCOMBOguided3 options: manual, guided, assistant_fill
merge_policyCOMBOreplace2 options: replace, append
character_budgetINT68001000–7000
cine_linxoptIAMCCS_SUPERNODE_LINXOptional upstream IAMCCS Cine H3 Vision Info. Its analyzed visual context is routed to the declared global/local targets without embedding image tensors in this node.
assistant_draftoptSTRINGOptional structured draft from any text source. In Assistant Fill mode it fills only empty structured boxes.

Outputs (4)

NameTypeDescription
cine_linxIAMCCS_SUPERNODE_LINX
final_promptSTRING
project_jsonSTRING
reportSTRING