ComfyUI Node
MiniMax H3 GGUF Prompt Enhancer
Run an existing GGUF through a managed llama-server bound to loopback. No binary or model is downloaded, and the server is terminated after every queued invocation.
MiniMax H3 GGUF Prompt Enhancer
- enhanced_prompt
- validation_report
- enhancement_manifest
- duration_seconds
- aspect_ratio
- treatment_warnings
- width
- height
◄basic_prompt►
◄modeauto►
◄duration_seconds5.00►
◄reference_context►
◄llama_server_path►
◄gguf_model_path►
◄registered_model_dirs►
◄gpu_layersauto►
◄context_size32768►
◄threads0►
◄temperature0.20►
◄max_tokens8192►
◄request_timeout300►
◄startup_timeout180►
◄repair_attempts2►
◄disable_thinkingtrue►
◄creative_latitudeenhanced_production►
◄keep_server_loadedfalse►
◄ambience_foley_policyauto►
◄background_score_policyfollow_prompt►
◄instrumental_description►
◄voice_performanceaudible►
◄aspect_ratioauto►
◄media_manifest►
◄multishot_shot_count0►
◄frame_count0►
◄multishot_identity_lock►
◄multishot_voice_lock►
◄multishot_setting_lock►
◄show_advanced_controlsfalse►
◄creative_treatment_json►
◄shot_plan_json►
◄cinematography_json►
◄instrumental_stylenone►
◄acoustic_spacenone►
◄dialogue_coverageoff►
◄always_re_enhancefalse►
◄delivery_targetlocal►
◄dialogue_languageauto►
◄visual_style_presetnone►
◄target_megapixels0.00►
◄editing_intentnone►
◄lora_trigger_words►
CategoryMiniMax H3/Prompting
Inputs (43)
| Name | Type | Default | Description |
|---|---|---|---|
| basic_prompt | STRING | — | |
| mode | COMBO | auto | 7 options: auto, t2va, i2va, fl2va, l2va, ref2va, +1 |
| duration_seconds | FLOAT | 5.004–150 | 4-150 seconds. H3 was trained around 5-15 seconds; longer generations are experimental and require much more memory. |
| reference_context | STRING | Optional plain-language notes describing referenced pictures, videos, audio, identities, or roles. Usually needed only for Ref2VA. | |
| llama_server_path | STRING | Existing llama-server executable; never downloaded automatically | |
| gguf_model_path | STRING | Existing GGUF under a registered model directory | |
| registered_model_dirs | STRING | Optional additional roots separated by the OS path separator; ComfyUI and LM Studio model roots are automatic | |
| gpu_layers | STRING | auto | auto, all, -1, or an exact layer count |
| context_size | INT | 327680–131072 | 0 uses the safe 32768-token default |
| threads | INT | 00–256 | 0 uses llama-server's default |
| temperature | FLOAT | 0.200–2 | — |
| max_tokens | INT | 8192512–32768 | — |
| request_timeout | INT | 30010–1800 | — |
| startup_timeout | INT | 1800–1800 | 0 uses the safe 180-second default |
| repair_attempts | INT | 20–4 | — |
| disable_thinking | BOOLEAN | true | — |
| creative_latitude | COMBO | enhanced_production | How far beyond your text the writer may go. verbatim_source: none - keep your wording, facts and terseness as written; only reformat into H3 sections, apply the selected style and translate delivery marks. conservative_grounded: only the minimum structure the H3 mode requires. enhanced_production: resolve unspecified production decisions - composition, blocking, lighting, micro-performance. invented_production: treat your text as a premise and build the world around it. Quoted dialogue, reference identities, duration, shot count, ending and gore level stay locked at every level. |
| keep_server_loaded | BOOLEAN | false | — |
| ambience_foley_policyopt | COMBO | auto | Scene sounds other than speech or music: ambience plus physical action sounds such as footsteps, clothing, doors, impacts, and engines. |
| background_score_policyopt | COMBO | follow_prompt | 3 options: follow_prompt, add_instrumental, off |
| instrumental_descriptionopt | STRING | Describe concrete instrumentation, tempo, rhythm, and dynamics; mood words are translated into audible parameters. | |
| voice_performanceopt | COMBO | audible | 3 options: audible, silent_mouth_acting_experimental, none |
| aspect_ratioopt | COMBO | auto | 7 options: auto, 21:9, 16:9, 4:3, 1:1, 3:4, +1 |
| media_manifestopt | STRING | Advanced structured JSON for connected reference media. | |
| multishot_shot_countopt | INT | 00–64 | — |
| frame_countopt | INT | 00–3600 | Leave 0 to use Duration. A nonzero exact count must follow 17 × n + 5. Above about 362 frames (~15 s) is experimental. |
| multishot_identity_lockopt | STRING | — | |
| multishot_voice_lockopt | STRING | — | |
| multishot_setting_lockopt | STRING | — | |
| show_advanced_controlsopt | BOOLEAN | false | Show structured reference metadata and exact frame controls |
| creative_treatment_jsonopt | STRING | Stable schema-v2 storage for genre, visual language, world aesthetic, and tone. Legacy v1 remains runtime-compatible; blank is neutral. | |
| shot_plan_jsonopt | STRING | Optional authoritative shot plan. Schema v1 remains compatible; v2 adds generations, presence, states, environments and start/path/end camera. Blank preserves automatic planning. | |
| cinematography_jsonopt | STRING | Optional schema-v2 manual color, camera, optics, focus, texture, and motion-rendering controls. Legacy v1 remains runtime-compatible; blank is neutral. | |
| instrumental_styleopt | COMBO | none | When instrumental score is enabled, adapt its arrangement to this musical language while preserving compatible user direction. |
| acoustic_spaceopt | COMBO | none | Diegetic sound space for the permitted ambience, foley, and voices. It renders existing sounds; it never adds a source. |
| dialogue_coverageopt | COMBO | off | Keep every speaking character's mouth and eyes unobstructed, in focus, and framed at medium close-up or tighter for the whole line. |
| always_re_enhanceopt | BOOLEAN | false | Re-run the LLM on every queue even when the inputs are unchanged. Disabled reuses the cached enhancement, so requeueing an unchanged prompt no longer forces the H3 sampler to regenerate the video. |
| delivery_targetopt | COMBO | local | API v2 makes the 7000-character text-block limit repairable and hard. |
| dialogue_languageopt | COMBO | auto | Target dialogue language. 'auto' automatically detects language from prompt context/dialogue. |
| visual_style_presetopt | COMBO | none | Quick visual style preset. When selected, automatically applies this visual language unless overridden in creative treatment JSON. |
| target_megapixelsopt | FLOAT | 0.00 | Target resolution in Megapixels (MP), e.g. 0.2, 0.3, 0.5, 0.92 (720p), 2.0 (1080p). Leave 0.0 for standard defaults; Custom accepts any positive finite value. |
| editing_intentopt | COMBO | none | Quick video editing intent preset for Ref2VA (Character Swap, Wardrobe Transfer, Voice/Dialogue Swap, Background Change, Motion Transfer, Custom Editing). Automatically enforces video editing summary and retention policies. |
| lora_trigger_wordsopt | STRING | Trigger tokens for the LoRAs loaded elsewhere in the graph. Appended verbatim to the end of the description after enhancement and validation, so they never pass through the LLM and survive character for character. |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| enhanced_prompt | STRING | — |
| validation_report | STRING | — |
| enhancement_manifest | STRING | — |
| duration_seconds | FLOAT | — |
| aspect_ratio | STRING | — |
| treatment_warnings | STRING | — |
| width | INT | — |
| height | INT | — |