Nodes/ComfyUI-MiniMax-H3-Prompt-Enhancer/MiniMax H3 GGUF Prompt Enhancer
ComfyUI Node

MiniMax H3 GGUF Prompt Enhancer

Run an existing GGUF through a managed llama-server bound to loopback. No binary or model is downloaded, and the server is terminated after every queued invocation.

By hyukudan·Created 20 days ago·Updated a day ago· 15
MiniMax H3 GGUF Prompt Enhancer
    • enhanced_prompt
    • validation_report
    • enhancement_manifest
    • duration_seconds
    • aspect_ratio
    • treatment_warnings
    • width
    • height
    basic_prompt
    modeauto
    duration_seconds5.00
    reference_context
    llama_server_path
    gguf_model_path
    registered_model_dirs
    gpu_layersauto
    context_size32768
    threads0
    temperature0.20
    max_tokens8192
    request_timeout300
    startup_timeout180
    repair_attempts2
    disable_thinkingtrue
    creative_latitudeenhanced_production
    keep_server_loadedfalse
    ambience_foley_policyauto
    background_score_policyfollow_prompt
    instrumental_description
    voice_performanceaudible
    aspect_ratioauto
    media_manifest
    multishot_shot_count0
    frame_count0
    multishot_identity_lock
    multishot_voice_lock
    multishot_setting_lock
    show_advanced_controlsfalse
    creative_treatment_json
    shot_plan_json
    cinematography_json
    instrumental_stylenone
    acoustic_spacenone
    dialogue_coverageoff
    always_re_enhancefalse
    delivery_targetlocal
    dialogue_languageauto
    visual_style_presetnone
    target_megapixels0.00
    editing_intentnone
    lora_trigger_words
    CategoryMiniMax H3/Prompting

    Inputs (43)

    NameTypeDefaultDescription
    basic_promptSTRING
    modeCOMBOauto7 options: auto, t2va, i2va, fl2va, l2va, ref2va, +1
    duration_secondsFLOAT5.004–1504-150 seconds. H3 was trained around 5-15 seconds; longer generations are experimental and require much more memory.
    reference_contextSTRINGOptional plain-language notes describing referenced pictures, videos, audio, identities, or roles. Usually needed only for Ref2VA.
    llama_server_pathSTRINGExisting llama-server executable; never downloaded automatically
    gguf_model_pathSTRINGExisting GGUF under a registered model directory
    registered_model_dirsSTRINGOptional additional roots separated by the OS path separator; ComfyUI and LM Studio model roots are automatic
    gpu_layersSTRINGautoauto, all, -1, or an exact layer count
    context_sizeINT327680–1310720 uses the safe 32768-token default
    threadsINT00–2560 uses llama-server's default
    temperatureFLOAT0.200–2
    max_tokensINT8192512–32768
    request_timeoutINT30010–1800
    startup_timeoutINT1800–18000 uses the safe 180-second default
    repair_attemptsINT20–4
    disable_thinkingBOOLEANtrue
    creative_latitudeCOMBOenhanced_productionHow far beyond your text the writer may go. verbatim_source: none - keep your wording, facts and terseness as written; only reformat into H3 sections, apply the selected style and translate delivery marks. conservative_grounded: only the minimum structure the H3 mode requires. enhanced_production: resolve unspecified production decisions - composition, blocking, lighting, micro-performance. invented_production: treat your text as a premise and build the world around it. Quoted dialogue, reference identities, duration, shot count, ending and gore level stay locked at every level.
    keep_server_loadedBOOLEANfalse
    ambience_foley_policyoptCOMBOautoScene sounds other than speech or music: ambience plus physical action sounds such as footsteps, clothing, doors, impacts, and engines.
    background_score_policyoptCOMBOfollow_prompt3 options: follow_prompt, add_instrumental, off
    instrumental_descriptionoptSTRINGDescribe concrete instrumentation, tempo, rhythm, and dynamics; mood words are translated into audible parameters.
    voice_performanceoptCOMBOaudible3 options: audible, silent_mouth_acting_experimental, none
    aspect_ratiooptCOMBOauto7 options: auto, 21:9, 16:9, 4:3, 1:1, 3:4, +1
    media_manifestoptSTRINGAdvanced structured JSON for connected reference media.
    multishot_shot_countoptINT00–64
    frame_countoptINT00–3600Leave 0 to use Duration. A nonzero exact count must follow 17 × n + 5. Above about 362 frames (~15 s) is experimental.
    multishot_identity_lockoptSTRING
    multishot_voice_lockoptSTRING
    multishot_setting_lockoptSTRING
    show_advanced_controlsoptBOOLEANfalseShow structured reference metadata and exact frame controls
    creative_treatment_jsonoptSTRINGStable schema-v2 storage for genre, visual language, world aesthetic, and tone. Legacy v1 remains runtime-compatible; blank is neutral.
    shot_plan_jsonoptSTRINGOptional authoritative shot plan. Schema v1 remains compatible; v2 adds generations, presence, states, environments and start/path/end camera. Blank preserves automatic planning.
    cinematography_jsonoptSTRINGOptional schema-v2 manual color, camera, optics, focus, texture, and motion-rendering controls. Legacy v1 remains runtime-compatible; blank is neutral.
    instrumental_styleoptCOMBOnoneWhen instrumental score is enabled, adapt its arrangement to this musical language while preserving compatible user direction.
    acoustic_spaceoptCOMBOnoneDiegetic sound space for the permitted ambience, foley, and voices. It renders existing sounds; it never adds a source.
    dialogue_coverageoptCOMBOoffKeep every speaking character's mouth and eyes unobstructed, in focus, and framed at medium close-up or tighter for the whole line.
    always_re_enhanceoptBOOLEANfalseRe-run the LLM on every queue even when the inputs are unchanged. Disabled reuses the cached enhancement, so requeueing an unchanged prompt no longer forces the H3 sampler to regenerate the video.
    delivery_targetoptCOMBOlocalAPI v2 makes the 7000-character text-block limit repairable and hard.
    dialogue_languageoptCOMBOautoTarget dialogue language. 'auto' automatically detects language from prompt context/dialogue.
    visual_style_presetoptCOMBOnoneQuick visual style preset. When selected, automatically applies this visual language unless overridden in creative treatment JSON.
    target_megapixelsoptFLOAT0.00Target resolution in Megapixels (MP), e.g. 0.2, 0.3, 0.5, 0.92 (720p), 2.0 (1080p). Leave 0.0 for standard defaults; Custom accepts any positive finite value.
    editing_intentoptCOMBOnoneQuick video editing intent preset for Ref2VA (Character Swap, Wardrobe Transfer, Voice/Dialogue Swap, Background Change, Motion Transfer, Custom Editing). Automatically enforces video editing summary and retention policies.
    lora_trigger_wordsoptSTRINGTrigger tokens for the LoRAs loaded elsewhere in the graph. Appended verbatim to the end of the description after enhancement and validation, so they never pass through the LLM and survive character for character.

    Outputs (8)

    NameTypeDescription
    enhanced_promptSTRING
    validation_reportSTRING
    enhancement_manifestSTRING
    duration_secondsFLOAT
    aspect_ratioSTRING
    treatment_warningsSTRING
    widthINT
    heightINT