Nodes/MiniMax H3 Audio/MiniMax H3 Audio Prompt
ComfyUI Node

MiniMax H3 Audio Prompt

Turn a music idea into the exact T2VA prompt MiniMax H3 actually wants

By t22m003·Created 2 months ago·Updated 2 months ago· 1
MiniMax H3 Audio Prompt
    • prompt
    ◄non_diegetic_musicwarm lo-fi hip-hop beat at 86 BPM, dusty vinyl crackle, mellow electric piano chords, deep rounded bass, soft snare brushes, gentle swing groove, stereo width, no vocals►
    ◄overall_soundscapeN/A►
    ◄sceneCinematic night city skyline with soft neon reflections on wet asphalt; no people on screen; the camera holds a slow static wide shot while the soundtrack carries the emotion.►
    ◄seconds15.0►

    MiniMax H3 is an omni-modal model - text, image, video, and audio all go through one transformer, and the audio comes out with the picture instead of being bolted on afterwards. That's great, but it means the text encoder has opinions about prompt format. It doesn't want "make me a chill beat." It wants a structured T2VA prompt with named fields. MiniMaxH3AudioPrompt is a formatter that builds that string from plain-English boxes you can actually read.

    It's a one-output utility node: three or four text fields in, one prompt STRING out. You wire that into MiniMaxH3TextToAudio.prompt, and that's the whole job.

    How it works

    Peek at the source and it's honestly just string assembly, but the details matter. It composes the three fixed T2VA fields in H3's expected order:

    • integrated_multimodal_description - your scene, wrapped in [Shot 1], plus a duration hint if seconds is set
    • overall_soundscape - ambient/room tone
    • non_diegetic_music - the actual music description

    Field names and order are fixed; the author is explicit that you shouldn't invent alternate keys. If overall_soundscape is left blank it falls back to N/A, which is the correct way to say "no ambience" in this format.

    The inputs that matter

    • non_diegetic_music - the one you'll actually spend time on. Instruments, tempo, dynamics: "warm lo-fi hip-hop beat at 86 BPM, dusty vinyl crackle, mellow electric piano chords…". This is the only field that can't be empty - validation hard-fails on a blank string.
    • scene - a visual description, even though you're only making audio. H3 correlates music with the described picture, so a scene helps steer mood. The default ("cinematic night city skyline…") is a good template.
    • overall_soundscape - ambient sound like "distant traffic hush, light rain on pavement". Leave as N/A if you want clean BGM.
    • seconds - here's the trap: this is only a text hint. It bakes "Duration is approximately 15.00 seconds" into the prompt. The actual clip length is decided by MiniMaxH3TextToAudio.seconds. Change this one, nothing about the output's length changes - don't be confused when it doesn't.

    Install and the wider workflow

    Install the pack once - ComfyUI Manager → Custom Nodes → search "ComfyUI-MiniMaxH3Audio", or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/t22m003/ComfyUI-MiniMaxH3Audio
    

    Then restart ComfyUI. No pip dependencies - pyproject.toml lists none. What you do need is the H3 stack itself: the minimax_h3_fl2va_bf16 UNet, the qwen3vl_32b_minimax_h3_int8_convrot CLIP (loaded with type=minimax), and the minimax_h3_audio_vae_fp32 VAE, all from ComfyUI's official MiniMax H3 support. A working chain looks like: CLIPLoader → MiniMaxH3AudioPrompt → MiniMaxH3TextToAudio → BasicGuider → SamplerCustomAdvanced → VAEDecodeAudio → SaveAudioMP3.

    Two gotchas worth knowing before you fight them. BasicGuider has no CFG, so negative prompts are ignored - put your exclusions ("no vocals") directly in the music text, which is what the default does. And the H3 weights run under the MiniMax H3 Community License, which excludes the US, EU, UK, and South Korea - worth a glance at the terms before you build a workflow on it.

    This node won't do anything on its own; it's the boring, load-bearing glue that keeps you from hand-typing a fragile prompt format. Pair it with its sibling MiniMaxH3TextToAudio and you've got a real text-to-music pipeline.

    Categoryaudio/minimax_h3

    Inputs (4)

    NameTypeDefaultDescription
    non_diegetic_musicSTRINGwarm lo-fi hip-hop beat at 86 BPM, dusty vinyl crackle, mellow electric piano chords, deep rounded bass, soft snare brushes, gentle swing groove, stereo width, no vocals楽器・テンポ・ダイナミクス(観客だけが聴く BGM)
    overall_soundscapeSTRINGN/A環境音の要約。無ければ N/A
    sceneSTRINGCinematic night city skyline with soft neon reflections on wet asphalt; no people on screen; the camera holds a slow static wide shot while the soundtrack carries the emotion.integrated_multimodal_description に入れる情景(黒画面より情景込みが有利な場合あり)
    secondsFLOAT15.05.16–150尺ヒント(プロンプト文言にのみ反映。実際の latent 尺は TextToAudio.seconds)

    Outputs (1)

    NameTypeDescription
    promptSTRING—