MiniMax H3 Audio Prompt
Turn a music idea into the exact T2VA prompt MiniMax H3 actually wants
- prompt
MiniMax H3 is an omni-modal model - text, image, video, and audio all go through one transformer, and the audio comes out with the picture instead of being bolted on afterwards. That's great, but it means the text encoder has opinions about prompt format. It doesn't want "make me a chill beat." It wants a structured T2VA prompt with named fields. MiniMaxH3AudioPrompt is a formatter that builds that string from plain-English boxes you can actually read.
It's a one-output utility node: three or four text fields in, one prompt STRING out. You wire that into MiniMaxH3TextToAudio.prompt, and that's the whole job.
How it works
Peek at the source and it's honestly just string assembly, but the details matter. It composes the three fixed T2VA fields in H3's expected order:
integrated_multimodal_description- yourscene, wrapped in[Shot 1], plus a duration hint ifsecondsis setoverall_soundscape- ambient/room tonenon_diegetic_music- the actual music description
Field names and order are fixed; the author is explicit that you shouldn't invent alternate keys. If overall_soundscape is left blank it falls back to N/A, which is the correct way to say "no ambience" in this format.
The inputs that matter
non_diegetic_music- the one you'll actually spend time on. Instruments, tempo, dynamics: "warm lo-fi hip-hop beat at 86 BPM, dusty vinyl crackle, mellow electric piano chords…". This is the only field that can't be empty - validation hard-fails on a blank string.scene- a visual description, even though you're only making audio. H3 correlates music with the described picture, so a scene helps steer mood. The default ("cinematic night city skyline…") is a good template.overall_soundscape- ambient sound like "distant traffic hush, light rain on pavement". Leave asN/Aif you want clean BGM.seconds- here's the trap: this is only a text hint. It bakes "Duration is approximately 15.00 seconds" into the prompt. The actual clip length is decided byMiniMaxH3TextToAudio.seconds. Change this one, nothing about the output's length changes - don't be confused when it doesn't.
Install and the wider workflow
Install the pack once - ComfyUI Manager → Custom Nodes → search "ComfyUI-MiniMaxH3Audio", or:
cd ComfyUI/custom_nodes
git clone https://github.com/t22m003/ComfyUI-MiniMaxH3Audio
Then restart ComfyUI. No pip dependencies - pyproject.toml lists none. What you do need is the H3 stack itself: the minimax_h3_fl2va_bf16 UNet, the qwen3vl_32b_minimax_h3_int8_convrot CLIP (loaded with type=minimax), and the minimax_h3_audio_vae_fp32 VAE, all from ComfyUI's official MiniMax H3 support. A working chain looks like: CLIPLoader → MiniMaxH3AudioPrompt → MiniMaxH3TextToAudio → BasicGuider → SamplerCustomAdvanced → VAEDecodeAudio → SaveAudioMP3.
Two gotchas worth knowing before you fight them. BasicGuider has no CFG, so negative prompts are ignored - put your exclusions ("no vocals") directly in the music text, which is what the default does. And the H3 weights run under the MiniMax H3 Community License, which excludes the US, EU, UK, and South Korea - worth a glance at the terms before you build a workflow on it.
This node won't do anything on its own; it's the boring, load-bearing glue that keeps you from hand-typing a fragile prompt format. Pair it with its sibling MiniMaxH3TextToAudio and you've got a real text-to-music pipeline.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| non_diegetic_music | STRING | warm lo-fi hip-hop beat at 86 BPM, dusty vinyl crackle, mellow electric piano chords, deep rounded bass, soft snare brushes, gentle swing groove, stereo width, no vocals | 楽器・テンポ・ダイナミクス(観客だけが聴く BGM) |
| overall_soundscape | STRING | N/A | 環境音の要約。無ければ N/A |
| scene | STRING | Cinematic night city skyline with soft neon reflections on wet asphalt; no people on screen; the camera holds a slow static wide shot while the soundtrack carries the emotion. | integrated_multimodal_description に入れる情景(黒画面より情景込みが有利な場合あり) |
| seconds | FLOAT | 15.05.16–150 | 尺ヒント(プロンプト文言にのみ反映。実際の latent 尺は TextToAudio.seconds) |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| prompt | STRING | — |