MiniMax H3 Text to Audio
MiniMax H3's audio stream, without paying for pixels you'll throw away
- clip
- positive
- latent
MiniMax H3 is a video model that generates sound and picture together - which is exactly the problem if all you want is a music track. Run it normally and you're paying (in VRAM and seconds) for a 1344×768 picture you're going to delete. MiniMaxH3TextToAudio is the workaround: it feeds H3 a dummy video canvas shrunk to 224×128 while keeping the audio latent at full quality, so the shared DiT spends its tokens on the sound and the video becomes cheap padding. The result is text-to-music at a practical cost, decoded by core ComfyUI's VAEDecodeAudio - no separate video VAE in the loop.
What it actually does
Take a clip, a prompt, and a length in seconds, and you get two outputs:
positive- a CONDITIONING tensor, wired into aBasicGuiderlatent- an audio+video latent (a ComfyUINestedTensor), wired intoSamplerCustomAdvanced.latent_image
The prompt is tokenized and encoded on the spot (clip.tokenize + encode_from_tokens_scheduled), so you don't need a separate text-encode step.
The mechanism, ground truth
The source builds an empty NestedTensor with two branches: a video branch of zeros sized [1, 24, latent_t, h//16, w//16] at your canvas resolution, and a stereo audio branch [1, 32, 2, audio_t] at 40 fps. Your requested seconds get snapped up to the 17k+5 frame grid at 24 fps - the minimum real length is 124 frames, about 5.17 seconds, which is why seconds has a 5.16 floor. Two quirks baked into the code: batch is fixed at 1 (H3's NestedTensor path rejects bigger batches), and there's a warning logged once you cross 15.1s, because the ~5–15s window is the trained range and anything longer is experimental.
The inputs that matter
clip- fromCLIPLoader, theqwen3vl_32b_minimax_h3_int8_convrotclip loaded with typeminimax. This is the tokenizer/encoder that actually understands the prompt format.prompt- wire inMiniMaxH3AudioPrompt's output. H3 wants its structured T2VA format (integrated_multimodal_description / overall_soundscape / non_diegetic_music), not a loose sentence; feeding a raw sentence is how you get mush.seconds- default 15. The trained range is ~5–15s; past 15.1s the node warns you're outside it. This is the field that actually sets the clip length (unlike the same-named field on the prompt node).video_canvas- the dummy video resolution, a dropdown of 7:4 options from112x64up to native1344x768. Leave it at224x128: it's the default for a reason, and the point of this node is not to generate video. Bigger canvases only slow you down for audio-only work.
Install and the model files
cd ComfyUI/custom_nodes
git clone https://github.com/t22m003/ComfyUI-MiniMaxH3Audio
Then restart. (Or use ComfyUI Manager and search "ComfyUI-MiniMaxH3Audio".) There are no Python dependencies in this pack - but you need the three H3 checkpoint files sitting in your models folders: minimax_h3_fl2va_bf16 (UNet), qwen3vl_32b_minimax_h3_int8_convrot (CLIP), and minimax_h3_audio_vae_fp32 (VAE). These come from ComfyUI's official MiniMax H3 integration, and they're big.
The trap to remember
BasicGuider has no CFG, so a negative prompt does nothing. Put exclusions in the positive music text instead - "no vocals", "no drums" - that's the author's explicit recommendation. And before you commit to H3 locally, check the MiniMax H3 Community License: it excludes the US, EU, UK, and South Korea, outputs included. If you're in one of those regions, this whole local pipeline isn't licensed for you, whatever the hardware can handle.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| clip | CLIP | — | |
| prompt | STRING | H3 T2VA 形式のプロンプト(MiniMaxH3AudioPrompt の出力を推奨) | |
| seconds | FLOAT | 15.005.16–150 | 目標尺(秒)。24fps の 17k+5 グリッドへ切り上げ。最小実尺は 124f≈5.17s。15.1s 超は学習範囲外・experimental |
| video_canvas | COMBO | 224x128 | ダミー video キャンバス(1344×768 と同アスペクト 7:4 推奨。既定は較正勝者) |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | — |
| latent | LATENT | — |