DIGIT MiniMax Video
MiniMax H3 video with native audio, from inside ComfyUI
- first_frame
- last_frame
- reference_image1
- reference_image2
- reference_image3
- reference_image4
- reference_image5
- reference_image6
- reference_image7
- reference_image8
- reference_image9
- reference_video1
- reference_video2
- reference_video3
- reference_audio1
- reference_audio2
- reference_audio3
- video
- video_paths
- status
MiniMax H3 (the lab's omni-modal video model) makes clips that come with sound already baked in - synchronized audio, not a silent video you have to score afterwards. That's the headline feature and it's the reason people route around Veo for a certain kind of job. This node gives you H3 through your choice of hosted providers (fal by default, or MUAPI), with mode auto-detection like the rest of the DIGIT video family: prompt only → text-to-video, first frame connected → image-to-video, first + last frame → interpolation, reference inputs → reference-to-video.
Worth being clear about up front: this is the API path, not the local H3 weights. H3's open-weights release has a license that excludes the US, EU, UK and South Korea - but calling it through fal or MUAPI sidesteps the local-weights license entirely. You're paying per second instead.
How it works
The node builds the request, submits it to the provider, polls until the job completes, downloads the clip, and hands it back as a VIDEO tensor plus file paths. There's a live cost strip on the node that updates as you change provider, resolution, and duration - fal's is a static price table, MUAPI's proxies a live estimate endpoint.
The inputs that matter:
- prompt - required.
- provider -
fal(default) ormuapi. Replicate stays hidden until the model is published there. fal has the safety checker and prompt expansion; MUAPI is 2K-only today. - resolution -
768P,2K,4Kon fal; 2K only on muapi. Pick 2K and leave it unless you need 4K. - aspect_ratio - fixed ratios for text-to-video (
adaptiveis rejected there); image-to-video and first/last-frame follow the source image; reference mode supportsadaptive. - duration - 4 to 15 seconds. Billed per second (fal lists ~$0.26/s at 2K).
- batch_count - 1–8 generations, all submitted before polling.
- enable_prompt_expansion and enable_safety_checker - fal only, both default on.
- first_frame / last_frame - image-to-video and interpolation. Mutually exclusive with reference inputs.
- reference_image1…9, reference_video1…3, reference_audio1…3 - reference mode. Cite them in the prompt as
Image 1,Video 1,Audio 1etc. Audio alone isn't enough to trigger reference mode. - No seed input - H3's fal and MUAPI APIs don't expose seed control. What you get is what you get.
Outputs: video (first clip), video_paths (all clips, feed into a saver), and status (provider, route, cost, per-job request IDs). H3 outputs native stereo audio on every generation - there's no audio toggle to flip, it's always there.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
export FAL_KEY=your_fal_key # or MUAPIAPP_API_KEY=...
(Or ComfyUI Manager → search comfyui-digit → install.) Restart ComfyUI, look under DIGIT.
Common issues
The README's troubleshooting table is short and accurate: FAL_KEY environment variable is not set → export it. MUAPI supports 2K only → set resolution to 2K. Reference-to-video requires at least one reference_image or reference_video → you connected only audio; it needs an image or video to reference. Duration must be between 4 and 15 → read the dropdown. A Refusing to download from untrusted URL host error means the provider returned an unexpected CDN host - that's a real bug to report, and the request ID is in status.
The trap to watch: this is a metered video API, and video is where cloud sessions get expensive fast. Eight 15-second clips at 2K is a real line item on fal's bill. Use the cost strip before you queue, not after.
Inputs (25)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | — | |
| provider | COMBO | fal | fal — Hosted MiniMax H3 with safety checker and prompt expansion. muapi — MiniMax H3 via unified async API; 2K output today. replicate — MiniMax H3 (available when Replicate publishes the model). replicate is hidden until MiniMax H3 is published on Replicate. |
| resolution | COMBO | 2K | MUAPI supports 2K only. fal supports 768P, 2K, and 4K. |
| aspect_ratio | COMBO | 16:9 | T2V requires a fixed ratio. I2V/FLF follow the source image. R2V supports adaptive. |
| duration | COMBO | 5 | Output length (4-15s). |
| batch_count | INT | 11–8 | Submits this many generations before polling. |
| enable_prompt_expansion | BOOLEAN | true | fal only. |
| enable_safety_checker | BOOLEAN | true | fal only. |
| first_frameopt | IMAGE | Image-to-video. Mutually exclusive with reference inputs. | |
| last_frameopt | IMAGE | Optional end frame for first-to-last interpolation. | |
| reference_image1opt | IMAGE | — | |
| reference_image2opt | IMAGE | — | |
| reference_image3opt | IMAGE | — | |
| reference_image4opt | IMAGE | — | |
| reference_image5opt | IMAGE | — | |
| reference_image6opt | IMAGE | — | |
| reference_image7opt | IMAGE | — | |
| reference_image8opt | IMAGE | — | |
| reference_image9opt | IMAGE | — | |
| reference_video1opt | VIDEO | — | |
| reference_video2opt | VIDEO | — | |
| reference_video3opt | VIDEO | — | |
| reference_audio1opt | AUDIO | — | |
| reference_audio2opt | AUDIO | — | |
| reference_audio3opt | AUDIO | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| video_paths | VIDEO_PATHS | — |
| status | STRING | — |