ComfyUI Node
arkennemasis fal · MiniMax H3 Reference to Video · $0.05-0.16/s
MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio. Pricing (from fal, 2026-09-30): Video costs $0.05 per second at 480p, $0.06 per second at 768p, $0.13 per second at 2K and $0.16 per second at 4K; the first 5 reference images are free and each additional image costs $0.08. Model page: https://fal.ai/models/minimax/h3/reference-to-video Key: FAL_KEY in the .env next to run_nvidia_gpu.bat. Results are saved in output/fal/minimax_h3_reference-to-video/.
arkennemasis fal · MiniMax H3 Reference to Video · $0.05-0.16/s
- reference_image_1
- reference_image_2
- reference_image_3
- reference_image_4
- reference_image_5
- reference_image_6
- reference_video_1
- reference_video_2
- reference_video_3
- reference_audio_1
- reference_audio_2
- reference_audio_3
- video
- video_path
- expanded_prompt
- info
◄prompt►
◄duration5►
◄resolution2K►
◄seed-1►
◄enable_safety_checkertrue►
◄prompt_expansion_modebalanced►
◄aspect_ratioadaptive►
◄max_cost_usd20.0►
◄reuse_identical_runtrue►
◄extra_json►
◄max_concurrent1►
Categoryarkennemasis/fal/Video/MiniMax
Inputs (23)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | Text prompt for video generation. Refer to reference assets by their modality and order in the reference lists: Image 1, Image 2, Video 1, Audio 1, and so on. | |
| duration | INT | 55–15 | The duration of the video in seconds. |
| resolution | COMBO | 2K | The resolution of the generated video. 480P and 768P are native generation modes; 2K and 4K upscale a 768P base result. |
| seed | INT | -1-1–2147483647 | Random seed. A random seed is selected when omitted. -1 leaves it to the model. |
| enable_safety_checker | BOOLEAN | true | If set to true, the safety checker will be enabled. |
| prompt_expansion_mode | STRING | balanced | How much effort to spend rewriting the prompt before generation. 'disabled' skips prompt expansion. 'fast' returns in about a second. 'balanced' picks per request. 'quality' spends up to ~30s on a richer prompt. Empty leaves it out. |
| aspect_ratio | COMBO | adaptive | The aspect ratio of the generated video. |
| reference_image_1opt | IMAGE | URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_image_2opt | IMAGE | URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_image_3opt | IMAGE | URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_image_4opt | IMAGE | URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_image_5opt | IMAGE | URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_image_6opt | IMAGE | URLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_video_1opt | VIDEO | URLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_video_2opt | VIDEO | URLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_video_3opt | VIDEO | URLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_audio_1opt | AUDIO | URLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_audio_2opt | AUDIO | URLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| reference_audio_3opt | AUDIO | URLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1). | |
| max_cost_usdopt | FLOAT | 20.00–100000 | Safety cap. If the estimated cost of this run is above this many US dollars the node stops BEFORE sending anything to fal. 0 = no cap. |
| reuse_identical_runopt | BOOLEAN | true | If these exact inputs (the settings AND the same pictures/clips/audio) were already paid for and the files are still in output/fal, return them instead of paying again - also after a ComfyUI restart. Turn off for a fresh take with the same settings. |
| extra_jsonopt | STRING | Advanced: a JSON object merged into the request last, for any fal field this node has no box for. Example: {"seed": 7}. Leave empty normally. | |
| max_concurrentopt | INT | 11–32 | How many paid fal calls may run at the same time in one run, across every fal node. 1 (default) = one after another. This node starts its paid part only while fewer than this many fal calls are running. Once any fal node fails, the ones still waiting are not started, so nothing more is billed. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| video_path | STRING | — |
| expanded_prompt | STRING | — |
| info | STRING | JSON: request id, cost estimate, saved files and fal's full answer. |