Nodes/comfyui-arkennemasis/arkennemasis fal · MiniMax H3 Max Reference to Video · $0.05-0.16/s
ComfyUI Node

arkennemasis fal · MiniMax H3 Max Reference to Video · $0.05-0.16/s

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality Pricing (from fal, 2026-09-30): Billing uses the requested output duration at $0.05 per second for 480p, $0.08 for 768p, and $0.16 for 1080p. Each request includes 4,096 reference tokens, shared across all reference images, videos, and audio clips; additional usage costs $0.02 per 1,000 tokens, prorated. Square reference images, including 1024 × 1024 and 2048 × 2048 uploads, contribute 1,024 tokens each. With no video or audio references, the first four square images add no reference charge, and each additional square image adds $0.02048 (about $0.02). For example, a 5-second 768p output costs $0.40 with up to four square reference images, or $0.42048 with five. See the pricing details below for other image shapes, video references, and audio references. Model page: https://fal.ai/models/minimax/h3-max/reference-to-video Key: FAL_KEY in the .env next to run_nvidia_gpu.bat. Results are saved in output/fal/minimax_h3-max_reference-to-video/.

By Hishamahmer·Created 2 months ago·Updated 4 days ago· 10
arkennemasis fal · MiniMax H3 Max Reference to Video · $0.05-0.16/s
  • reference_image_1
  • reference_image_2
  • reference_image_3
  • reference_image_4
  • reference_image_5
  • reference_image_6
  • reference_video_1
  • reference_video_2
  • reference_video_3
  • reference_audio_1
  • reference_audio_2
  • reference_audio_3
  • video
  • video_path
  • timings
  • expanded_prompt
  • seed
  • info
◄prompt►
◄duration5►
◄resolution768P►
◄seed-1►
◄enable_safety_checkertrue►
◄prompt_expansion_modebalanced►
◄aspect_ratioadaptive►
◄max_cost_usd20.0►
◄reuse_identical_runtrue►
◄extra_json►
◄max_concurrent1►
Categoryarkennemasis/fal/Video/MiniMax

Inputs (23)

NameTypeDefaultDescription
promptSTRINGText prompt for video generation. Refer to reference assets by their modality and order in the reference lists: Image 1, Image 2, Video 1, Audio 1, and so on.
durationINT55–15The duration of the video in seconds.
resolutionCOMBO768PThe native generation resolution, or 1080P latent refinement from a native 768P source.
seedINT-1-1–2147483647Random seed. A random seed is selected when omitted. -1 leaves it to the model.
enable_safety_checkerBOOLEANtrueIf set to true, the safety checker will be enabled.
prompt_expansion_modeSTRINGbalancedHow much effort to spend rewriting the prompt before generation. 'disabled' skips prompt expansion. 'balanced' returns in about a second. 'quality' spends up to ~30s on a richer prompt.
aspect_ratioCOMBOadaptiveThe aspect ratio of the generated video.
reference_image_1optIMAGEURLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_image_2optIMAGEURLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_image_3optIMAGEURLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_image_4optIMAGEURLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_image_5optIMAGEURLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_image_6optIMAGEURLs of subject/style reference images, referenced in the prompt as Image 1, Image 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_video_1optVIDEOURLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_video_2optVIDEOURLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_video_3optVIDEOURLs of motion/reference video clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Video 1, Video 2, and so on. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_audio_1optAUDIOURLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_audio_2optAUDIOURLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
reference_audio_3optAUDIOURLs of reference audio clips (2-15 seconds each, combined duration at most 15 seconds), referenced in the prompt as Audio 1, Audio 2, and so on. Images, videos, and audio can be provided individually or together. Reference images, videos, and audio clips must add up to at most 12 files. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is number 1).
max_cost_usdoptFLOAT20.00–100000Safety cap. If the estimated cost of this run is above this many US dollars the node stops BEFORE sending anything to fal. 0 = no cap.
reuse_identical_runoptBOOLEANtrueIf these exact inputs (the settings AND the same pictures/clips/audio) were already paid for and the files are still in output/fal, return them instead of paying again - also after a ComfyUI restart. Turn off for a fresh take with the same settings.
extra_jsonoptSTRINGAdvanced: a JSON object merged into the request last, for any fal field this node has no box for. Example: {"seed": 7}. Leave empty normally.
max_concurrentoptINT11–32How many paid fal calls may run at the same time in one run, across every fal node. 1 (default) = one after another. This node starts its paid part only while fewer than this many fal calls are running. Once any fal node fails, the ones still waiting are not started, so nothing more is billed.

Outputs (6)

NameTypeDescription
videoVIDEO—
video_pathSTRING—
timingsSTRING—
expanded_promptSTRING—
seedINT—
infoSTRINGJSON: request id, cost estimate, saved files and fal's full answer.