Nodes/comfyui-arkennemasis/arkennemasis fal · MiniMax H3 Max 3D to Video · $0.05-0.16/s
ComfyUI Node

arkennemasis fal · MiniMax H3 Max 3D to Video · $0.05-0.16/s

Transform Blender renders and 3D previs into photorealistic video. H3 Max uses the source clip to guide scene layout, camera movement, and timing, with optional image references for appearance. Pricing (from fal, 2026-09-29): For every second of video you generate you will be charged $0.05 at 480p, $0.08 at 768p or $0.16 at 1080p (minimum 5 seconds), plus reference token usage for your input video and images, the same as H3 Max Reference to Video. If you don't provide reference images, up to max_generated_reference_images (default 2) are generated automatically at about $0.02–$0.04 each. For example, a 5 second 16:9 video at 768p with your own reference images costs about $0.96. Model page: https://fal.ai/models/minimax/h3-max/3d-to-video Key: FAL_KEY in the .env next to run_nvidia_gpu.bat. Results are saved in output/fal/minimax_h3-max_3d-to-video/.

By Hishamahmer·Created 2 months ago·Updated 4 days ago· 10
arkennemasis fal · MiniMax H3 Max 3D to Video · $0.05-0.16/s
  • video
  • reference_image_1
  • reference_image_2
  • reference_image_3
  • reference_image_4
  • reference_image_5
  • reference_image_6
  • video
  • video_path
  • info
◄prompt►
◄max_generated_reference_images2►
◄resolution768P►
◄max_cost_usd20.0►
◄reuse_identical_runtrue►
◄extra_json►
◄max_concurrent1►
Categoryarkennemasis/fal/Video/MiniMax

Inputs (14)

NameTypeDefaultDescription
videoVIDEOBlender video, up to 15 seconds and 32 shots, at a public HTTPS URL.
promptSTRINGOptional clarification of what your proxies represent or how they move, for example: 'The moving block represents a running person.' Camera, trajectories, timing and object count remain defined by the video. Empty leaves it out.
max_generated_reference_imagesINT21–8Maximum NEW images to generate when no reference images are supplied. Ignored when reference images are supplied. Uses fewer when enough views are covered. Uncovered shots stay in the video.
resolutionCOMBO768POutput quality. Duration is taken automatically from the video.
reference_image_1optIMAGEOptional environment, subject, interior or detail references. When supplied, uses these images directly without visual planning or generating new images; Scene Intent is optional. Without references, automatically creates appearance references within your new-image limit. Camera and movement always come from the video. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is numb
reference_image_2optIMAGEOptional environment, subject, interior or detail references. When supplied, uses these images directly without visual planning or generating new images; Scene Intent is optional. Without references, automatically creates appearance references within your new-image limit. Camera and movement always come from the video. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is numb
reference_image_3optIMAGEOptional environment, subject, interior or detail references. When supplied, uses these images directly without visual planning or generating new images; Scene Intent is optional. Without references, automatically creates appearance references within your new-image limit. Camera and movement always come from the video. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is numb
reference_image_4optIMAGEOptional environment, subject, interior or detail references. When supplied, uses these images directly without visual planning or generating new images; Scene Intent is optional. Without references, automatically creates appearance references within your new-image limit. Camera and movement always come from the video. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is numb
reference_image_5optIMAGEOptional environment, subject, interior or detail references. When supplied, uses these images directly without visual planning or generating new images; Scene Intent is optional. Without references, automatically creates appearance references within your new-image limit. Camera and movement always come from the video. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is numb
reference_image_6optIMAGEOptional environment, subject, interior or detail references. When supplied, uses these images directly without visual planning or generating new images; Scene Intent is optional. Without references, automatically creates appearance references within your new-image limit. Camera and movement always come from the video. Each socket may carry a batch; they are sent in socket order (the first picture of socket 1 is numb
max_cost_usdoptFLOAT20.00–100000Safety cap. If the estimated cost of this run is above this many US dollars the node stops BEFORE sending anything to fal. 0 = no cap.
reuse_identical_runoptBOOLEANtrueIf these exact inputs (the settings AND the same pictures/clips/audio) were already paid for and the files are still in output/fal, return them instead of paying again - also after a ComfyUI restart. Turn off for a fresh take with the same settings.
extra_jsonoptSTRINGAdvanced: a JSON object merged into the request last, for any fal field this node has no box for. Example: {"seed": 7}. Leave empty normally.
max_concurrentoptINT11–32How many paid fal calls may run at the same time in one run, across every fal node. 1 (default) = one after another. This node starts its paid part only while fewer than this many fal calls are running. Once any fal node fails, the ones still waiting are not started, so nothing more is billed.

Outputs (3)

NameTypeDescription
videoVIDEO—
video_pathSTRING—
infoSTRINGJSON: request id, cost estimate, saved files and fal's full answer.