AMD MiniMax H3 First-Last-Frame to Video
MiniMax H3 First-Last-Frame to Video, Explained
- first_frame
- last_frame
- VIDEO
Most image-to-video tools give you a first frame and a hope. This one takes two images and animates the journey between them, ending on the exact frame you specified - the difference between "a clip that drifts away from my composition" and "a shot that lands where the next shot starts". If you're cutting a sequence together, or doing any kind of morph or reveal, that endpoint control is the whole feature.
The catch: you're not running H3 locally. This node calls MiniMax H3 - Hailuo's model, open-weighted in August 2026 but at ~42.5GB with a licence that geofences the US, EU, UK and South Korea out of running the weights at all - through an AMD gateway at h3.oneclickamd.ai. No model file, no VRAM, no GPU. Just a key.
How it works
Your frames don't get uploaded to a storage bucket. The gateway has no upload endpoint, so the node base64-encodes each image as a PNG data URI and inlines it in the request body, tagged with a role: your first image as first_frame, the optional second as last_frame. Alongside that it sends at least one text content item, because the gateway requires exactly one non-empty prompt even when the images carry all the visual information. That last bit surprises people.
The job goes to /v2/video_generation with model: "MiniMax-H3" plus resolution and duration, then the node polls a query endpoint every two seconds and downloads the finished clip as a VIDEO output. It's built on ComfyUI's own comfy_api_nodes Hailuo request models and helpers rather than a bespoke HTTP client - hence zero dependencies and a ComfyUI version floor.
Inputs and outputs
first_frame(IMAGE, required) - the starting frame. Anything that produces an IMAGE tensor works, so you can generate it locally and animate it without ever saving a file.prompt(multiline, required) - describes the motion. Not optional, even though the images tell the model a lot.last_frame(IMAGE, optional) - wire it up and the clip is pinned at both ends. Leave it empty and this is plain image-to-video: one still, set in motion.duration- integer slider, 4 to 15 seconds, default 5.resolution-768Ponly. The model goes higher; this gateway doesn't.
There's no ratio widget, and the tooltip explains why: the video keeps the aspect ratio of your images. Convenient - and the reason to keep your two frames the same shape and roughly the same framing. Feeding it a 16:9 opening and a square ending is asking for a letterboxed drift in the middle.
Output is a single VIDEO. The shipped example runs it into core SaveVideo ("video/AMD_MiniMax_H3_flf"), which is the simplest thing that works.
Installing it
ComfyUI Manager, search ComfyUI-AMD-MiniMaxH3. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/wangxunx/ComfyUI-AMD-MiniMaxH3.git
Restart, grab a key at https://h3.oneclickamd.ai, and set it once per machine - it applies to every workflow after that:
export H3_GATEWAY_API_KEY="your-key-here"
# or drop it in ComfyUI/user/h3_gateway_api_key.txt
H3_GATEWAY_BASE_URL if you're running your own gateway. The key is read from the environment first, then that file - never from a widget, on purpose, since widget values get baked into workflow JSON and into exported video metadata.
Where people get burned
The empty prompt error. Instinct says "the images are the prompt, I'll leave the text blank". The gateway refuses that. Even a short line about how the frames connect does the job, and it genuinely helps - the model uses it to choose the motion, and camera moves in particular need describing.
Old ComfyUI. The pack needs >=0.30.0, because it imports the Hailuo03 request models that only exist from that release. On an older build the nodes just don't load. Update before you debug anything else.
Frames that don't belong together. Two images of the same subject at different angles work; two unrelated pictures produce a clip where something turns into something else. That can be the effect you want - it's a respectable morph tool - but know that's what you asked for.
Big source images. Each frame is inlined as base64 PNG, so a 4K start frame becomes a multi-megabyte payload over a consumer uplink. Downscale first; the result is 768P, so you lose nothing.
Assuming free means unlimited. The pack metadata advertises free H3 access through the AMD gateway, with no documented quota. Great for experimenting, bad to hard-code into an overnight batch - hosted video inference is metered everywhere.
Skim it first. One author, one commit, no community footprint, and it holds a credential and calls the network by design - the category that already shipped a credential-stealing node once. The key handling in nodes.py is careful and the file is short. Read it, then paste your key.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| first_frame | IMAGE | First frame of the output video. | |
| prompt | STRING | Text prompt describing how the images should animate. | |
| resolution | COMBO | 768P | Resolution of the output video. The gateway currently renders one size, 768 pixels on the short edge. |
| duration | INT | 54–15 | Duration of the output video in seconds (4-15). |
| last_frameopt | IMAGE | Optional last frame of the output video. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| VIDEO | VIDEO | — |