🎬 GrsAI MiniMax H3 - 480p
The cheap tier for finding shots before you pay for pixels
- image_1
- image_2
- image_3
- image_4
- image_5
- image_6
- image_7
- image_8
- image_9
- audio_1
- audio_2
- audio_3
- video
- status
- api_task_ids
What it is, and the honest case for 480p
MiniMax H3 is the 33B omni-modal video model MiniMax opened up in August 2026 - the one the community called "very, VERY good" the day it landed, and the one that generates its audio jointly with the picture instead of bolting a soundtrack on afterwards. Running it locally means ~42.5GB of weights and a licence that excludes the US, EU, UK and South Korea from where you're allowed to use it. So for a lot of people the hosted API isn't the lazy option, it's the only licensed option, and that's the gap this node fills: Grsai_MiniMaxH3_480p is a thin ComfyUI wrapper that posts your prompt to the GrsAI reseller API and hands back a native VIDEO.
480p is the bottom rung, and it's the one you'll use more than you expect. Video prompting is a guessing game - you don't know if the camera move reads until you watch it. At 480p a reference image plus a camera direction is a cheap way to audition ten framings, pick the one that works, and re-run just that one at 768p or 1080p. Nobody hangs 480p on a wall. Everybody uses it to find the shot.
Caveat worth stating plainly: GrsAI is a reseller with no community footprint at all - the corpus has zero threads mentioning it, and the pack's README doesn't even document the MiniMax nodes yet. You're trusting an unvetted vendor handle with your key, and your prompt and reference media leave your machine.
How it works
One submit, then polling. The node encodes your prompt and any reference media, POSTs to /v1/api/generate with model: "minimax-h3", resolution: "480p", duration, aspectRatio and replyType: "async", and gets a task ID back immediately. It then polls /v1/api/result every two seconds for up to an hour, streams the finished MP4 down, and wraps it in ComfyUI's native VIDEO type via VideoFromFile. The video isn't a folder of frames - it's a file handle ComfyUI decodes lazily.
The client validates before it spends anything: resolution must be 480p/768p/1080p, duration 1–15s, at most nine reference images and three reference audio clips. References are re-encoded on your side - images become base64 PNG, audio becomes 16-bit PCM WAV - which is why a heavy reference set makes the request feel slow before the API has done anything.
Inputs you actually touch
- prompt - multiline, and the shipped default is a decent Chinese starting prompt (cinematic realism, stable camera movement, consistent subject, audio synced to picture). Try it before you replace it.
- apikey - your GrsAI key,
sk-prefixed. The widget is the only place this node reads the key from; the pack's.envapplies to its legacy Flux nodes, not this one. - model -
minimax-h3, the only option. Present so the workflow file is self-documenting. - aspect_ratio -
portraitorlandscape, defaults to landscape. That's the whole choice: no numbers, the model picks the framing. 480p landscape is the cheapskate's wide shot. - duration - 1 to 15 seconds, default 10. H3 is documented as a 4–15s generator, so 1–3s requests are the ones most likely to behave oddly; the node will still let you ask.
- seed - 0 to 9999999999, and it has the standard
control_after_generatewidget. Fix it when you're iterating on prompt wording so you're comparing words, not noise. - image_1 … image_9 - optional IMAGE inputs. Only the first frame of each is sent, so these are keyframe references, not video inputs.
- audio_1 … audio_3 - up to three AUDIO inputs, sent as WAV references. Since H3's pitch is unified text/image/audio context, this is how you steer the audio side; the pack itself doesn't document it beyond "reference audio", so test rather than assume.
Outputs
video is the native VIDEO socket - wire it to ComfyUI's core Save Video node to get an MP4 on disk, or a video preview node to just watch it. A node expecting an IMAGE batch won't take this socket, which trips people up when they try to reuse their image workflow. status gives you model, orientation, resolution, duration and the reference counts in one line. api_task_ids holds the server task ID, which is the thing you quote if you need to argue about a failed or charged-but-empty job.
Install
Manager → Install via Git URL → https://github.com/31702160136/ComfyUI-GrsAI.git, or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/31702160136/ComfyUI-GrsAI.git
pip install requests python-dotenv httpx httpcore
Skip requirements.txt: it drags in torch, fal-client, Pillow and numpy, and nothing in the node code uses fal-client. The README's portable-Windows line adds --force-reinstall, which is a great way to re-resolve torch underneath a working ComfyUI. You also need a ComfyUI new enough to expose the native VIDEO type - the node errors out with "current ComfyUI doesn't support native VIDEO output" rather than pretending, which is at least a clear failure. Update ComfyUI if you see it.
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | 电影级写实风格,镜头运动稳定流畅,主体外观前后一致,环境音与画面自然同步。 | — |
| apikey | STRING | 请输入您的APIKEY: sk-xxxxxxx | — |
| model | COMBO | minimax-h3 | 1 options: minimax-h3 |
| aspect_ratio | COMBO | landscape | 2 options: portrait, landscape |
| duration | INT | 101–15 | — |
| seed | INT | 00–9999999999 | — |
| image_1opt | IMAGE | — | |
| image_2opt | IMAGE | — | |
| image_3opt | IMAGE | — | |
| image_4opt | IMAGE | — | |
| image_5opt | IMAGE | — | |
| image_6opt | IMAGE | — | |
| image_7opt | IMAGE | — | |
| image_8opt | IMAGE | — | |
| image_9opt | IMAGE | — | |
| audio_1opt | AUDIO | — | |
| audio_2opt | AUDIO | — | |
| audio_3opt | AUDIO | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| status | STRING | — |
| api_task_ids | STRING | — |