Nodes/GrsAI api in ComfyUI/🎬 GrsAI MiniMax H3 - 768p
ComfyUI Node

🎬 GrsAI MiniMax H3 - 768p

The one you'd actually leave in your workflow

By 31702160136·Created about a year ago·Updated a day ago· 130
🎬 GrsAI MiniMax H3 - 768p
  • image_1
  • image_2
  • image_3
  • image_4
  • image_5
  • image_6
  • image_7
  • image_8
  • image_9
  • audio_1
  • audio_2
  • audio_3
  • video
  • status
  • api_task_ids
◄prompt电影级写实风格,镜头运动稳定流畅,主体外观前后一致,环境音与画面自然同步。►
◄apikey请输入您的APIKEY: sk-xxxxxxx►
◄modelminimax-h3►
◄aspect_ratiolandscape►
◄duration10►
◄seed0►

The middle tier, and why it's the default answer

Three nodes in this pack call MiniMax H3; they differ by exactly one string. Grsai_MiniMaxH3_768p sends resolution: "768p", and 768p is the tier the model was actually pitched around - MiniMax's own line about H3 was price-performance at 2K, with 768p as the practical everyday setting. 480p is for auditions. 1080p costs more and caps you at ten seconds. This one keeps the full 15-second ceiling and enough pixels that you can hand the clip to an upscaler or a video model afterwards and not be embarrassed.

It's worth being clear about what you're buying here, because "MiniMax H3" means two different things in 2026. The open weights exist, they're 42.5GB, and their community licence excludes the US, EU, UK and South Korea - if you're in one of those places, self-hosting is not a licensed option and the hosted API is. This node is the hosted path, through GrsAI, a reseller with a per-call price and no community reputation whatsoever (a corpus search for the name comes back empty). Most people met H3 through exactly this route: on launch day it was in ComfyUI via API access before local weights existed anywhere. That's the trade - you skip the download and the geofence question, and in exchange your prompt, your reference images and your reference audio go to a third party.

How it works

Submit asynchronously, poll, download, wrap. The client POSTs the job to /v1/api/generate with model: "minimax-h3", resolution: "768p", your duration, aspectRatio, seed, and replyType: "async", prints the task ID, then hits /v1/api/result every two seconds until the job reads succeeded - up to an hour before it gives up. Result URLs get streamed down over a long-lived HTTP connection and handed to VideoFromFile, which is what turns a downloaded MP4 into a native ComfyUI VIDEO that downstream nodes decode lazily rather than loading every frame into VRAM.

Before any of that happens the node validates locally: 480p/768p/1080p only, duration 1–15, at most nine reference images and three reference audio clips. Reference images are flattened to their first frame and base64'd as PNG; audio is converted to 16-bit PCM WAV. That encoding happens on your machine, so a fat reference set costs you upload time before the API has done a thing.

Inputs worth setting

  • prompt - the default is a real prompt, in Chinese: cinematic realism, stable camera movement, consistent subject appearance, ambient audio synced to the picture. It's a good template even if you write in English.
  • apikey - the GrsAI key, sk- prefixed. This widget is where the key comes from for this node; the .env file the README talks about is read by the pack's older Flux nodes only.
  • aspect_ratio - landscape (default) or portrait. No ratio numbers, no resolution picking beyond the node you chose.
  • duration - 1 to 15 seconds, default 10. The model is documented as a 4–15s generator; if a very short clip comes back strange, that's the likely reason.
  • seed - 0 to 9999999999 with control_after_generate. Pin it while you iterate on wording, then let it roll when you're hunting for a take.
  • image_1 … image_9 - optional IMAGE inputs, first frame only, up to nine. In practice one or two is the workflow: an image to hold the look, sometimes a second for a position or pose cue.
  • audio_1 … audio_3 - up to three AUDIO references, sent as WAV. H3's selling point is that it treats audio as part of the same context as the picture, so this is your lever on the audio side; the pack documents it only as "reference audio", so experiment.

Outputs

video is the VIDEO socket - into core Save Video for an MP4, or a preview node to eyeball it. Not an IMAGE batch: image-side nodes won't accept it, which is the most common wiring mistake people make. status reports orientation, resolution, duration and how many references actually went up. api_task_ids carries the server task ID per run; hold onto it if a job fails after being accepted, because a submitted-then-failed task is a billing conversation.

Install

Manager → Install via Git URL → https://github.com/31702160136/ComfyUI-GrsAI.git, or:

cd ComfyUI/custom_nodes
git clone https://github.com/31702160136/ComfyUI-GrsAI.git
pip install requests python-dotenv httpx httpcore

Install those four by name instead of running requirements.txt, which also lists torch, Pillow, numpy and fal-client. ComfyUI already has the first three, and fal-client is dead weight - no module in the pack imports it. The README's portable-Windows command passes --force-reinstall, which on a list containing torch can wreck a perfectly good environment. The native VIDEO socket also needs a reasonably current ComfyUI; older builds trigger the node's own "update ComfyUI" error.

CategoryGrsAI/Video

Inputs (18)

NameTypeDefaultDescription
promptSTRING电影级写实风格,镜头运动稳定流畅,主体外观前后一致,环境音与画面自然同步。—
apikeySTRING请输入您的APIKEY: sk-xxxxxxx—
modelCOMBOminimax-h31 options: minimax-h3
aspect_ratioCOMBOlandscape2 options: portrait, landscape
durationINT101–15—
seedINT00–9999999999—
image_1optIMAGE—
image_2optIMAGE—
image_3optIMAGE—
image_4optIMAGE—
image_5optIMAGE—
image_6optIMAGE—
image_7optIMAGE—
image_8optIMAGE—
image_9optIMAGE—
audio_1optAUDIO—
audio_2optAUDIO—
audio_3optAUDIO—

Outputs (3)

NameTypeDescription
videoVIDEO—
statusSTRING—
api_task_idsSTRING—