🎬 GrsAI MiniMax H3 - 1080p
1080p output, but the duration slider stops at ten
- image_1
- image_2
- image_3
- image_4
- image_5
- image_6
- image_7
- image_8
- image_9
- audio_1
- audio_2
- audio_3
- video
- status
- api_task_ids
The finished-output tier, with a shorter leash
Grsai_MiniMaxH3_1080p is the same node as its 480p and 768p siblings with one string changed - resolution: "1080p" - and one real difference you'll notice the moment you drag the slider: duration caps at 10 seconds instead of 15. That's not a UI quirk, it's enforced twice, in the node's own input range and again in the API client, and the server won't take a longer 1080p job. Everything else about the workflow is identical, so pick this tier when the clip is the deliverable rather than a test.
It's also the tier where the local-versus-cloud question gets interesting. H3's headline is 2K output, and the open weights are ~42.5GB under a licence that excludes the US, EU, UK and South Korea - so if you're outside that fence and have the VRAM, running it yourself is both cheaper per clip and fully under your control. This node is the other path: MiniMax H3 through the GrsAI reseller API, no download, no geofence question, works anywhere you can pay. What you give up is privacy and predictability of price - your prompt and every reference image or audio clip leaves your machine, and the reseller has no community footprint to check (the corpus contains zero mentions of the name). On launch day a lot of people met H3 exactly this way, through API access, because local weights didn't exist yet.
How it works
One async round trip. The client posts to /v1/api/generate with model: "minimax-h3", resolution: "1080p", duration, aspectRatio and replyType: "async", gets a task ID back, then polls /v1/api/result every two seconds - up to an hour - until the job reports succeeded. The result MP4 is streamed down and wrapped with VideoFromFile into ComfyUI's native VIDEO type, which stays a file handle until something downstream actually decodes it.
The reason the API feels like it's taking a while is upstream of your machine: H3 generates its audio in the same pass as the picture - dialogue, room tone and effects are part of one unified context rather than a second model stacked on top. That's the thing that made it land well on release, and it's why a 1080p ten-second job isn't a fast call. Reference media is encoded locally first (first frame of each image as base64 PNG, audio as 16-bit PCM WAV), so on a big reference set you're also paying upload cost before generation even starts.
Inputs
- prompt - multiline; the shipped default is a legit Chinese prompt about cinematic realism, steady camera work, subject consistency and synced ambient audio. At 1080p it's worth being specific about camera and lighting, since detail is the point of paying for this tier.
- apikey - the GrsAI key,
sk-prefixed, entered in the widget. This node reads the key from the field, not from the pack's.env(that's the legacy Flux nodes' route). - model -
minimax-h3, single option. - aspect_ratio -
landscapeorportrait, defaulting to landscape. - duration - 1 to 10 seconds, default 10. If you need 12 or 15, drop to the 768p node.
- seed - 0 to 9999999999, with
control_after_generateso you can lock a good take while you tinker. - image_1 … image_9 - up to nine IMAGE references, first frame only. Two is a normal load: one for look, one for framing.
- audio_1 … audio_3 - up to three AUDIO references sent as WAV. Given H3's unified-context design this is how you push the audio side; the pack documents the input as "reference audio" and nothing more, so it's a test-and-see.
Outputs
video goes to a video-capable core node - Save Video for an MP4 on disk, or a preview node to watch it. It will not plug into an IMAGE input; that socket mismatch is the standard beginner snag. status gives model, orientation, resolution, duration and reference counts; api_task_ids gives you the server task ID, which is what you'll need if a job is accepted and then fails.
Install
Manager → Install via Git URL → https://github.com/31702160136/ComfyUI-GrsAI.git, or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/31702160136/ComfyUI-GrsAI.git
pip install requests python-dotenv httpx httpcore
Note that the README is behind the code: it documents the image nodes and Nano Banana in detail and never mentions MiniMax H3, and its changelog stops at v1.0.8 while the pack ships v1.1.4. The dropdowns are your documentation. Also install the four real dependencies by name rather than pip install -r requirements.txt - that file additionally lists torch, numpy, Pillow and fal-client, and nothing imports fal-client. Pairing the README's portable-Windows --force-reinstall line with a requirements file containing torch is a genuinely bad idea. Finally, the node needs a ComfyUI new enough to have the native VIDEO type; otherwise it stops with its own "please update ComfyUI" error.
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | 电影级写实风格,镜头运动稳定流畅,主体外观前后一致,环境音与画面自然同步。 | — |
| apikey | STRING | 请输入您的APIKEY: sk-xxxxxxx | — |
| model | COMBO | minimax-h3 | 1 options: minimax-h3 |
| aspect_ratio | COMBO | landscape | 2 options: portrait, landscape |
| duration | INT | 101–10 | — |
| seed | INT | 00–9999999999 | — |
| image_1opt | IMAGE | — | |
| image_2opt | IMAGE | — | |
| image_3opt | IMAGE | — | |
| image_4opt | IMAGE | — | |
| image_5opt | IMAGE | — | |
| image_6opt | IMAGE | — | |
| image_7opt | IMAGE | — | |
| image_8opt | IMAGE | — | |
| image_9opt | IMAGE | — | |
| audio_1opt | AUDIO | — | |
| audio_2opt | AUDIO | — | |
| audio_3opt | AUDIO | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| status | STRING | — |
| api_task_ids | STRING | — |