MiniMax H3 Asset
Casting a scene with a file, a frame number and one wire
- assets
- image
- video
- audio
- assets
- image
- audio
- report
MiniMax H3 is an omni-modal model: it takes text, pictures, video and audio as one context. The catch is that all that context has to get into the graph somehow, and you need to say not just what a file is but where in the video it belongs and what job it does there. This node is that statement. One picture, clip or sound, plus a role, a scene and a frame - chained as many times as your shot list needs.
How it works
Each asset node declares one piece of media and its place in the video, then hands its address book to the next node in the chain. Chain them, wire the last node's assets output into assets on MiniMax H3 Conditioning, and every segment's prompt can reference what you declared: pictures as <Picture N>, clips as <Video N>, sounds as <Audio N>, numbered in the order you chained them.
The Prompt Timeline window on the conditioning node is the friendlier face of the same data - you can drop files onto its tracks and it writes the asset chain for you. Every edit in that window lands on real nodes, so the graph runs identically with the window closed.
The inputs that matter
file- which picture, clip or sound to read, ascast/alice.png [input]ortakes/shot_03.mp4 [output]. The bracketed part is ComfyUI's folder; the node reads ComfyUI's own directories plus anything you add underpaths.allow_readin the pack'sconfig.yaml. The value(wired input)makes the node use the sockets instead, and a wired socket wins over a chosen file either way.role- the job it does.first frameopens the segment on it,last framecloses on it,keyframepins it atframe, andreference picture,reference clipandreference audioturn it into a numbered reference the prompt names. That last set is where H3 gets interesting: it's how you keep a character's face, a location, or a specific take of dialogue consistent across scenes.segment- which segment it belongs to,1upward.0means every segment, which is what you want for a recurring cast picture.frame- where a keyframe lands, counted in that segment's new frames at 24 fps.0is its first new frame,48is two seconds in, and-1is its last frame. Negative values count from the end, so-1is the one you use for "close on this picture".- Optional extras:
clip_startandclip_secondsto take a slice out of a clip or sound (0reads the first 362 frames of a clip, about fifteen seconds, and the whole of a sound), and theimage,videoandaudiosockets when the media comes from a node rather than a filename.
Four outputs: assets for the next node in the chain, plus image and audio so you can preview or reuse what it holds, and a report saying what was read and where it went.
Installing it
Part of WAS Node Suite v3 - ComfyUI Manager, search WAS Node Suite v3, or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+ and Python 3.10+, and the pack itself installs nothing.
What the node needs around it is the H3 model set: a transformer in models/diffusion_models (the community quants like minimax_h3_ref2va_pruned_w6a8.safetensors live under an H3 subfolder), minimax_h3_video_vae_int8_convrot.safetensors and minimax_h3_audio_vae_fp32.safetensors in models/vae, the text encoder in models/text_encoders, and - if you use a language model for the writing tools - a VLM via Load CLIP. Before you build on any of it: H3's weights are under a Community License whose Applicable Territory excludes the US, the EU, the UK and South Korea. Hosted Hailuo has no such problem; the local weights do.
Where it goes wrong
Everything renders as though the pictures weren't there. The asset chain has to be attached to the conditioning node's assets input - a chain that ends nowhere is just a nicely formatted list. Also check segment: an asset pointing at segment 3 references nothing in segment 1.
A prompt references <Picture 2> and the model invents something. Either the second asset in the chain is in a different segment, or its role isn't a reference role. Only reference picture gets a <Picture N> number.
Keyframes land where you didn't ask. frame counts that segment's new frames, not frames of the finished video. Placing at 48 in segment 4 is two seconds into segment 4. The one exception is the (the video being made) entry in the file list, where the frame is the finished video's frame - useful for a callback shot near the end.
A wired socket loses to a dropdown. That's by design: if both a file and a socket are set, the socket wins. Set file to (wired input) to make what's happening obvious when you come back to the workflow in a month.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| file | COMBO | Which picture, clip or sound to read, as `cast/alice.png [input]` or `takes/shot_03.mp4 [output]`. `(wired input)` reads the image, video and audio sockets instead, and a wired socket always wins over the file. | |
| role | COMBO | first frame | What this does for its segment. `first frame` = the segment opens on it; `last frame` = closes on it; `keyframe` = pinned at `frame`, as a still, a clip, a sound, or a clip with its sound; `reference picture`, `reference clip` and `reference audio` = named `<Picture N>`, `<Video N>` and `<Audio N>` in that segment's prompt. |
| segment | INT | 10–24 | Which segment it belongs to, as `1` for the opening segment or `3` for Segment 3. `0` = every segment, as a cast picture the whole run references. |
| frame | INT | 0-3600–3600 | Where a `keyframe` lands, counted in the segment's new frames at 24 fps: `0` = its first new frame, `48` = two seconds in, `-1` = its last frame. With `(the video being made)`, the frame of the finished video to reference, as `120`. |
| assetsopt | WAS_H3_ASSETS | The assets before this one, from another MiniMax H3 Asset's assets output. This asset joins the end of that chain, so a run of any length is one wire into the next node. | |
| clip_startopt | FLOAT | 0.00–36000 | Seconds into the file a clip or sound starts, as `0` or `12.5`. |
| clip_secondsopt | FLOAT | 0.00–600 | Seconds of the file to read, as `5`, or `0` for the first 362 frames of a clip, about 15 seconds, and the whole of a sound. A reference clip is cut to the longest segment either way. |
| imageopt | IMAGE | A picture, or a batch of frames at 24 fps for a clip, in place of a file. | |
| videoopt | VIDEO | A video in place of a file: its frames, brought to 24 fps, and its sound. | |
| audioopt | AUDIO | A sound in place of a file, or the soundtrack of the wired image frames. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| assets | WAS_H3_ASSETS | The chain with this asset at its end, for the next MiniMax H3 Asset or for assets on MiniMax H3 Conditioning. |
| image | IMAGE | The picture or frames it holds, for a preview or another node. |
| audio | AUDIO | The sound it holds. |
| report | STRING | What was read and where it goes. |