Nodes/was-node-suite-comfyui/MiniMax H3 Asset
ComfyUI Node Runs on cloud

MiniMax H3 Asset

Casting a scene with a file, a frame number and one wire

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
MiniMax H3 Asset
  • assets
  • image
  • video
  • audio
  • assets
  • image
  • audio
  • report
◄file▾►
◄rolefirst frame►
◄segment1►
◄frame0►
◄clip_start0.0►
◄clip_seconds0.0►

MiniMax H3 is an omni-modal model: it takes text, pictures, video and audio as one context. The catch is that all that context has to get into the graph somehow, and you need to say not just what a file is but where in the video it belongs and what job it does there. This node is that statement. One picture, clip or sound, plus a role, a scene and a frame - chained as many times as your shot list needs.

How it works

Each asset node declares one piece of media and its place in the video, then hands its address book to the next node in the chain. Chain them, wire the last node's assets output into assets on MiniMax H3 Conditioning, and every segment's prompt can reference what you declared: pictures as <Picture N>, clips as <Video N>, sounds as <Audio N>, numbered in the order you chained them.

The Prompt Timeline window on the conditioning node is the friendlier face of the same data - you can drop files onto its tracks and it writes the asset chain for you. Every edit in that window lands on real nodes, so the graph runs identically with the window closed.

The inputs that matter

  • file - which picture, clip or sound to read, as cast/alice.png [input] or takes/shot_03.mp4 [output]. The bracketed part is ComfyUI's folder; the node reads ComfyUI's own directories plus anything you add under paths.allow_read in the pack's config.yaml. The value (wired input) makes the node use the sockets instead, and a wired socket wins over a chosen file either way.
  • role - the job it does. first frame opens the segment on it, last frame closes on it, keyframe pins it at frame, and reference picture, reference clip and reference audio turn it into a numbered reference the prompt names. That last set is where H3 gets interesting: it's how you keep a character's face, a location, or a specific take of dialogue consistent across scenes.
  • segment - which segment it belongs to, 1 upward. 0 means every segment, which is what you want for a recurring cast picture.
  • frame - where a keyframe lands, counted in that segment's new frames at 24 fps. 0 is its first new frame, 48 is two seconds in, and -1 is its last frame. Negative values count from the end, so -1 is the one you use for "close on this picture".
  • Optional extras: clip_start and clip_seconds to take a slice out of a clip or sound (0 reads the first 362 frames of a clip, about fifteen seconds, and the whole of a sound), and the image, video and audio sockets when the media comes from a node rather than a filename.

Four outputs: assets for the next node in the chain, plus image and audio so you can preview or reuse what it holds, and a report saying what was read and where it went.

Installing it

Part of WAS Node Suite v3 - ComfyUI Manager, search WAS Node Suite v3, or:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI 0.14.0+ and Python 3.10+, and the pack itself installs nothing.

What the node needs around it is the H3 model set: a transformer in models/diffusion_models (the community quants like minimax_h3_ref2va_pruned_w6a8.safetensors live under an H3 subfolder), minimax_h3_video_vae_int8_convrot.safetensors and minimax_h3_audio_vae_fp32.safetensors in models/vae, the text encoder in models/text_encoders, and - if you use a language model for the writing tools - a VLM via Load CLIP. Before you build on any of it: H3's weights are under a Community License whose Applicable Territory excludes the US, the EU, the UK and South Korea. Hosted Hailuo has no such problem; the local weights do.

Where it goes wrong

Everything renders as though the pictures weren't there. The asset chain has to be attached to the conditioning node's assets input - a chain that ends nowhere is just a nicely formatted list. Also check segment: an asset pointing at segment 3 references nothing in segment 1.

A prompt references <Picture 2> and the model invents something. Either the second asset in the chain is in a different segment, or its role isn't a reference role. Only reference picture gets a <Picture N> number.

Keyframes land where you didn't ask. frame counts that segment's new frames, not frames of the finished video. Placing at 48 in segment 4 is two seconds into segment 4. The one exception is the (the video being made) entry in the file list, where the frame is the finished video's frame - useful for a callback shot near the end.

A wired socket loses to a dropdown. That's by design: if both a file and a socket are set, the socket wins. Set file to (wired input) to make what's happening obvious when you come back to the workflow in a month.

CategoryWAS Suite/Latent/Video

Inputs (10)

NameTypeDefaultDescription
fileCOMBOWhich picture, clip or sound to read, as `cast/alice.png [input]` or `takes/shot_03.mp4 [output]`. `(wired input)` reads the image, video and audio sockets instead, and a wired socket always wins over the file.
roleCOMBOfirst frameWhat this does for its segment. `first frame` = the segment opens on it; `last frame` = closes on it; `keyframe` = pinned at `frame`, as a still, a clip, a sound, or a clip with its sound; `reference picture`, `reference clip` and `reference audio` = named `<Picture N>`, `<Video N>` and `<Audio N>` in that segment's prompt.
segmentINT10–24Which segment it belongs to, as `1` for the opening segment or `3` for Segment 3. `0` = every segment, as a cast picture the whole run references.
frameINT0-3600–3600Where a `keyframe` lands, counted in the segment's new frames at 24 fps: `0` = its first new frame, `48` = two seconds in, `-1` = its last frame. With `(the video being made)`, the frame of the finished video to reference, as `120`.
assetsoptWAS_H3_ASSETSThe assets before this one, from another MiniMax H3 Asset's assets output. This asset joins the end of that chain, so a run of any length is one wire into the next node.
clip_startoptFLOAT0.00–36000Seconds into the file a clip or sound starts, as `0` or `12.5`.
clip_secondsoptFLOAT0.00–600Seconds of the file to read, as `5`, or `0` for the first 362 frames of a clip, about 15 seconds, and the whole of a sound. A reference clip is cut to the longest segment either way.
imageoptIMAGEA picture, or a batch of frames at 24 fps for a clip, in place of a file.
videooptVIDEOA video in place of a file: its frames, brought to 24 fps, and its sound.
audiooptAUDIOA sound in place of a file, or the soundtrack of the wired image frames.

Outputs (4)

NameTypeDescription
assetsWAS_H3_ASSETSThe chain with this asset at its end, for the next MiniMax H3 Asset or for assets on MiniMax H3 Conditioning.
imageIMAGEThe picture or frames it holds, for a preview or another node.
audioAUDIOThe sound it holds.
reportSTRINGWhat was read and where it goes.