Nodes/was-node-suite-comfyui/MiniMax H3 Clip Select
ComfyUI Node Runs on cloud

MiniMax H3 Clip Select

One clip at a time, out of a prompt bundle

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
MiniMax H3 Clip Select
  • prompts
  • positive
  • latent
  • frames
  • report
  • model
◄index0►

Every MiniMax H3 workflow eventually splits into two jobs, and they are not the same job. One is continuation: shots that carry on from each other into one long take. The other is a set of separate clips - the same character, the same style, five different scenes that you'll cut together in an editor. MiniMax H3 Clip Select is the second one.

It takes the prompt bundle from MiniMax H3 Conditioning, an index, and gives you back that clip's prompt and that clip's own empty latent. Nothing passes between clips. Each one is sampled from scratch, which is exactly what you want when they're separate shots, and exactly what you don't want when they're supposed to be one shot.

Inputs and outputs

prompts is the WAS_H3_PROMPTS bundle - the only thing that produces it in this pack is the conditioning node, which is where the prompts, durations and overlaps are actually written. index picks the row: from 0, up to 999. Drive it from a While Loop's index and you step through the clips one per iteration, which is the intended shape of the workflow.

positive is that clip's encoded prompt, for the guider that samples it. latent is its own empty latent, sized for the clip's own frame count and the canvas the conditioning node worked out. frames gives you the frame count for the clip - 124, say - which is what you feed to a video node's length. report is a string saying which clip was taken and how long it runs, and it's more useful than it sounds: it's how you confirm that the durations you typed survived snapping to H3's frame grid.

Note what's not here: no negative, and no start-image input. The mode you chose on the conditioning node decided what the clip is conditioned on; this node just hands you the row.

Why this exists at all

Because batching separate clips in one graph is otherwise miserable. Without it you'd duplicate the whole conditioning chain per clip, or manually re-prompt between runs, and you'd lose the one property that makes a prompt-bundle-and-index design worth having: the prompts live in one editable list, and swapping the loop's index range changes how many clips you get without touching a single connection. Add a row on the conditioning node, raise the loop count, done.

The trade is that there's no continuity to lean on. Two clips generated from the same bundle will share whatever the prompts share - style, subject, wardrobe - and no actual pixel history, because there isn't any. If that's not enough consistency for the shot you want, the answer isn't this node; it's ref2va mode on the conditioning node so every clip is built on the same reference pictures, or H3 Extend Window if the clips are really one continuous take.

Installing it

It ships in WAS Node Suite v3 - MIT, WASasquatch, 468 nodes - so you install the pack:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI Manager's WAS Node Suite v3 entry is the recommended install. You'll need ComfyUI 0.14.0+ and Python 3.10+, and that's it from the pack's side: no pip packages installed, nothing fetched, nothing built. The first start after installing (and after any update) runs a little longer while it writes its config and compiles to bytecode under <ComfyUI user dir>/was-node-suite/. Ignore the v2-era "Cannot import … custom_nodes\was-node-suite-comfyui" posts; those were pip dependencies that v3 removed.

Where people get stuck

The index is zero-based and the rows start at prompt_1, which is index 0. Off-by-one here is the most common mistake, and the symptom is a blank first clip rather than an error - a row is only live if it has a prompt in it.

Decode inside the loop. The whole design assumes each iteration samples, decodes and saves its own clip, rather than accumulating latents to decode at the end. Latents from different rows aren't frames of one sequence, and stacking them will produce something that looks like a video and isn't.

And be realistic about hardware. H3 is a 33B omni-modal model with reported weights around 42.5 GB; "one clip per loop iteration" is honest about cost in a way that a hidden batch never is. Ten clips is ten generations, and you'll feel it.

CategoryWAS Suite/Latent/Video

Inputs (2)

NameTypeDefaultDescription
promptsWAS_H3_PROMPTSEvery clip's prompt and length, from MiniMax H3 Conditioning.
indexINT00–999Which clip to take, from `0`. Wire a While Loop Open's index in to step through them one per iteration.

Outputs (5)

NameTypeDescription
positiveCONDITIONINGThat clip's prompt, for the guider that samples it.
latentLATENTThat clip's own empty latent, for the sampler's latent input.
framesINTFrames that clip runs for, as `124`.
reportSTRINGWhich clip was taken and how long it is.
modelMODELThe model that clip's row chose, from the ones wired into MiniMax H3 Conditioning, for the guider that samples it. Blocked with a message where none is wired.