H3 Cast Board
Your reference images, uploaded into the node and tagged for you
- ref_images
- ref_audios
- ref_videos
- cast
- image_1
- image_2
- image_3
- image_4
- image_5
- image_6
- image_7
- image_8
- audio_1
- audio_2
- audio_3
- audio_4
- video_1
- video_2
- video_3
- video_4
- report
H3 Cast Board is the answer to a boring problem that eats an afternoon: you have a rapper photo, a product shot, a wardrobe reference, a track and a first frame, and every one of them needs to be a LoadImage or LoadAudio feeding a numbered ref_image_0..N socket on the sampler, with a matching <Subject N> / <Picture N> tag written into the prompt by hand. Get one number wrong and H3 confidently puts the product shot on someone's face.
This node holds all of that. You drop files onto the node itself, assign each one a role, and it uploads them, tags them, numbers them and hands out the wired outputs.
Roles decide the tags, and this is the bit people get wrong
The H3 guide is explicit that a character image is not automatically a <Picture N>. If the image only supplies reusable identity - a face, a product, a style, a room - it's a <Subject N>. <Picture N> is reserved for a concrete frame: first frame, last frame, keyframe, composition anchor.
So the board maps twelve roles onto four tag kinds:
- character, product, style, wardrobe, environment, prop →
<Subject N> - first_frame, last_frame, keyframe, composition →
<Picture N> - video →
<Video N> - audio →
<Audio N>
The four kinds are numbered independently. <Subject 1> and <Picture 1> can be completely different assets, and that's correct behaviour rather than a bug.
How it behaves, mechanically
Files land in input/h3_planner/, which ComfyUI already serves, so a saved workflow reopens with its cast intact - no "missing image" red nodes next time.
Image, audio and video cards play in place, and audio and video cards have in and out points, so you can trim a reference without leaving the node. The trimmed file is what leaves the board, not the original.
Numbering follows card order within each kind. Reorder the board and the tags renumber - and the wired outputs move with them, which is the point: if a prompt says <Subject 2>, that prompt keeps pointing at the slot it was written for. Copy tags dumps the whole tag block to your clipboard, and JSON lets you edit cast_json by hand.
The three reference inputs grow on their own: wire ref_image_0 and ref_image_1 appears. That's ComfyUI's own Autogrow, the same mechanism H3's sampler uses, so the two nodes behave identically. Wired references are numbered on after the uploaded cards, so a wired image doesn't steal <Subject 1> from an upload.
Inputs and outputs
The only required input is cast_json - managed by the node's UI, editable as raw JSON if you like. Optional ref_images, ref_audios and ref_videos are the autogrow families.
Outputs: cast (the structured object every other planner node reads), image_1…image_8, audio_1…audio_4, video_1…video_4 (IMAGE frames - which is exactly what H3's ref_video_ inputs want), and report.
The slots are a fixed set rather than autogrowing, and that's deliberate: an output link is bound to a slot index, so a node that adds or hides output sockets silently repoints every connection downstream. A card past the last slot of its kind doesn't get tagged with nothing behind it - it's left out of the cast and named in the report.
Install
Nothing extra. Install the pack and restart:
cd ComfyUI/custom_nodes
git clone https://github.com/AIJigyasa/ComfyUI-H3-Planner
Or search MiniMax H3 Planner (the registry name) in ComfyUI Manager. Developed against ComfyUI 0.35.0, and it needs a ComfyUI with the MiniMax H3 nodes - the board's outputs wire into MiniMaxH3ReferenceToVideo.
Where people get burned
Don't hand-edit tags in a prompt after you reorder the board. Reordering is safe because the outputs move with it, but a prompt you typed <Subject 3> into by hand now points somewhere else. Either write prompts from the board's tags, or wire cast into the planner nodes and let them bind the tags for you.
A video card gives you IMAGE frames, not a video object. That's what the sampler wants; don't go looking for a VIDEO output to convert.
The board doesn't sample. It has no model, no weights and no VRAM cost - it's the reference layer, and the pack is explicit that your checkpoint, LoRAs, sigmas and latent upscaler stay exactly where they were in your own graph.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| cast_json | STRING | {"entries": []} | managed by the Cast Board UI; press JSON on the node to edit it by hand |
| ref_imagesopt | COMFY_AUTOGROW_V3 | — | |
| ref_audiosopt | COMFY_AUTOGROW_V3 | — | |
| ref_videosopt | COMFY_AUTOGROW_V3 | — |
Outputs (18)
| Name | Type | Description |
|---|---|---|
| cast | H3_CAST | — |
| image_1 | IMAGE | — |
| image_2 | IMAGE | — |
| image_3 | IMAGE | — |
| image_4 | IMAGE | — |
| image_5 | IMAGE | — |
| image_6 | IMAGE | — |
| image_7 | IMAGE | — |
| image_8 | IMAGE | — |
| audio_1 | AUDIO | — |
| audio_2 | AUDIO | — |
| audio_3 | AUDIO | — |
| audio_4 | AUDIO | — |
| video_1 | IMAGE | — |
| video_2 | IMAGE | — |
| video_3 | IMAGE | — |
| video_4 | IMAGE | — |
| report | STRING | — |