Decode H3 Character Sheet
Four, Six or Eight Views of the Same Person
- samples
- vae
- sheet
- selected_views
- all_frames
- info
Character reference sheets - front, three-quarter, profile, back, all recognisably one person - are exactly the kind of thing that used to mean six separate generations and six slightly different faces. The trick everyone converged on is to stop generating stills and run a camera around the subject instead, then keep the frames you want. This node is the "keep the frames you want" half of that.
Why one pass beats six prompts
The encoder compiles a locked-character camera prompt and asks H3 to orbit the subject in a single continuous sequence - 73, 124 or 171 frames depending on the profile. Because identity, palette and proportions all share one denoising trajectory, they don't drift between panels the way six independent generations do. You're paying for a video-sample, which is the cost: a 124-frame REF2VA pass is substantially slower and heavier than the still profiles. Worth it. The alternative is a sheet where the jawline changes between column one and column three.
The decoder's other job is that it knows where the good frames are. The camera settles at specific moments during the orbit, and those moments are fixed indices, not something you hunt for:
4 panels- 73 frames, indices2, 24, 45, 68→ 2x2 sheet6 panels- 124 frames, indices2, 21, 42, 63, 84, 113→ 3x2 sheet8 panels- 171 frames, indices2, 21, 42, 63, 84, 108, 131, 160→ 4x2 sheet
Nine-ish seconds of video for eight calibrated angles, in one queue.
Inputs and outputs
Required: samples (the character-sheet latent from the encoder), vae (the H3 video VAE), layout, gutter_px (0–256, default 6) and gutter_color (black, neutral gray, white).
layout is the only one with a real decision in it. auto from encoded profile reads the metadata the encoder stamped on the latent and does the right thing - that's what you want, and it's the default. The explicit options exist so you can force 6 panels | 3x2 or 8 panels | 4x2, but forcing a layout that doesn't match the encoded frame count just earns you an error.
Four outputs, and the first two are the ones you'll use: sheet (the single stitched image), selected_views (those same four/six/eight frames as an IMAGE batch, in case you'd rather save them individually or feed them somewhere else), all_frames (every decoded orbit frame - manual frame-hunting, or a safety net when a panel came out with an odd expression), and info, which reports the resolved layout, decoded frame count, extraction indices and output size.
Install
Manager → search the pack title, or:
cd ComfyUI/custom_nodes
git clone https://github.com/ethanfel/ComfyUI-MiniMax-H3-Edit
Restart. No extra Python dependencies - the pack's requirements file installs nothing beyond ComfyUI's own runtime, which is rare enough to be worth saying twice.
Then load example_workflows/H3_Character_Sheet_6_Panel.json. It wires three chained native references through Add H3 Edit Reference into the encoder, uses 768x1344 frames, Euler with the linear-quadratic scheduler at 25 steps, and a REF2VA checkpoint. Replace all three placeholder images and rewrite the per-picture assignment prompt before you queue - Picture 1 defines identity and proportions, Picture 2 the outfit, Picture 3 the rear construction you can't see from the front. There's also H3_Character_Clothing_6_View.json if you only care about garments, which swaps the facial push-ins for six constant-scale head-to-knee captures at 0°, 60°, 120°, 180°, 240° and 300°.
Use a generation role, not the strong edit anchor. Native Picture 1 plus native builders is the default high-fidelity route; an all-semantic set runs Qwen-only on FL2VA with no input VAE encoding at all. Mixing semantic and native is flagged experimental for a reason.
Where people get burned
- Auto layout on the wrong latent. Feed it a still-image latent and it refuses: auto layout needs a 73-, 124- or 171-frame encoded profile. That's a feature - better an error than a 2x2 made from re-posed stills.
- Doing the encoder config by hand. The character-sheet compilers write a locked-camera prompt themselves; your text only needs to assign a job and an ignore list to each
<Picture N>. Writing a full camera description on top of that is how you get a sheet with an unrequested zoom in panel four. - Memory. 171 frames is a lot of latent to sample and decode in one go. The 4-panel profile is the honest starting point if a 124-frame run is already tight.
- Panel mismatch on re-decode. If you decode the same latent twice with different
layoutvalues, the wrong one errors out rather than silently cropping - but you'll lose the queue time, so set it right the first time (or leave it on auto).
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| samples | LATENT | Sampled H3 character-sheet latent from Text Encode H3 Edit / Generate. | |
| vae | VAE | MiniMax H3 video VAE. | |
| layout | COMBO | auto from encoded profile | Auto reads the encoded 73/124/171-frame profile metadata. |
| gutter_px | INT | 60–256 | — |
| gutter_color | COMBO | black | 3 options: black, neutral gray, white |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| sheet | IMAGE | One stitched 2x2, 3x2, or 4x2 character sheet. |
| selected_views | IMAGE | The four, six, or eight calibrated views as an IMAGE batch. |
| all_frames | IMAGE | Every decoded orbit frame for manual inspection or alternate picks. |
| info | STRING | Resolved layout, decoded frame count, extraction indices, and output size. |