H3 Prompt & Size - SatoDive
The messy front half of an H3 graph, collapsed into one node
- clip
- vae
- ref_image_1
- ref_image_2
- ref_image_3
- ref_image_4
- ref_image_5
- ref_image_6
- ref_image_7
- ref_image_8
- ref_image_9
- positive
- latent
Every ComfyUI workflow is mostly plumbing, and with MiniMax H3 the front of the graph is the worst of it: a text encoder, a size you have to compute in multiples of 32, reference images that need a different conditioning path than a plain text prompt, and a latent that isn't shaped like any normal image latent. H3 Prompt & Size does all of that and hands you two clean outputs - positive and latent.
Reach for it when you want to build your own sampler chain instead of using the pack's all-in-one nodes, or when you want the same conditioning feeding more than one branch. It's the front half of the graph, pre-assembled.
What it actually builds
Two things, and the second one is the interesting one.
The conditioning comes from ComfyUI's native H3 nodes. If any ref_image_* slot is connected it uses MiniMaxH3ReferenceToVideo (H3's reference/edit pathway, which is what makes <Picture N> work); with no refs it falls back to MiniMaxH3ImageToVideo. Either way it's called at length=5, the minimum H3 context, because for a still the conditioning doesn't depend on length - so you're not paying for frames you'll never sample.
The latent is a single-frame H3 still latent: one video frame with an even latent grid (the model's DiT patchifies 2×2, so the grid has to stay even) plus a small audio latent, wrapped together as a nested tensor. That's not an Image latent for a normal KSampler - it's an H3 latent, and it's the thing you'd otherwise be assembling by hand. H3 only supports batch size 1, so it's one latent, one image per pass.
The inputs you set
promptand theref_image_1–ref_image_9slots. The rule, straight from the tooltip: "N counts only connected slots, in slot order." Unused slots in the middle are invisible to numbering, so<Picture 1>is your first connected image.aspect+megapixels, oraspect = Customwithwidth/height(snapped to multiples of 32). AspectCustomalso ignoresmegapixels.ref_image_size-matchscales refs down (never up) to the output area;maxkeeps them at a 2048px short edge for the best identity at a real speed and VRAM cost. The default here ismatch, which is the friendlier one while you're still hunting a composition.
There's no negative conditioning output and no CFG field, because H3 doesn't use one - the pack's samplers drive it with a single-conditioning guider. Don't go looking for a prompt-weight workaround; it isn't there.
Outputs
positive (CONDITIONING) goes to a guider or a sampler's positive input. latent (LATENT) is your starting image. Wire them into SamplerCustomAdvanced or into this pack's H3 Image Generation / H3 Draft Grid, which is what most people do.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/SatoDive/ComfyUI-H3-IMG-Gen-SatoDive
ComfyUI Manager will find it under ComfyUI-H3-IMG-Gen-SatoDive (MiniMax H3 Image Gen - SatoDive); restart after installing. Needs a ComfyUI build with the native H3 nodes, plus your H3 model, text encoder and VAE. No extra Python packages - the pack declares no dependencies at all.
Gotchas
Size maths is deliberately boring. Pick 16:9 at 3 MP and you get a multiple-of-32 pair that's near 3 MP, not exactly 3 MP - that's how the model wants it, and the exact-size fix happens downstream in the Simple or Final Size nodes.
ref_image_size is the main speed dial. Going max on a laptop GPU is the fastest way to make a workflow feel broken: the refs stay big through the text-and-VAE encode, which is exactly where small cards start swapping models in and out. Try match, then move back once the composition is locked.
One conditioning, one reuse. This node re-encodes prompt and refs every run. That's the right behaviour for a building block, but if you're re-rolling seeds in a hand-built graph, the pack's Simple node - which caches the last encode - will feel dramatically faster. That difference is the node, not your GPU.
Inputs (17)
| Name | Type | Default | Description |
|---|---|---|---|
| clip | CLIP | — | |
| vae | VAE | — | |
| prompt | STRING | — | |
| aspect | COMBO | 16:9 | Custom = use width and height below. |
| megapixels | COMBO | 2.5 | Ignored when aspect = Custom. |
| width | INT | 1920256–8192 | Used when aspect = Custom (snapped to a multiple of 32). |
| height | INT | 1088256–8192 | Used when aspect = Custom (snapped to a multiple of 32). |
| ref_image_1opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_2opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_3opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_4opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_5opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_6opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_7opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_8opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_9opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_sizeopt | COMBO | match | match = refs scaled (down only) to the output area; max = 2048px short edge, best identity but slower. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | — |
| latent | LATENT | — |