H3 Image (Simple) - SatoDive
One MiniMax H3 node, one pass, the size you actually typed
- model
- clip
- vae
- ref_image_1
- ref_image_2
- ref_image_3
- ref_image_4
- ref_image_5
- ref_image_6
- ref_image_7
- ref_image_8
- ref_image_9
- image
- width
- height
MiniMax H3 is a 33B video model with native audio. It's also quietly excellent at stills, because its image convention is a single latent frame - and this is the node that lets you make stills with it without hand-building the plumbing yourself.
The pitch is in the README and it's unusually literal: one node, one pass, what you set is what you get. No automatic refine, no hidden upscale, and megapixels is the final size. If you've been burned by a wrapper that quietly ran three passes and handed you something you didn't ask for, you'll get why that's the selling point.
How it works
H3 lives in ComfyUI as native nodes, and this pack drives them. The node builds the still latent - one video frame plus a small audio latent, because H3 is omni-modal and even a "still" carries an audio stream. It then gets conditioning from MiniMaxH3ReferenceToVideo if you connected any reference image, or MiniMaxH3ImageToVideo if you didn't, both at length=5, the shortest context H3 accepts. Sampling runs through ComfyUI's advanced sampler path with a plain guider on the positive conditioning: no negative, no CFG slider. That's normal for a turbo/distilled model, and it means if a prompt isn't landing you rewrite the prompt rather than hunting for a negative field that isn't there.
Then there's the decode, which is the part worth knowing about. ComfyUI's stock VAE Decode produces banding on a lone H3 latent frame, so the pack ships its own: it replicates the single frame into a full 5-frame group, decodes that, and keeps pixel frame 3, just past the decoder's causal lead-in. That code is adapted from ComfyUI-Fizgig-H3-Still (MIT), and it's why this isn't just the stock nodes with a nicer panel.
The inputs that actually matter
model, clip, vae come off whatever H3 loader you use. Then:
promptplusref_image_1–ref_image_9. Point at connected references with<Picture N>, where N counts only the connected slots, in slot order - that's the author's own tooltip, and it's the one thing people get wrong. Connect refs 1 and 4 and your second image is<Picture 2>.- Size:
size_modeis eitherAspect + megapixels(eight ratios, 0.25–16 MP) orCustom size(width/height).exact_sizeon means the result is resized by the few leftover pixels to hit your number exactly, since the model works in multiples of 32. stepsandlora_name/lora_strength. The guidance is worth repeating: a 4-step turbo LoRA wants ~4 steps, other turbo LoRAs ~20, no LoRA ~50. More steps than your LoRA was trained for is just a slower way to make artifacts.ref_image_size-maxis 2048px short edge and gives the best likeness but is slower and hungrier on VRAM;matchscales refs down to the output area. Start onmatchwhile you're iterating.detail_model/detail_strength- an upscale model used only to add fine detail, at the same resolution. See the gotcha below.
Outputs are image, plus width and height as INTs - the real size, after any exact-size resize, so you can wire them straight into a filename or a note node.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/SatoDive/ComfyUI-H3-IMG-Gen-SatoDive
Or find it in ComfyUI Manager (search the pack title, ComfyUI-H3-IMG-Gen-SatoDive, or the display name MiniMax H3 Image Gen - SatoDive), then restart. There are no Python dependencies beyond ComfyUI itself - but you do need a ComfyUI new enough to have comfy_extras/nodes_minimax_h3.py, plus your H3 model, text encoder and VAE. A turbo LoRA is optional.
Where people get burned
Old ComfyUI. If your build doesn't ship the native H3 nodes, the pack raises its own message - "This ComfyUI has no native MiniMax H3 nodes… Update ComfyUI to a version that includes MiniMax H3." That's the pack being friendly; update and move on.
Expecting the detail pass to redraw. It doesn't. It upscales with an ESRGAN-style model, shrinks back to your resolution with area filtering, and blends the result in by detail_strength - supersampling, i.e. a sharpener. It makes skin and distant faces crisper; it will not fix a bad face. At very low strength it does almost nothing, which is by design.
The conditioning cache is one entry. The node hashes prompt, size and ref-image tensors and skips re-encoding when nothing changed - and re-encoding is the genuinely slow part, especially on a small GPU where it forces model swaps. But it holds exactly one entry, so alternating between two prompts thrashes it. Fine, just don't expect magic.
The elephant. The H3 weights are ~42.5 GB (reported), there's no published consumer-VRAM floor, and the Community License excludes the US, EU, UK and South Korea - including for outputs. If you're in an excluded region, the local weights aren't licensed to you at all, and no node pack changes that.
Inputs (28)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| clip | CLIP | — | |
| vae | VAE | — | |
| prompt | STRING | — | |
| size_mode | COMBO | Aspect + megapixels | Aspect + megapixels: pick a ratio and an area. Custom size: type the exact width and height. |
| aspect | COMBO | 16:9 | Used only in 'Aspect + megapixels' mode. |
| megapixels | FLOAT | 3.000.25–16 | Used only in 'Aspect + megapixels' mode. This IS the final size, no hidden upscale. |
| width | INT | 192064–8192 | Used only in 'Custom size' mode. |
| height | INT | 108864–8192 | Used only in 'Custom size' mode. |
| exact_size | BOOLEAN | true | The model works in multiples of 32. ON: the result is resized by the few leftover pixels so it is EXACTLY the size you typed. OFF: keep the native multiple-of-32 size. |
| seed | INT | 00–18446744073709550000 | — |
| steps | INT | 201–200 | Match your LoRA: a 4-step turbo LoRA needs ~4 steps (more is just slower). Other turbo LoRA: ~20. No LoRA: ~50. |
| lora_name | COMBO | Turbo LoRA (optional). | |
| lora_strength | FLOAT | 0.380–2 | — |
| sampler_name | COMBO | er_sde | 44 options: euler, euler_cfg_pp, euler_ancestral, euler_ancestral_cfg_pp, heun, heunpp2, +38 |
| scheduler | COMBO | simple | 9 options: simple, sgm_uniform, karras, exponential, ddim_uniform, beta, +3 |
| detail_model | COMBO | Optional. Upscale model used ONLY to add fine detail (faces far away, textures). The resolution does NOT change. None = off. | |
| detail_strength | FLOAT | 0.700–1 | How much of the detail pass is mixed in. 0 = nothing, 1 = full. Lower keeps the image closer to the original. |
| ref_image_1opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_2opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_3opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_4opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_5opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_6opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_7opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_8opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_9opt | IMAGE | Reference image. Use <Picture N> in the prompt; N counts only connected slots, in slot order. | |
| ref_image_sizeopt | COMBO | max | max = references at full quality (2048px short edge, best identity, SLOWER and heavier on VRAM). match = scaled down to the output area (much faster, weaker likeness). |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| width | INT | — |
| height | INT | — |