H3 Image Generation - SatoDive
Three presets that hide the fiddly parts of MiniMax H3
- model
- vae
- positive
- latent
- IMAGE
This is the "just get me an image" node: it takes the positive and latent that H3 Prompt & Size produced, stacks a turbo LoRA, samples, optionally runs a second low-noise refine pass, decodes the H3 still, and optionally pixel-upscales it. One node, one IMAGE output.
The reason it exists as a separate node from H3 Image (Simple) is the presets. Simple gives you one pass and full manual control; this one gives you three canned recipes plus a Custom mode, which is the faster route when you don't yet know what H3 likes.
The presets, and what they quietly do
Pick a preset and it overrides your steps, lora_strength and refine widgets. That's the trap and the feature:
- Fast 2-pass - LoRA strength 0.38, 20 steps, then a refine pass at scale 1.5 with 10 steps. The default, and a sensible first look.
- Fast draft - same sampling, no refine pass. Composition hunting.
- Max quality - LoRA strength 0.0 (i.e. no LoRA at all), 50 steps, no refine. This is the base-model path; it's slow, and if you have a good turbo LoRA it's rarely worth it.
- Custom - your widgets are respected, and a refine pass runs only if
refine_scaleis above 1.0.
So if you drag lora_strength to 0.9 on Fast 2-pass and nothing changes, that's why. Switch to Custom.
How the refine pass works
Refine means: bump the latent's spatial grid by refine_scale (1.0–2.5), re-sample it briefly, decode. The clever bit is refine_strength, which is "the noise level (sigma) the refine pass starts from," not a denoise fraction. That distinction matters - under a flow shift, a denoise of 0.45 at shift 12 corresponds to a sigma around 0.91, which is nothing like what you meant. So the node builds a full sigma schedule, finds the point where sigma dips below your value, and keeps the tail. 0 means auto, derived from the scale: roughly 0.40 + 0.30 × (scale − 1), capped at 0.85. At the default 1.5 that's about 0.55. Lower keeps more of the draft; higher lets the refine pass invent more.
The refine pass runs on seed + 1, so it doesn't collide with the draft's noise, and upscale_model / upscale_to_mp then optionally run an ESRGAN-style pixel upscale after decode - native keeps the model's own factor, otherwise the result is resized to roughly the megapixel target you pick.
Inputs worth naming
model, vae, positive, latent come from your loader and H3 Prompt & Size. Then preset, seed, lora_name, steps, lora_strength, sampler_name (defaults to er_sde), scheduler (defaults to simple), refine_scale, refine_steps, refine_strength, upscale_model, upscale_to_mp. Output is a single IMAGE.
Note what's not here: no negative, no CFG. H3 runs single-conditioning, so there's no CFG to tune and no negative prompt to fight - the turbo LoRA plus ~20 steps is doing that job.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/SatoDive/ComfyUI-H3-IMG-Gen-SatoDive
Or ComfyUI Manager → ComfyUI-H3-IMG-Gen-SatoDive (MiniMax H3 Image Gen - SatoDive), then restart. No extra Python dependencies, but you need a ComfyUI with the native MiniMax H3 nodes. Optional extras: a turbo LoRA in models/loras, and an upscale model in models/upscale_models if you want the pixel-upscale stage.
Where it bites
Step count must match the LoRA. A 4-step distilled LoRA at 20 steps isn't better, it's artefacts - the model was trained to make four large jumps and you're asking for twenty small ones. Presets assume a ~20-step turbo LoRA; if yours is 4- or 8-step, use Custom and say so.
The refine pass is where VRAM goes. Latent upscale at 1.5 means 2.25× the pixels in the second sampling pass. On a small card, Fast draft plus a pixel upscale afterwards often beats Fast 2-pass, especially while you're still deciding whether you like the image.
A 4× upscale model on a 3 MP still is a ~48 MP intermediate before the resize back down. The node does that in VRAM, in one shot. It's a real, unhurried "why did my machine stop responding" moment - drop upscale_to_mp to 8 or 12 rather than letting native run, or skip the upscale stage entirely.
Inputs (16)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| vae | VAE | — | |
| positive | CONDITIONING | — | |
| latent | LATENT | — | |
| preset | COMBO | Fast 2-pass | 4 options: Fast 2-pass, Fast draft, Max quality, Custom |
| seed | INT | 00–18446744073709550000 | — |
| lora_name | COMBO | 1 options: None | |
| steps | INT | 201–200 | Custom preset only. |
| lora_strength | FLOAT | 0.380–2 | Custom preset only. |
| sampler_name | COMBO | er_sde | 44 options: euler, euler_cfg_pp, euler_ancestral, euler_ancestral_cfg_pp, heun, heunpp2, +38 |
| scheduler | COMBO | simple | 9 options: simple, sgm_uniform, karras, exponential, ddim_uniform, beta, +3 |
| refine_scale | FLOAT | 1.501–2.5 | Custom: 1.0 = no refine pass. |
| refine_steps | INT | 101–100 | — |
| refine_strength | FLOAT | 0.000–0.95 | Noise level (sigma) the refine pass starts from. 0 = auto (from refine_scale). Lower keeps more of the draft. |
| upscale_model | COMBO | Optional pixel upscale (ESRGAN-style) applied after decode. | |
| upscale_to_mp | COMBO | native | native = the model's own factor; otherwise resize the result to about this many megapixels. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| IMAGE | IMAGE | — |