H3 Latent Upscale - SatoDive
The one-slider refine step, and the two cases where it does nothing
- samples
- LATENT
This is the pack's smallest node and its most misunderstood one, because it doesn't upscale your image. It upscales the latent of an H3 still, which is only useful if you're about to sample it again.
That's the classic two-pass trick: sample a composition at a comfortable size, blow up the latent, then re-sample briefly at the bigger resolution so the model paints detail into the new pixels. The node's own description says what to pair it with - "a second sampler at denoise ~0.4-0.5". If you were expecting a Lanczos-style sharpening pass, you want a different node (H3 Final Size, or any ESRGAN upscaler); this one is a step inside a pipeline, not a finishing move.
What it does, precisely
H3's still latents are nested tensors - a video stream plus an audio stream, because H3 is a video/audio model underneath. This node takes that pair apart, spatially resizes the video stream only, reassembles the pair, and leaves the audio stream exactly as it was.
Two details from the implementation that matter in practice:
- The latent grid stays even. H3's DiT patchifies 2×2, so a grid with an odd dimension breaks, and the node rounds the new size to an even number and floors at 2. Your
scaleis therefore a target, not a promise - scaling a small latent by 1.05 can round straight back to where you started. - It only acts on a single-frame still. If the latent isn't nested (i.e. it's a normal image latent from anything else in ComfyUI), or if the video stream has more than one frame (an actual H3 video clip), it returns your input untouched. No error, no warning. That's the "why isn't this doing anything" case.
The three inputs
Just samples (LATENT), scale (1.0–2.5, default 1.5), and method - bicubic, bilinear or nearest-exact. Bicubic is the default and the sensible one; interpolation choice matters far less here than it does in a pixel-space upscale, because the follow-up sampler is going to rewrite most of what you interpolated.
One output: LATENT, ready for a second sampler. Note that you then need a decode that understands a one-frame H3 latent - that's exactly why this pack bundles its own still decoder (adapted from ComfyUI-Fizgig-H3-Still), and why, in this pack, the node mostly gets used inside H3 Refine Winner and the Fast 2-pass preset of H3 Image Generation rather than wired up by hand. ComfyUI's stock VAE Decode bands a lone H3 latent frame; know that before you build the manual version.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/SatoDive/ComfyUI-H3-IMG-Gen-SatoDive
ComfyUI Manager → ComfyUI-H3-IMG-Gen-SatoDive (MiniMax H3 Image Gen - SatoDive), restart. No Python dependencies beyond ComfyUI, but a recent ComfyUI with native H3 support is required - this node imports the nested-tensor machinery at pack load.
Gotchas
Silent no-op is the failure mode. Two ways to hit it, and both look like a broken node: feeding it a non-H3 latent, or feeding it a multi-frame H3 video latent. Check that the upstream is the pack's still latent before you spend twenty minutes debugging a workflow.
A bigger scale is not automatically a better image. 1.5 is a fair default. At 2.0 and up the second pass gets enough room to change the composition, and if you also let the noise level run high you'll get a different picture rather than a sharper one. If you're using this alongside a sampler's denoise, start low and walk it up.
Pixel upscalers are still the cheaper answer for "I want it bigger". If the detail you want already exists and you just need more pixels, a 4× ESRGAN model or Lanczos is milliseconds and cannot invent anything. Come to latent upscaling when the image is soft at the size it's already at and you want the model to paint - that's a different job, and it costs a full sampling pass at the larger resolution.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| samples | LATENT | — | |
| scale | FLOAT | 1.501–2.5 | — |
| method | COMBO | bicubic | 3 options: bicubic, bilinear, nearest-exact |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| LATENT | LATENT | — |