MiniMax H3 一键高清放大 - Star7
Hi-res fix for a 33B video model, without re-sampling the whole clip
- sampled_av_latent
- h3_context
- hd_av_latent
- report
Sampling MiniMax H3 at 4K for ten seconds isn't a thing most cards can do, and sampling at 1MP and pretending it's 4K isn't either. The old image-gen answer to this was hi-res fix: generate at native resolution, upscale the latent, then run a few low-denoise steps to let the model re-commit detail at the new size. MiniMaxH3OneClickHDStar7 is that recipe, packaged for H3's audio-video latent, in one node.
What's inside the one click
Two stages, in this order.
A learned 3D latent upscale. The video latent gets normalised with the upscaler's own fixed 24-channel mean/std, pushed through a 3D latent upscaler checkpoint, then de-normalised. That's a real learned spatial-temporal resizer, not bicubic on a tensor - which matters, because H3's latent is where the motion and the detail both live.
A short low-noise refine. The upscaled latent gets a noise_mask of ones over video and zeroes over audio, and then goes through SamplerCustomAdvanced with simple sigmas that only cover the tail of the schedule - that's what refine_strength controls. The default preset is 2 steps at 0.25. Then it checks the result is finite and reassembles the nested latent with the original audio member. The audio is never resampled, never re-diffused, never touched. For a model whose whole selling point is jointly-generated dialogue and room tone, that's the correct decision, and it's the difference between "sharper" and "subtly out of sync."
The output is a latent, not pixels. VAE decode stays outside, so the intended chain is:
Sampler "latent" -> One-Click HD -> MiniMaxH3ChunkedDecodeStar7 -> video combine
All-in-one Conditioning "context" ----^
Inputs that matter
enable_hd is the master switch, and turning it off returns your latent untouched - no model load, no second pass, no downloads. That's your A/B control.
h3_context is required and comes from the pack's MiniMax H3 多合一条件载入 - Star7 (All-in-one Conditioning). Without a model and a positive in it, the node raises an "HD context is incomplete" error. It wants that context because it needs the model, the prompt, and - usefully - any first/last keyframes you're generating from, which it re-encodes with the VAE at the new resolution so you're not stamping low-res endpoint latents onto a high-res pass.
target_megapixels (0.2–36) is the real size knob, aspect preserved. If the target is at or below your source size the node returns the latent unchanged and says so - it will not shrink anything.
preset is a paired steps-and-strength recipe, and the smart part is that it's profile-aware. Turbo-style checkpoints get 2–3 steps at 0.18–0.30; base or plain-LoRA H3 gets 4–6 steps, because it isn't distilled and shouldn't be treated as a few-step model. It detects Turbo/PDD by reading upstream names and mounted patches. With a PDD model it uses the trained partition tail and Euler; if a PDD file was loaded with a normal LoRA loader it errors instead of producing mush. The status line shows the exact sigma ladder used, from the actual model - no fixed shift guesswork, and sigmas minus one is your step count.
refine_steps / refine_strength / seed are what "自定义" (custom) unlocks when the presets aren't it.
upscale_model defaults to minimax_h3_latent_upscaler_3d_fp16.safetensors and downloads itself from the HF mirror, then Hugging Face, SHA-256 verified into ComfyUI/models/latent_upscale_models. Only that one file is exposed, deliberately - so the node can't silently grab an LTX upscaler or a BF16/FP32 sibling out of that shared folder and call it a different quality tier.
Tiling, and the attention override
enable_tiling is a pure VRAM lever for the refine pass. It splits each high-res prediction into an aspect-aware grid of overlapping 2D tiles - 16:9 asked for 16 tiles becomes 3×6 (18), square-ish gets 4×4 - with boundaries snapped to H3's 2×2 latent patches and cosine-weighted blending across tile_overlap. tile_count is treated as a floor, not an exact number. Tiling cuts peak refine VRAM and costs you model calls; it doesn't change sigmas, sampler, or audio. It has nothing to do with the decode node's internal VAE tiling.
second_pass_attention defaults to 继承一采 - inherit the first pass. The rest of the list is the pack's CK/SLA/Sol/Hybrid menu if you want the refine pass on a different kernel. Note it restores the pack's global attention config after the refine finishes, so an override here won't leak into a cached chunk node on your next queued run.
Install and where it bites
Same pack, one install: ComfyUI Manager → MiniMax H3 Activation Chunk - Star7, or
cd ComfyUI/custom_nodes
git clone https://github.com/star7code/minimax-h3-chunk-star7.git
Restart, then put the upscaler in models/latent_upscale_models or let the node fetch it. Pack dependencies (scipy, scenedetect, ultralytics) come along whether or not this node needs them.
The failures you'll actually meet: a missing or incomplete context wire; a non-nested latent handed in (it wants [B,24,T,H,W] video plus audio, and says so); an upscaler that fails its SHA check, with an error naming the directory and every source that failed; and time. Tiles and refine steps both cost real GPU minutes on a 33B video model - change one, then read the report string. For an A/B you don't need to delete the node; flip enable_hd off and compare.
Inputs (15)
| Name | Type | Default | Description |
|---|---|---|---|
| enable_hd | BOOLEAN | true | — |
| sampled_av_latent | LATENT | — | |
| h3_context | STAR7_H3_REFINE_CONTEXT | — | |
| upscale_model | COMBO | minimax_h3_latent_upscaler_3d_fp16.safetensors | 3D H3 latent-upscaler checkpoint in models/latent_upscale_models. The pinned FP16 default downloads automatically when missing. |
| second_pass_lora | COMBO | 继承一采 | Inherit the first-pass model unchanged, or apply the selected model-only LoRA for this HD refinement only. |
| second_pass_lora_strength | FLOAT | 1.00-100–100 | Model strength for the selected second-pass LoRA. |
| second_pass_attention | COMBO | 继承一采 | 20 options: 继承一采, existing, comfy_kitchen_int8, sla_sm75_qk_int8_pv_fp16, sla_sm75_all_int8, sol_sm75_all_int8, +14 |
| preset | COMBO | 平衡高清 | 5 options: 平衡高清, 高质量, 远景小脸, 高速运动, 自定义 |
| target_megapixels | FLOAT | 1.000.2–36 | — |
| refine_steps | INT | 21–50 | — |
| refine_strength | FLOAT | 0.250–0.5 | Fraction of the denoising path replayed by refinement. Higher values start from noisier latents and allow more repainting; 0 disables refinement. PDD uses its trained tail boundaries. |
| seed | INT | 00–18446744073709550000 | — |
| enable_tiling | BOOLEAN | false | — |
| tile_count | INT | 22–64 | — |
| tile_overlap | INT | 12832–512 | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| hd_av_latent | LATENT | — |
| report | STRING | — |