Nodes/comfyui-minimax-h3-audio-T8/H3 HyperFlow P7 · Original Learned 3D Lift (T8 EXP)
ComfyUI Node

H3 HyperFlow P7 · Original Learned 3D Lift (T8 EXP)

The upscaler that isn't an upscaler, and why audio survives it

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 HyperFlow P7 · Original Learned 3D Lift (T8 EXP)
  • low_result
  • learned_lift
  • enlarged_av
  • report_json
◄upscaler_model▾►

Two-pass generation has always leaned on a resize step between passes - the second pass needs something bigger to work on. The interesting thing about H3 is that "bigger" isn't just a picture. The latent is joint: video and audio in one tensor. So the resize has to be a learned 3D resampler that knows what it's looking at, and the audio has to come through it intact.

That's this node. It runs the original P7 learned 3D lift on the LOW phase's predicted clean output (not its final sampler output - more on that below), targets the HIGH dimensions the chain was configured with, and preserves the audio. It's a standalone operation: no diffusion, no noise, no second model.

Inputs

low_result - the typed bound LOW result. It must be a bound result, not a loose latent, because the lift needs to know that the LOW phase actually completed and which parent it belongs to.

upscaler_model - a combo listing whatever's in models/latent_upscale_models. This is the H3 learned 3D latent upscaler, not a pixel-space upscale model and not an SDXL ESRGAN-style thing. If your dropdown is empty, that directory is empty, and nothing downstream will work regardless of how the rest of the graph is wired.

Outputs

learned_lift is the typed handoff object for Original HIGH Audio Reconcile. enlarged_av is the LATENT itself, if you're wiring by hand or want to inspect geometry. report_json records the target dimensions and the resize contract.

The trap this node is placed to avoid

The LOW phase produces two outputs, and they are not interchangeable. Socket 0 is the sampler's terminal output; socket 1 is the predicted clean x0 - the denoised output for the 0:4 interval. The P7 lift wants the predicted clean x0. Wire the terminal output in and you get a "clean" picture that is not actually clean, and the failure is quiet: the HIGH pass will run, produce something, and look slightly worse than it should for reasons you'll never find by staring at the sampler settings.

The pack handles this by making the choice for you - Bind Completed LOW Result returns the right socket as low_x0 and the lift consumes the bound result - which is exactly why you should use the bind node rather than reaching for the sampler's own outputs. The same socket-swap mistake exists in the pack's other upscale routes and the docs call it out explicitly, so it's a known sharp edge rather than a hypothetical.

Where it sits

LOW 0:4 Stage Sampler → Bind Completed LOW Result → Original Learned 3D Lift
                                                    → Original HIGH Audio Reconcile

One lift per segment, and it happens between the two diffusion passes. If you're coming from the Fresh route's HyperFlowFreshLiftInputEXPT8, note that's a different route with a different typed contract - the names are similar, the objects don't cross, and the Fresh route distinguishes hyperflow_low_full8 (lift from completed output) from hyperflow_low_partial4 (lift from denoised_output). P7 is the partial-4 flavour, all the way down.

Install

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Full restart of ComfyUI, then refresh the browser. The pack's requirements.txt is deliberately empty - no pip installs, so it can't replace your Torch/CUDA build. Model-side you need the H3 base weights, Qwen text encoder, both VAEs, the HyperFlow adapter in models/hyperflow/loras/, and this node's upscaler in models/latent_upscale_models.

If it fails

"No upscaler selected" is the boring one. The more annoying one is a shape mismatch reported at the lift: that means LOW was run at a geometry whose 2× relationship to HIGH doesn't hold, so check that low_width/low_height in the segment selector are exactly half the HIGH values on the 32-grid. And because this is a GPU-side resize over a joint AV latent at full HIGH dimensions, it is not a free operation - if you're near your VRAM ceiling, the lift is a plausible place for the process to die, not just the HIGH diffusion pass that follows it.

CategoryT8/MiniMax H3/Modular Sampling/HyperFlow P7 Experimental

Inputs (2)

NameTypeDefaultDescription
low_resultT8_HYPERFLOW_P7_LOW_RESULT—
upscaler_modelCOMBO0 options:

Outputs (3)

NameTypeDescription
learned_liftT8_HYPERFLOW_P7_LEARNED_LIFT—
enlarged_avLATENT—
report_jsonSTRING—