Nodes/comfyui-minimax-h3-audio-T8/FastH3 V2 · Learned 3D Upscale + Origin (T8 EXP)
ComfyUI Node

FastH3 V2 · Learned 3D Upscale + Origin (T8 EXP)

The learned 3D lift with a paper trail

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
FastH3 V2 · Learned 3D Upscale + Origin (T8 EXP)
  • av_latent
  • av_latent
  • width
  • height
  • report_json
  • upscale_receipt
◄model_name▾►
◄size_modescale_by►
◄scale_by2.00►
◄target_megapixels0.70►
◄target_width1152►
◄target_height640►
◄aspect_policypreserve_source►
◄max_anisotropy1.05►
◄precisionfp16►
◄release_policyoffload_after►

What it is

The FastH3 V2 pipeline is LOW at 256×384, a learned 3D lift, then HIGH at 512×768 with a second 4-step pass. This node is the lift. It runs the pack's learned 3D latent upscaler once and attaches a typed receipt describing exactly what it fed in, what it produced, and what the report said.

Two design choices tell you what it's for. It runs once - no second upscale hiding inside, no HIGH sampling smuggled into the same node. And it emits upscale_receipt, which the handoff attestation downstream requires in order to prove the HIGH stage started from a lifted latent that really came from your LOW pass.

The distinction between upscaling pixels and upscaling latents is worth a sentence, because it's the reason this approach exists at all: a latent-space lift costs a fraction of a decode → upscale → re-encode round trip, and on a video model the round trip is where you lose the most time. The price is that you're trusting a learned model to place the extra detail in latent space rather than a familiar image upscaler placing it in pixels. That trade is H3's design, not this node's.

How it works

It wraps the pack's existing learned 3D upscaler with a provenance layer: it records the settings that matter, the input latent identity, the output latent identity, and the underlying report, then hashes the lot. What comes out the other side is both a LATENT you keep wiring and a receipt that nodes downstream can check. No second upscale call runs here, and no HIGH sampling - the node's description is explicit, and it's the reason you can put it in a resume graph without re-running the expensive half.

The settings you'll actually touch

  • av_latent - the LOW x0 from the completed LOW stage.
  • model_name - which learned upscaler checkpoint to load.
  • size_mode - scale_by by default; the alternatives are megapixels or explicit dimensions.
  • scale_by (default 2) - the multiplier when you're scaling by ratio.
  • target_megapixels, target_width, target_height (0.7, 1152, 640) - used when the size mode calls for them.
  • aspect_policy - preserve_source is the safe default; honor_dimensions_exp lets X and Y scale differently while still enforcing max_anisotropy.
  • max_anisotropy (default 1.05) - the guard rail against stretching your video into a funhouse mirror.
  • precision (fp16) and release_policy - the memory dial. offload_after keeps a CPU cache but frees GPU weights, clear_after drops the CPU cache too, and keep_loaded is opt-in and holds GPU memory on purpose.

Outputs: av_latent (wire into reconcile), width, height, report_json, and upscale_receipt.

The memory question, honestly

Pick release_policy with your card in mind, not with a hope. keep_loaded is the fastest if the upscaler stays resident, and it's the wrong answer on a 16GB GPU driving an H3 HIGH pass right after. offload_after is the reasonable default posture; clear_after when something else needs the space and you don't mind a reload. The pack's own cold/hot measurements showed FastH3 V2 being faster than the native 8-step route without any VRAM reduction - so don't expect the speed route to buy you headroom you don't have.

Install

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Quit ComfyUI completely, restart, refresh the page. The pack's requirements.txt deliberately installs nothing, so there's no pip step and no chance of it replacing ComfyUI's Torch/CUDA stack. You'll need a recent ComfyUI with native H3 support, H3 in models/diffusion_models, Qwen in models/text_encoders, both VAEs in models/vae, the FastH3 V2 ConvRot INT8 checkpoint, plus the learned 3D upscaler weights this node references by name.

Common issues

If model_name comes up empty or the load fails, the upscaler weights aren't where the pack expects - that's a file-placement problem, not a graph problem, and the error will name the file.

The failure that costs people an hour is watching the downstream handoff attestation reject the upscale receipt. That happens when the latent you're legally handing forward isn't the one the upscale actually produced - usually because a node in between cloned or re-typed it, or because you re-ran a stage after the upscale without re-running the upscale. The receipt is bound to the real output; regenerate both together.

And if report_json shows a geometry you didn't ask for, check aspect_policy before blaming the upscaler. preserve_source will round to keep your aspect ratio, which is usually what you want and occasionally not what you typed.

CategoryT8/MiniMax H3/Modular Sampling/Continuation Experimental

Inputs (11)

NameTypeDefaultDescription
av_latentLATENT—
model_nameCOMBO0 options:
size_modeCOMBOscale_by3 options: scale_by, target_megapixels, target_dimensions
scale_byFLOAT2.001–4—
target_megapixelsFLOAT0.700.01–8—
target_widthINT115232–4096—
target_heightINT64032–4096—
aspect_policyCOMBOpreserve_sourcepreserve_source is the safe default. honor_dimensions_exp permits different X/Y scales but still enforces max_anisotropy.
max_anisotropyFLOAT1.051–2—
precisionCOMBOfp163 options: fp16, bf16, fp32
release_policyCOMBOoffload_afteroffload_after keeps a CPU cache but releases GPU weights; clear_after also removes the CPU cache; keep_loaded is opt-in and retains GPU memory.

Outputs (5)

NameTypeDescription
av_latentLATENT—
widthINT—
heightINT—
report_jsonSTRING—
upscale_receiptT8_FAST_H3_V2_UPSCALE_RECEIPT—