Nodes/comfyui-minimax-h3-audio-T8/H3 一采动态预览+安全取消(TAEH3 EXP/T8)
ComfyUI Node

H3 一采动态预览+安全取消(TAEH3 EXP/T8)

Watch the Motion Mid-Sample, Then Kill the Take You Hate

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 一采动态预览+安全取消(TAEH3 EXP/T8)
  • model
  • model
◄enabledtrue►
◄checkpoint▾►
◄phase▾►
◄update_every_steps2►
◄min_interval_ms500►
◄max_resolution256►
◄latent_prefix▾►
◄frames12►
◄fps12►
◄jpeg_quality75►

What it is

H3 generations are long. Five or ten minutes in, you often already know the shot is wrong, and ComfyUI's default preview gives you a blurry latent teaser that tells you almost nothing about motion. This node is the fix: it stretches a MODEL pass-through between your model loader and your existing sampler, watches the real x0 the sampler produces, and streams a tiny decoded video preview into a panel while the job runs. Bad take? Cancel it there and re-roll, instead of waiting out the full run.

It's the H3-specific version of a very old trick - tiny approximation VAEs feeding live previews during sampling. The community route into that world is well known: drop a small TAE model in vae_approx/, plug the preview node in, and you get a fast, low-quality look at what's happening. Kijai's tiny-VAE work and the general "preview during generation so you can cancel midway if the scene doesn't look right" pattern are the same idea, and this node is the H3 variant of it.

How it works

Your sampling is untouched. The sampler keeps its own sigmas and wiring; the node just observes the callback and decodes the head of the clip. Two things it is emphatically not: not an extra diffusion step, and not a look at the audio. There's a small decode and transfer cost per update, which is what the throttling knobs are for.

The temporal TAEH3 decoder has real temporal context, so a latent prefix maps to a short run of frames - roughly 5, 22 or 39 frames for prefixes of 2, 7 or 12. The optional 2D tiny model (Kijai's, saved under a different filename) has no temporal context: it decodes per-latent-frame approximations and the panel labels them as such, with no source fps, because calling 7 approximated latent frames a continuous 24 fps clip would be a lie.

Output is one socket, model, and it goes exactly where your sampler used to get the model from. Disable enabled and the original MODEL object is returned untouched.

The inputs that matter

  • checkpoint - the tiny decoder, picked from files starting with taeh3 in models/vae_approx/. taeh3.safetensors is the temporal default. If you also pull the 2D model, save it as taeh3_2d_kijai.safetensors and select it explicitly; overwriting the temporal file is the mistake this tooltip exists to prevent.
  • enabled, phase (low/high/all), update_every_steps (default 2), min_interval_ms (500), max_resolution (256), latent_prefix (default 7), frames (12), fps (12), jpeg_quality (75). Lower max_resolution or raise the update interval when you want the overhead out of the way; those are the only knobs that meaningfully change cost.
  • Phase labels come from reality. LOW/HIGH reflect the actual producer phase; a plain sampler gets labelled as unidentified rather than being guessed at as LOW. If you're on the dual-pass workflow, that distinction is the point.

The panel has a cancel button, and its scope is worth quoting precisely: it cancels the request bound to that panel via Core's per-request cancel API. It does not clear your queue, does not interrupt other jobs, and does not stop the service. Pause pauses playback, not sampling. And if the preview model is missing or unusable, your original sampling simply continues.

Installing

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Restart fully and refresh the page, or search MiniMax H3 Audio T8 in Manager. Then grab t8star/Taeh3-Comfy and merge its models/ folder into ComfyUI/models/, keeping the vae_approx/ directory intact. The pack puts the temporal file at about 22.7 MB and the 2D file just under 10 MB - these are preview weights, not a new training or quantisation run, and they neither replace your final VAE nor produce any sound.

Reality checks

The preview shows the head of the clip. It's an honest motion check, not a quality check: not the whole action, not the final VAE's output, not audio, and not seam quality in a dual-pass run. Small-frame grids can round the aspect ratio in the JPEG you see; that's cosmetic and doesn't touch the sampled latents.

One more thing the pack is straight about: last time the author tried to prove preview on/off was numerically identical in the same process, it wasn't - the difference traced back to Core's LoRA residency/casting paths, not to the preview decoding. Fresh independent processes matched exactly. So if you're chasing bit-identical A/B results while this node is installed, run the comparison in a clean process rather than in a session that's already loaded and unloaded LoRAs.

CategoryT8/MiniMax H3/Experimental

Inputs (11)

NameTypeDefaultDescription
modelMODEL—
enabledBOOLEANtrue—
checkpointCOMBO下载:https://huggingface.co/t8star/Taeh3-Comfy;放models/vae_approx。taeh3.safetensors为时序版;2D另存taeh3_2d_kijai.safetensors,不覆盖时序版。
phaseCOMBO3 options: low, high, all
update_every_stepsINT21–1000—
min_interval_msINT5000–60000—
max_resolutionINT25664–512—
latent_prefixCOMBO3 options: 7, 2, 12
framesINT121–24—
fpsINT121–24—
jpeg_qualityINT7530–95—

Outputs (1)

NameTypeDescription
modelMODEL—