comfyui-minimax-h3-audio-T8
A ComfyUI extension with 549 custom nodes.
Nodes (549)
The H3 second-pass sampler that quietly fixes your audio
The H3 'memory saver' that honestly tells you it might save nothing
When other packs' attention patch nodes won't talk to H3, this shim translates
The one conditioning node that runs every MiniMax H3 video task
An audio QC pass that never touches the audio
Lock your voice into the latent so H3 doesn't reinvent it
Blend source and generated audio without leaving the graph
Kill it without touching the timing
Catching the 'voice suddenly sounds far away' failure
The bouncer for H3's audio-refine chain — check everything before you burn a run
The audio-refine middleman that builds the 4-step tail and won't overclaim it
It needs to know exactly what you already ran
The node that hands SamplerCustomAdvanced a ready-made 4-step audio tail
This H3 Audio Refine Audit Refuses To Approve Your Candidate — That's The Point
The Gatekeeper Before Your H3 Audio Resample
One Boolean That Stops Your Long H3 Video Saving Context It Can't Back Up
Turns the refine plan into noise, guider, sampler and sigmas
Picking The Right H3 Audio Refine Audit
The Flexible End Of H3 Audio Refine
The strict sibling of the dual-clock setup
The dual_model Half Of H3's Audio Refine Audits
Handing H3's Audio Refine To A Second Model
Catching A Wrong-Length H3 Checkpoint Before It Costs You A Render
Resume H3 Audio Refine From A Checkpoint Someone Can Actually Prove
The split that stops your next long-video segment from inheriting yesterday's audio fix
When your audio-refine pass needs a different model, this router validates the pairing
A strict four-step tail, and only two denoise points allowed
A signed, deterministic plan for re-sampling just H3's audio tail
Your refine only counts if you actually listened
The Last Automated Gate Before Your Refined H3 Audio Tail Ships
Is that your recording or generated audio? This node answers it
Trim your source audio to exactly the grid H3 wants
Your H3 Avatar Video Should Ship The Original Recording, Not A Model's Impression Of It
Handing H3 Avatar Off To HIGH Without Touching The Audio Anchor
H3 Avatar's Cheap Step, With The Recording Bound
Make a portrait talk with a recording you already have — no cloning, no retraining
H3 Avatar's Smallest And Most Important Node
Check the decode's VRAM bill before the VAE ever runs
Turning the joint AV latent back into frames and sound
The two-plug adapter that welds MiniMax H3's picture and sound latents together
Split the joint latent without paying for a decode
Reshape the sigma schedule without changing the step count
Spend extra steps exactly where the tail ends
One extra step at the very end of H3 sampling
CADS for H3 references — noise on the picture, hands off the audio
The Learned 3D Lift Inside A Chunked H3 PASS2 Graph
Attach One Stage EAV Config To Exactly One PASS2 Segment (Or Don't Bother)
Auditing EAV's Actual Calls On A Chunked H3 PASS2
Preparing PASS2 State For Chunked H3
Who Counts The Relay Calls When EAV Is Also In The Graph?
Pairing An External Prompt Relay With One Full-Clip H3 PASS2
Refine one H3 time segment at a time — and actually mean it
Slice one first-pass segment without sampling anything
The v2 plan that fixed chunked two-pass H3
H3 low-sigma two-pass
Mask-preserving low-sigma refinement for H3
The planner for H3's chunked two-pass upscale — and why full-frame is the safe default
Run H3's chunked two-pass upscale, and never touch the audio tensor
Did your prompt-relay actually run? This node answers, for one segment
Projecting a whole-clip prompt timeline onto one chunk without breaking it
Resuming a chunked H3 render from a SHA, not from faith
Freezing one chunk so you never resample it again
Bolt a video-enhancement attention pass onto one H3 window
Counting the EAV calls your window actually made
One upscale pass for the whole clip, then never touch it again
The node that actually renders your chunked H3 windows
One noise draw for the whole clip — the v5 windows all share it
The call counter for per-window prompt relay
Your 141-frame prompt timeline can't fit a 136-frame window — fix that here
Proving your partial4 source is really a finished native checkpoint
Resuming window 2 without resampling window 1
Checkpointing one refined window so a rerun doesn't cost you the clip
The tiny patch you reach for when 16GB is barely enough
Check your ClipProj setup without loading a single matrix
The deterministic half of H3's scene-understanding pipeline
A reference-semantics node that defaults to fully offline
The continuation node where you actually write the prompt — twice
Turning the last frames of your accepted clip into motion guides
Handing off a finished continuation without pretending to be an editor
Applying the video-enhancement pass to your continuation's HIGH phase only
Where the upscaled latent becomes something a sampler can finish
The second half of the segment, with its own model and its own prompts
The one-continuation detail boost, applied to half a segment
Run the cheap half of a long-video segment, nothing else
The node that decides which 4 of your 8 steps run cheap
Give the model a sense of time without touching the other half
Point your prompt timeline at the right second of a continued video
The node that stops your long video from continuing the wrong clip
Quarantine candidate files instead
Picking the next shot from what the background job finished
Bind a Creator plan to the Long Video background controller
Where were we? The resume node that never queues anything
A delete list that refuses to delete anything
Log what actually happened, so the plan can trust it
A cache *planner* for long H3 shoots — tells you what's cached without touching a file
The shot card you can edit without rebuilding the timeline
Side-by-side A/B review without the guesswork
Pick the shot and seed you're about to burn VRAM on
'render shots 3–7, seeds 0–2, and don't touch my files'
H3 Dance Motion Source
Four quality knobs that all start at zero
'Did that line actually play here?' — finding speech boundaries locally
A dialogue mix that refuses to eat your words
From a two-column script to a dialogue plan
Rerun one line, not the whole scene
Obsidian Director D1 is a storyboard tool, not a generator — and that's the point
Your H3 clip is 24fps and the motion looks juddery — this double it with DLSS, not RIFE
NVIDIA's denoiser-plus-upscaler as a ComfyUI post node
Why nothing in this pack runs before the audit says READY
The file-streaming DLSS-NR node
One persistent DLSS process across all your frames
An AYS schedule that knows it isn't an image-model AYS schedule
One clock for video, one for audio
The 4+4 long-video node
Turn seconds into valid H3 frame counts (22, 124, 362…)
CFG on H3 is expensive — this guider makes you say yes first
How many forwards did that guider actually burn?
Enhance-A-Video's receipts, checked after sampling
One measurement, not fifty
Enhance-A-Video that survives segment boundaries
When Prompt Relay and Enhance-A-Video want the same attention slot
Enhance the target, leave the reference alone
Enhance-A-Video plus SageAttention without the fallback roulette
FETA plus skip-block guidance on H3, with the extra forwards audited
Temporal sharpening for H3 clips, injected straight into attention
Diagnose your H3 install before you blame the workflow
Feed BlockSwap settings to the external H3 streaming sampler
Assemble a 2–3 person cast from your character profiles
Turn reference photos into a reusable character profile
When the only prompts that matter belong to the repaired window
Proving the per-frame repair belongs to the clip you just posted
Per-frame denoise for faces that only drift in some frames
Swap the face crops into your H3 latent, audio untouched
The manual-verification gate that checks you built it the author's way
The node that forces a human to look before a face fix lands
Inject the parity crop latent, keep the audio exactly as-is
Replicate the reference face-refine recipe, exactly
The reference-mechanism stitch, bit-exact outside the mask
Denoise that scales with face size — 30px faces get more help
Plan the face-repair crops before you touch the sampler
An automatic 'don't ship the worse face' check before export
The H3 Face Refine v1.1.1 Sampler-Mask Patch
The low-denoise dual-clock sampler for the face-refine pass
Paste the refined face back — with a bit-exact audit of what you touched
Carve one legal H3 window out of your video — context and all
Tell H3 exactly which frames are ugly — it plans the repair window
The atomic 'yes' that makes a face fix durable
Rebuild the final video from a manifest — H3 never even loads
Start a crash-safe face-fix run across many windows
Keep your timestamped prompts while a face pass runs on top
The node that checks your repair is still repairing the right thing
Give the face repair its own sampler instead of hijacking the whole video pass
Label every face track with the right character — or override by hand
Read this before you decode and review a repaired face crop
Repairing one face in one stretch of frames, without re-rendering the film
FastH3's 4-step contract — apply the LoRA first, then let this wire the sampler
Hand an accepted clip's tail to a distilled 8-step pipeline
The trim maths nobody wants to do by hand
Lock the old segment in before the distilled tail runs
Picking which half's audio survives a 4+4 continuation
Sliding a Relay plan onto the accepted frame clock
90 frames or 124? The FastH3 V2 remainder window, and the 34 frames it throws away
Writing the colour-matched version without destroying the original
Pulling a finished LOW pass out of FastH3 V2
The relay conditioner that leaves a receipt — FastH3 V2 condition provenance
Re-checking both accepted segments, read-only
The explicit accept step for FastH3 V2 segments
The FastH3 V2 candidate writer
Proving LOW → lift → HIGH actually happened in that order
Hanging today's render off an old accepted parent — carefully
Getting 68 usable frames out of a 124-frame render
A fingerprint of the graph as it is right now — FastH3 V2's Current Recipe
Prove which sampler actually ran — FastH3 V2 stage attestation
The 4+4 loop that tries to hide its own seam
Matching segment two's colour to the accepted parent
No LOW model, no LOW sampler, no guessing
Freeze the LOW half so tomorrow's HIGH pass has nothing to guess
Same as the stage attest — but now it also certifies where the conditioning came from
Did the sparse attention actually run? The 30-second check people skip
FastH3 V2 is not a LoRA — it's the 8-step recipe node for a different student model
One DMD stage, no hidden loop behind the curtain
The learned 3D lift with a paper trail
H3's turnaround sheet, sampled in a single pass
Why you can't just VAE-decode a character sheet
The T8 FlashVSR strategy node
What it actually checks (and what it won't fake)
2× or 4×, audio untouched
Stop H3 from round-tripping to the CPU every step
FreeNoise for long H3 videos — one shared noise pool, permuted per segment
Steer an H3 shot with a depth, pose or edge video — structure in, motion out
Load the ControlNet that steers MiniMax H3's video with depth, pose or edges
Does your ComfyUI actually have Generic Loops?
The tiled-VAE fix that H3 tested and rejected — now a read-only audit
Bolting EAV onto a single H16 PASS2 window
Your effect node's counter is lying to you — this one counts the real forwards
Split your H3 second pass into windows before you burn 40 minutes on one 124-frame run
One H3 window, five things you can edit — the PASS2 node that stops hiding the sampler
Proof that your prompt actually changed mid-clip — the Relay audit node
Prompt Relay was built for a whole clip. This squeezes it into one H3 window.
Start an H3 window chain from a checkpoint — but only if the SHA checks out
Resume your H3 clip at window 4 without re-sampling the first four
Freeze one H3 window so you never sample those 34 frames again
Build a tiny content-addressed model patch, not a whole copy
Quarantine, restore and repair your hybrid model patches safely
Your Hybrid setup is probably broken somewhere — this node finds it before you burn a render
Load a clean quality base, then bolt one small hybrid artifact onto a clone — in that order
Two pruned checkpoints walk into a render — check their SHAs before you trust the blend
Feed 4 steps, then upscale
Which socket feeds the upscaler? This tiny node exists so you stop guessing
Not a LoRA loader — the node that keeps your HyperFlow adapter honest
The pass-through node that tells you what your H3 stage actually executed
Load a frozen H3 HyperFlow stage by SHA
Full8, partial4, or the HIGH half — one node that sets up an H3 HyperFlow stage
Did the effect actually run on your HEAD half? This node reads the record
Bind an attention effect to the first half of a continuous H3 run
Resume a continuous H3 run at the halfway point — and delete the first half from your graph
The first half of the trajectory, and why it isn't just half the sigmas
Save the raw x_sigma, not the pretty version — freezing a HyperFlow HEAD
The HyperFlow HEAD stage
Latent parity between a split and a single-pass H3 run
632 tensors of real adapter, and it won't fake it for you
Where the seam lands, and how resume works
Continue an 8-second H3 film from the segment you actually approved
Nothing continues until you press accept — reviewing a P7 long-video segment
The node that writes your segment — and still refuses to call it final
The P7 node you instantiate twice per segment
Where the upscaled picture meets the HIGH audio policy
Resume a finished HIGH without spending a single diffusion step
The bouncer at the end of the HIGH pass
Four hours of sampling, saved as a path and a hash
The HIGH pass, wired but not fired
Segment zero is the easy one, and this node keeps it that way
The upscaler that isn't an upscaler, and why audio survives it
Reuse a frozen LOW pass — and re-prove it still belongs to this chain
Pick the right socket and your HIGH pass gets better for free
Bank the cheap half so you can re-roll the expensive one
4 clock
Segment two needs a memory of segment one — here's exactly what it takes
Relay routing only counts if the sampler actually did it
Projecting Relay onto the right frames
After setup, and only if you opt in
The cheap upscale route, in eight evaluations
The plan is bound to the model, not the schedule
8 + 4 evaluations, and no pretending it's seamless
The 8-NFE path, wired into a normal SamplerCustomAdvanced
Native 8-step HyperFlow at full size
Split the 8-step HyperFlow run across two models without breaking the trajectory
Ask the result, not the node, whether your effect actually ran
Wire the second half of a split generation — without asking for new noise
Decode a finished clip on a machine with no H3 model loaded at all
The re-noised second half that finishes an upscale
A save node that refuses to overwrite anything — and that's the point
The second half of a HyperFlow split, with no fresh noise
Two characters, one render, zero identity guarantees — the honest multi-speaker dialogue experiment
Pin a middle keyframe, pick a spot on the timeline, and let H3 interpolate around it
Stitch the repair back onto the original — repaired pixels in, untouched sound everywhere else
Preparing the mask before the sampler
Upscale your H3 latent without wrecking the aspect ratio
Real detail on the video, native audio left untouched
The sigma plan that makes 4+4 actually work
This audit fails unless every H3 forward really used the sparse path
When SLA and KJ Sage both want the attention path, this composer picks one owner per call
4-step H3 with 85% of the attention blocks skipped — the LightX2V SLA path, properly pinned
The renderer that plays every scene in order
The V2 vocal-lock renderer
The renderer behind the accepted 32-second MV
The review gate that keeps your long video honest
Load the accepted previous segment and the exact parent identity your next render depends on
The terminal node that accepts, composes, and queues exactly one next segment
Register the prompt before you sample, so a crash or OOM doesn't nuke the whole long video run
Save a segment without touching your accepted history — the candidate half of long video
Killing the color seams in a multi-segment H3 long video
Verify every accepted segment, then stream them into one MP4 without blowing up memory
The conditioning that turns 'one more segment' into 'a continuation' — context overlap done right
The node that gives your long H3 video a memory of the last segment
Save the tail of this H3 segment, and nothing else — that's the job
Prompt Relay timelines plus Enhance-A-Video
H3 long video that renders every segment in a single execution
Type a total duration, and this node decides every H3 segment for you
MiniMax H3 doesn't do 60-second clips, so it plans them in pieces
Refine the tail of every long-video segment without rebuilding the loop
The audit that catches H3 segments going green (or dark) at the seams
Keep the right voice on the right character across long H3 segments
The first-shot approval door for voice-plan releases
An H3 LoRA loader that stops being picky about where the LoRA came from
Head-grouping to shave H3's attention peak
Converting a latent without touching a single pixel
The audit that stops you lying to yourself about the LTX refiner
An adapter boundary, not an upscaler — the H3→LTX learned handoff
It wants the sampler's receipts, not your word
A sampler that keeps score of whether it actually sampled
Your H3 audio survives the LTX pass — here's how to get it back
Checking the RGB→LTX handoff, and keeping the original audio honest
Frames you already have, audio you keep
Two passes you wire yourself, exactly like the old runner did
A real geometry preview before you spend a render on it
Meridian's First Node Doesn't Load a Model — It Audits Your Files
Three Forwards, No CFG, and a Full Decode Before You Get the File
One Image or One Video. Wiring Both Gets You Nowhere.
Sneak a tiny time-bias into H3's sigma — no extra steps, no noise, just bias
The lazy gate that skips the whole repair graph when H3 got it right
Loading a first pass you can actually prove is the first pass
The read-only motion analyzer that decides if H3 botched the movement
A dependency-free audit that flags shaky, frozen or mushy H3 footage
Put the repaired clip back on the exact world clock — audio untouched by default
The composer that turns your stretched latent into a controlled second pass
Pairing Prompt Relay with a retimed second pass — without fooling yourself
The one-number node that keeps Prompt Relay honest
Turn an audit's flagged ranges into a repair plan, without touching the media
Stretch the bad frames before H3 re-samples them — that's the recovery trick
Slice one repair plan into VRAM-sized windows instead of one giant pass
Prove the repaired take is the take you asked for (and don't ship the wrong audio)
Bind the repair pass, then let it touch nothing else
Assemble the repaired windows, and rerun only the one that died
Stitch each repaired face back onto the untouched original
One character, one shot, one repair job — H3's answer to multi-person face fixes
Before you stitch that repaired face in, check whose face it is
One node per character. Yes, really — that's how multi-face repair works
First frame, last frame, and the keyframes in between — one conditioning node
Different step counts for video and audio — with one honest caveat about cost
The prompt node that refuses to invent lyrics
The six sections MiniMax actually expects, compiled by hand — on purpose
The scene planner that needs two audio tracks — and why that's the point
Directing each shot like it's a contract
The CPU-only storyboard planner
The audio
One stage of the dual recipe at a time, without the loop hiding under you
Reload a finished H3 latent, and let it prove nothing got swapped on you
Save a finished H3 latent as a tamper-checked file — and it won't overwrite anything
Splicing H3 Long Video segments by removing the exact 5/22/39-frame head
Proving your reloaded H3 latent is bit-for-bit the one you saved
Joining two H3 clips in latent space so you don't decode doubled frames
Hard-lock the previous segment's native latent tail
The cheap way to give a graph you already trust a stage identity
Checkpointing H3 mid-render at exact step boundaries, then resuming in a fresh process
The immutable JSON receipt that makes NFE resume safe to trust
Which Copy of That Node Is Actually Running?
The 'is my H3 setup actually broken' node (and why 'unknown' is an answer)
Trimming an H3 render by seconds, video and audio together
The PDD setup node that isn't an ordinary LoRA loader
Split a Distilled 8-Step Run in Half
The five-second check that stops an H3 render before it OOMs at step 40
Reading a Prepared Tao/LTX Input Bundle
Where the Split Chain Finally Becomes a File
The Sampler Half of a Split LTX Refinement
Reattach to a Run After ComfyUI Died
H3's Serial Video Node
Effects on the Full-Size Pass, Anchored to the Clean Source
Prove the Second Pass Ran What You Asked For
Get the Second Pass's Noise Right, or Watch It Drift
Decode Last Night's Render Without Starting a Sampler
Write the Finished Latent With Its Receipt Attached
The Euler Tail, With Nothing Hidden Behind It
Hand Your Half-Denoised Latent to Your Own Upscaler
An entire long-video pipeline hiding in one node
Effects on the First Pass Only
Did Your EAV Actually Fire, or Just Sit There?
Build the Graph That Never Touches LOW Again
Bank the Expensive Half and Never Re-Run It
The First Pass, and Why It Won't Hand You a Video
Prompt Relay, Packed Per Stage
The H3 sampler that skipped 37% of my render time
One clock, two passes, zero generation
Prep a Prompt for Exactly One Stage
The Node That Decides Where Your First Pass Ends
The 7000-character reality check your H3 prompt needs before you queue it
One structured brief that compiles for H3, Wan 2.2, or LTX-Video
From the Studio compiler's packet to a relay timeline — without rewriting a word
Prompt enhancement with zero network by default — and guarded keys if you opt in
The node that actually installs the relay — encode the plan, patch a cloned model
One action per node — chain these to build a timed H3 scene
Segment-level relay conditioning that keeps Long Video's motion context intact
Keeping relay events from restarting at every Long Video segment
One global prompt, several timed local actions — the heart of the relay system
See every event's frame range before a single model loads
The switch that extends relay routing from video to audio — paper-faithful by default
Guess H3 packed-row counts before you commit — and don't call it a VRAM certificate
Stop Re-Sending the Scene You Already Shot
The local 8B prompt rewriter that runs on 16GB and unloads after every use
The VRAM release button for MiniMax H3's 8B prompt rewriter
Stop your prompt rewriter from quietly changing the scene
Get your <Picture 1> / <Video 1> / <Audio 1> tags right before H3 eats them
See whether your H3 Qwen prefix cache is actually saving you anything
Cache H3's expensive vision prefix so repeated references don't re-encode
Carry one reviewed mask through the shot — and make it stop at the cut
A read-only RAFT audit that tells you if your H3 shot actually moves
A loader that refuses to load — until it's sure the machine can survive
T2VA-only, or nothing
One knob panel for the whole RAVEN streaming chain
RealBasicVSR temporal restore — cleanup after H3, with audio treated as untouchable
Re-noising H3's clean output and descending again
Render the reel to MP4 without ever holding it all in RAM
Plan the whole reel before you render a single frame of it
Pick which segment you're going to redo, without touching the rest
Who decides how much VRAM H3 gets to hog? A node that plans, but doesn't touch
Set Up the First Pass So You Can Actually Freeze It
This node does nothing — and the RF restart is broken if you skip it
The 3-step second pass that keeps your audio intact
A save node that refuses to write a corrupt H3 MP4
Track two or three people through a shot with SAM3.1 — colors are not identity
Inject your drive audio into H3's denoise window, on schedule
The 'yes, I actually reviewed this' handshake in the repair chain
Lock a repair segment to an immutable manifest revision
Render the repaired timeline without ever rebuilding the original
Preview a repair candidate without committing a single byte
Build the do-over list for your long video, non-destructively
The node with two inputs that isn't optional
A 5120→5120 conditioning adapter, not a LoRA and not a speedup
A non-generative skin finish that never touches a pixel you didn't sign off
Dichromatic specular attenuation
The frequency-split rebuild
The node that doesn't frontalize faces
Per-person skin masks in a two-person shot, without rerunning SAM
Finish two people's skin in a long clip without loading SAM twice
The node that executes the cast sheet, track by track
The little node that builds the per-person skin-fix plan
Actually look at a skin-finish candidate before you accept it
The two-pass quality stream
A fail-closed audit that isn't a beauty score
A skin mask that knows eyes from cheeks, so your finish doesn't smear
The frequency split that cares about highlights (specular-aware, not a replacement)
The guided-filter skin finisher that refuses to eat the texture
The honest way to 'finish' H3 skin without a face-fix trip
A mechanical guard so your skin finish can't flatten shadows or clip highlights
One keyframe at a time
Keyframed skin finish, per person, per shot
The only Skin Finish node that re-encodes your video without touching the audio
Skin finish on a long H3 clip without ever materializing all the frames
The SLA LoRA loader that refuses to re-quantize your FP8 base
The audit node that calls the render a failure when the math didn't happen
The precision-first H3 attention patch
Before you bolt Sol-Attn onto H3, let this node check the patch ownership
Get your H3 draft ready for the LTX-2.5 refiner — audio stays out
The low-sigma LTX refiner for when the official route changes the face too much
The official LTX-2.5 3-step refiner — apply the 0.8 LoRA first, then wire the sampler
The fast decode that ends H3 Super — TAEHV wide, not the full LTX VAE
The TAEHV encoder you should mostly not use — legacy round-trip, diagnostic only
The one-input loader that hands H3 Super its fast codec
Plan dialogue, music, ambience and SFX on one H3 sound timeline
Assemble video and audio latents for a redraw, with denoise masks H3 understands
Cut source video into the exact H3 window the VAE can encode
STG for H3, tuned for its joint AV transformer and its conflict-prone patches
Catch ambiguous multi-speaker dialogue routing before you burn a render
Make a voiceover land on the exact frame count the scene needs
Stitch rendered speech turns into one track and get the SRT for free
The node that turns 'speak this line in this voice' into H3 conditioning
Pull just the audio out of an H3 latent, no video decode
The bookend that releases your VRAM after a speech render
A seatbelt for H3 speech renders that crash or get cancelled
Commit a good H3 segment before the next render — and don't lose it when ComfyUI dies
Sew your accepted H3 segments into one clean audio track — subtitles included
Cancel, reset, or just check on a long-form speech job — without touching the graph
Render one segment at a time, resume after a crash
Emotion, pace, pitch — with the honest caveat that none of it is calibrated
Turn a script into an H3 voice performance — without guessing how long it'll take
Condition, sample, decode, release
ASR verification and exact-target trimming
A VRAM reality check before the 33B speech sampler runs — not a guarantee, a gate
Strict 24fps, 17n+5, no stretching
The DCT trick that hands one SPEED stage to the next (and its own audio clock)
Noise that doesn't change the audio when you change the video canvas
Prove which H3 checkpoint and VAE you actually ran — with a SHA-256, not a filename
SPEED's plan, without the hand-waving
Prompt Relay, one SPEED stage at a time — and it does nothing by default
Progressive-resolution H3 with audio riding along
Hold the raw H3 inputs so every SPEED resolution stage can re-encode its own references
Accumulate H3 spectrum statistics across runs without hoarding latents in VRAM
Persist your accumulated H3 spectrum dataset — atomic save, no silent overwrites
Turn a hundred accumulated H3 clips into a spectrum profile — only if the fit earns it
Measure H3's own latent spectrum instead of borrowing WAN or Flux constants
Resume a SPEED clip from disk — path, SHA, and nothing else
One sampling call, no loops, and the output you actually need for the next stage
Freeze one SPEED stage to disk (and keep the SHA — you will need it)
Canvas, conditioning, schedule — but no sampling
Applying Enhance-A-Video to one H3 stage without breaking your attention backend
The audit node that tells you your effect didn't actually run
The tau that does nothing at 0, and the window that isn't per-stage
Loading a frozen H3 stage from disk
Noise as an explicit, inspectable node
One sampler call, and a receipt that survives the restart
The path and SHA are the point
Load the second model after the first stage finishes — not alongside it
Tell it what to change, not which pixels
Steal a Still From an H3 Latent Without Decoding the Whole Clip
Check Your H3 Still Contract Before You Burn a 15-Second Generate on It
Pick One Shot Out of Your Timeline and Render Just That
Plan a Whole H3 Video as Shots, Not One Giant Prompt
Subject-safe RGB compositing done honestly
Why your H3 progress preview is one frame — and how to check if it could be better
Watch the Motion Mid-Sample, Then Kill the Take You Hate
Sharpen H3 Footage Without Making the Motion Flicker
Lock a Music Bed Into an H3 Clip Past the Dialogue
Pin a reference image to an exact second — without spending an H3 slot
Give H3 a whole video of motion cues, timed to the frame
The Topaz node that stops you guessing
The Official Topaz Interpolation Node
Topaz enhancement inside ComfyUI — and why it writes a lossless master, not an MP4
Finish the Second Half of an H3 Sample You Started Elsewhere
Save Your Half-Finished H3 Latent — Only When You Say So
Plot an object's path across the clip — without touching attention
Turn your trajectory plan into conditioning H3 Fun Control can actually eat
Split an H3 Sample in Half So You Can Resume It Later
Before you burn an hour compiling TensorRT engines, run this 5-second check
12GB free VRAM, 24GB RAM, and a long coffee break
TensorRT VAE decode — the recommended TRT node, and the benchmark that says 'no total win'
The full TRT VAE (encode + decode), and why its own author tells you not to use it
A research paper's temporal fix as an optional switch
The router that refuses to let one pretend to be the other
A Guardrail That Stops Pass 2 From Silently Rewriting Your Audio
Tail, Bias, STG, Restart
Stitch the Two H3 Passes Together Without Breaking the Audio Clock
The 4+4 Sigma Split That Doubles as an H3 Hires Fix
One Character Sheet for Your Whole H3 Video
The plan node that won't let you misstep
Grafting a distilled branch onto your H3 model
The node that plans VDN's second pass so you don't have to
Real temporal bias on a VDN stage — and no extra NFE for it
Did the VDN relay actually fire? Read this, not the wiring
A 4.28GB audit that never actually loads the 4.28GB branch
Setting up one VDN stage without hidden sampling
Fix the Sound, Leave the Video Alone
Generate one window, look at it, then decide
The two-second check that saves you an hour of sampling
Save the outpaint — and it will refuse if the first frame moved
Decode, paste back, and hand you an MP4 that actually plays
Sample the other 21 windows under the settings you approved
See the crop before you spend an hour generating it
Per-side prompts for the new canvas — and an honest person-box audit
Reopen a candidate you generated yesterday, without re-sampling it
Sampling finished, the save crashed. Don't re-run it.
Reconnect to a prepared run without re-encoding a single frame
Get your confirmed candidate back after a restart — by ID alone
Decide the new canvas before you burn an hour of GPU
The regional-prompt version of Prepare — keep it on its own run name
Encode the source, audio and prompt once — then never again
Binding region prompts onto the H3 model — and what it refuses to bind
Every window in one queue, resumable if your PC dies
The one-click approval gate — and why it defaults to off
The One-Knob Fix for H3's Over-Smoothed Reference Textures
Delete a Voice the Way You Should — Into a Recoverable Trash Folder
Pull a Saved Voice Back Into Any Workflow
Save a Voice You Actually Want to Reuse (Explicitly)
Give H3 a Voice It Can Remember — Describe It or Clone It
A Reserve Policy That Actually Runs Before Load
Scripting H3-World's 37-step action timeline
One scene prompt, 37 bound actions
The H3-World LoRA loader that also installs the secret runtime
The H3-World save node that refuses to write a corrupt MP4
Block-sparse attention, and the three knobs that matter
MiniMax H3 Audio T8
用于 ComfyUI 的 MiniMax H3 视频与声音节点:参考图/音频、双采与长视频、人物口型、运镜编辑,以及可选高清后处理。
简体中文 | English · 当前版本:1.92.0 · 更新日志
安装
需要较新的 ComfyUI 原生 H3 支持。手动安装:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
安装或更新后,完全退出并重启 ComfyUI,再刷新网页。Manager 可搜索 MiniMax H3 Audio T8;Registry 与 GitHub 发布进度独立,未显示最新版时使用 GitHub 安装。
选一个工作流开始
先替换自己的模型和素材,不要直接运行示例里的占位文件。
|你想做什么|入口|
|---|---|
|第一次跑 H3/首帧图生视频|基础生成|
|音频参考、人物说话/唱歌|音频控制 · 原生音色/情绪与 Avatar|
|双模型 4+4、分段长视频|长视频|
|独立一采/二采、外置 EAV/Relay|分离采样 · Veda 8 步|
|FreeVideo FP8:8+2/真正4+4(EXP)|六张分离工作流 · 独立环境|
|Kijai 九原件、Acc8 Full/Pruned 分离双采|18张工作流 · 模型配对|
|外片续拍、可见脸 MASK、无脸旁路及曲线双采|RADAR 新工作流与限制|
|可视化编排镜头、素材和声音|曜石导演台(也可点 ComfyUI 左侧独立的 T8 曜石导演台)|
|OpenVDN、FastH3、低显存|加速工作流 · 低显存说明|
|Meridian 图片/视频运镜、源时间编辑|四节点工作流|
|一采动态预览、定向取消|TAEH3 预览|
|成片高清放大/插帧|Topaz · DLSS-NR · DLSS 插帧|
|其他功能与详细操作|完整工作流目录 · 详细使用说明|
模型下载与位置
下载前阅读对应模型仓库的 README 和许可。不要把不同版本同名权重互相覆盖。
|模型/资源|下载与 ComfyUI 目录|
|---|---|
|H3 主模型、Qwen、视频/音频 VAE|分别放 models/diffusion_models、models/text_encoders、models/vae;基础安装说明|
|TAEH3 时序/2D 预览模型|Taeh3-Comfy → models/vae_approx/;2D 文件另名保存|
|Meridian ConvRot INT8 + Omega 1B512|Meridian-Comfy → models/meridian/;Omega 在 vggt-omega/checkpoints/,源码/assets 和 H3 VAE 另备|
|Semantic Bridge|Semantic-Bridge-Comfy → models/semantic_bridge/t8_compat/|
|OpenVDN 完整包|Vdn-Minimax-H3-Comfy;保持仓库目录结构|
|H3-World 动作 LoRA|Minimax-H3-World-Comfy;接线说明|
Meridian 的 DMD 已合并,不要再加载同一 DMD LoRA。Omega 是官方原始 PT 几何模型,不是 INT8 或普通 VGGT。TAEH3 只做近似预览,不替换最终 VAE,不生成声音。
使用注意
Advanced/EXP使用对应工作流;指定样片通过不代表所有素材、显卡或长窗口都能得到相同效果。- 旧工作流不强制迁移;未验证的 LoRA/Sage/Sol 组合只提示边界,不一刀切禁止连接。见补丁共存策略。
- 不保证精确声纹、逐字时序、无接缝或通用 16GB 性能。双采仍有轻微接缝变色的已知限制。
- Topaz/DLSS 是独立后处理;各自的程序、资源和运行条件见对应文档,不能只下载 H3 权重。
节点变红、声音异常或模型不兼容时,先看常见问题与高级配置。反馈请附完整报错、工作流与 Core/节点版本:提交 Issue,移除密钥和隐私素材。
T8 链接
|平台|入口| |---|---| |B站/YouTube|B站 · YouTube| |API/在线 AI 应用|API(推广链接) · RunningHub(邀请链接)| |整合包/模型网盘|ComfyUI 整合包 · 模型网盘| |模型/节点|Hugging Face · 节点 GitHub|
许可与文档
节点源码:GPL-3.0-or-later。本 GitHub 不含模型权重;模型及外部资源各自遵守原许可。Meridian 衍生权重遵守 MiniMax H3 Community License,Omega 单独遵守 FAIR 非商用研究许可;转换不改变这些边界。