FastH3 V2 · Decode + Trim Bound HIGH (T8 EXP)
Getting 68 usable frames out of a 124-frame render
- job_binding
- high_result
- video_vae
- audio_vae
- frames
- audio
- media_receipt
- report_json
What it is
The part of a video workflow nobody writes blog posts about, and the part that decides whether your continuation lands cleanly: taking a completed HIGH stage and turning it into actual frames and audio, once, with the right window.
This node calls the pack's native AV decode and output trim a single time, against an authenticated completed-HIGH result, and produces media plus a typed source receipt. It does not save a candidate and does not accept one. It hands you frames and audio for preview or saving, and media_receipt for the writers downstream.
Why "once" is in the spec: decoding a 22-frame-overlap window twice is how you end up with a two-frame drift that nobody notices until the seam. One decode, one trim, one identity.
How it works
The job_binding input is the gate. It carries the proof from the previous step - that this completed two-stage job is bound to the accepted parent it claims - and the node uses it to work out which part of the decoded AV timeline is actually yours to deliver. That's the whole 124 + 22 + 68 story in one place: segment 1 renders a window that includes 22 frames of context you already accepted, and the fresh frames are the ones past that overlap.
high_result is the completed HIGH stage result; video_vae and audio_vae are the native H3 VAEs used to decode it. Both are required, which is worth stressing: H3's audio isn't a bonus track, it's part of the same latent, and this node decodes video and audio from the same completed result so they stay in sync.
The three trim settings
start_seconds(default 0) - where in the decoded AV timeline to begin.duration_seconds(default 124/24 ≈ 5.1667) - how much to take, expressed in seconds at 24fps.fps(default 24) - the frame-rate the seconds are converted with.
All three are exposed as inputs rather than free-typed widgets, so they're meant to be driven by the graph - including the exact trim coordinates the continuation-side delivery port publishes. If you're hand-typing them, you'd better be sure the numbers match the render window you actually queued; the frame count is the thing that matters, and 124 frames at 24fps is 5.1667 seconds, not a round number you can guess.
Outputs: frames (IMAGE), audio (AUDIO), media_receipt (T8_FAST_H3_V2_CURRENT_MEDIA_RECEIPT) and report_json. Wire media_receipt into the candidate save - that's the only route that produces an accepted-able candidate.
Why the receipt matters
Once the video leaves this node it's just pixels and samples. The receipt is what carries the "where did this come from" story to the writers: which chain, which segment, which parent, which frames. The colour-match node requires it, the candidate writer requires it, and the review gate checks it against the file on disk. Skip it and you've made a nicely-encoded video the pack refuses to accept - which is fine if all you wanted was a preview.
Install
ComfyUI Manager → MiniMax H3 Audio T8, or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Quit ComfyUI fully, restart, refresh the browser. No pip step: the pack ships an intentionally empty requirements.txt so nothing can replace ComfyUI's Torch/CUDA stack, and optional EXP features pull their own dependencies only when used. You'll need a recent ComfyUI with native H3 support, H3 weights in models/diffusion_models, Qwen in models/text_encoders, and both VAEs in models/vae - plus the FastH3 V2 ConvRot INT8 checkpoint for this chain.
Common issues
A missing job_binding is the most frequent stop: you can't run this node before the binding step, and a binding from a different segment will refuse. The second is VAE placement - if decode throws a shape or channel error, it's almost always the wrong VAE or the native ones not being where the pack expects.
And the honest caveat from the pack's own documentation: dual-pass H3 continuation still has a known slight seam colour change. This node gives you a clean, exactly-trimmed delivery of 68 fresh frames from a 124-frame window. It doesn't make the model invent a seamless join, and if you're expecting that, the colour-match node downstream is where the pack tries to address it - with a real algorithm, and a receipt saying what it did.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| job_binding | T8_FAST_H3_V2_CURRENT_JOB_BINDING | — | |
| high_result | T8_STAGE_RESULT | — | |
| video_vae | VAE | — | |
| audio_vae | VAE | — | |
| start_seconds | FLOAT | 0.00 | — |
| duration_seconds | FLOAT | 5.17 | — |
| fps | FLOAT | 24.00 | — |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| frames | IMAGE | — |
| audio | AUDIO | — |
| media_receipt | T8_FAST_H3_V2_CURRENT_MEDIA_RECEIPT | — |
| report_json | STRING | — |