FastH3 V2 · Dual MODEL4+Upscale+4 Loop (T8 EXP)
The 4+4 loop that tries to hide its own seam
- model_pass1
- model_pass2
- clip
- video_vae
- audio_vae
- prompt_relay_plan
- drive_audio
- final_audio
- first_frame
- last_frame
- persistent_identity_image
- ref_images
- ref_videos
- ref_video_audios
- ref_audios
- source_motion
- semantic_bridge
- semantic_bridge_pass1
- semantic_bridge_pass2
- video
- video_path
- manifest_path
- completed_segments
- status
- report_json
What it is
This one node is a whole rendering pipeline. Pass 1 runs four steps at low resolution on one FastH3 V2 student. The latent result goes through the learned 3D latent upscaler. Pass 2 runs the remaining four steps at final resolution on a second student. Segments are generated serially inside the node, overlapped for continuity, then stitched and saved - with a manifest, so an interrupted run can pick up where it stopped.
If you only want one 8-second clip, use the recipe node instead. Reach for this when you want the cheap-first-then-refine shape and a resumable multi-segment render: the low pass is where the composition gets decided, and refine at final resolution is where detail gets spent. Chaining clip-to-clip by hand is the standard long-video trick everywhere (Wan workflows live on it), and it's exactly where identity drift and visible joins come from. This node is T8's attempt to do that loop with overlap and colour matching built in rather than in your head.
It's an output node, so it saves for you: video, video_path, manifest_path, completed_segments, status and report_json.
The 4+4 split, exactly
It's built on the pack's dual-model long-video node but pre-bound to the trained V2 profile - coarse 4, refine 4, and shifts 10/3 on both passes, with those six widgets removed so you can't drift back onto the old 12/3 grid. The first pass takes rungs 0 through 4, the second takes 4 through 8: no sigma reset, same ladder, same clocks. The first four steps' audio isn't finished, so second_audio_source: auto keeps the second half doing joint audio-video rather than freezing coarse audio as your final track.
And the geometry people miss: the render window is 124 frames with 22 frames of context overlap. For an 8-second render the join lands around 5.17 seconds, not a tidy halfway point. Watch and listen across that moment, not just the last frame.
Inputs you actually set
profile is trained_vsa_exp (default) or dense_compat_exp. Only the dense profile can carry Prompt Relay - prompt_relay_mode: apply_exp requires it explicitly, and the author's point is that timeline bias is never silently dropped. model_pass1 and model_pass2 are two bare full students; per-pass LoRAs go on their own branch. Do not wire the recipe node's latent-bound MODEL output here. low_width/low_height (512×288) set the first pass, width/height (1024×576) the final output, and upscaler_model picks the learned upscaler from models/latent_upscale_models. Match your reference image's aspect rather than stretching - the reviewed combo is 256×384 into 512×768 for a 2:3 still. total_duration_seconds starts at 8, and minimum_free_vram_mib (512) is rechecked before every segment; it's a start floor, not a peak guarantee.
Then the cache knobs: chain_id plus resume_existing. Caching binds to content - model, LoRA, components, code, recipe, initialisation, first-pass identity - not filenames. Change anything and you start a new chain; never drag an old chain's stage files across.
Seams and colour are the honest part
T8's own notes describe the journey: low_context_source: accepted_picture_low_context_v1 re-encodes the accepted tail of the previous segment to guide the next low pass (a short VAE encode, no extra sampling steps), which visibly improved background continuity while leaving a slight colour jump. color_match_mode: bounded_spatial_temporal_exp stabilises the first 12 frames of a continuation; the motion-colour variant fixes confident local colour outliers on top. Slight seam tinting is still listed as a known limitation, and switching either option requires a new chain_id. Leave color_match on.
Install and the traps
Same pack, same install - Manager, search MiniMax H3 Audio T8, full restart, or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
You need the FastH3 student in models/diffusion_models, the latent upscaler in models/latent_upscale_models, a current ComfyUI, and ffmpeg on PATH for saving. KJNodes is not required for the dual 4+4 workflow - v1.82.0 added one specifically so it isn't.
Traps worth naming. Don't put the LowVRAM node's head_chunks=4 and ChunkFFN's chunks=2 on both branches expecting a free win: in the first V2 probe that combo measured slower and with higher occupancy (79.24s sampling versus 43.97s for the default h1/c1). Two students plus the Qwen encoder in residence eats a lot of system RAM; 16 GB of VRAM doesn't imply the rest of the machine copes. Don't stack old EMA/Turbo acceleration LoRAs, and don't migrate old workflows - nothing carries over by design. And the licence caveat applies here as everywhere in this pack: H3 and its derivatives are geofenced out of the EU, UK, Korea and the US (panel).
Last note on expectations: the accepted 8-second loops are specific samples the author's user base reviewed, not a promise about your prompts, and the pack is essentially invisible on Reddit - zero threads mention it. You're early, and docs/FAST_H3_V2_EXP.md in the repo is the actual documentation.
Inputs (66)
| Name | Type | Default | Description |
|---|---|---|---|
| profile | COMBO | trained_vsa_exp | 2 options: trained_vsa_exp, dense_compat_exp |
| model_pass1 | MODEL | — | |
| model_pass2 | MODEL | — | |
| low_width | INT | 51232–16384 | — |
| low_height | INT | 28832–16384 | — |
| upscaler_model | COMBO | 0 options: | |
| second_audio_source | COMBO | auto | 4 options: auto, legacy_policy, first_pass, highres_template |
| second_audio_strength | FLOAT | 0.000–1 | — |
| clip | CLIP | Native MiniMax H3 Qwen3-VL CLIP. | |
| video_vae | VAE | — | |
| audio_vae | VAE | — | |
| chain_id | STRING | fasth3_v2_dual_4plus4_exp | — |
| total_duration_seconds | FLOAT | 8.00 | — |
| width | INT | 102432–16384 | — |
| height | INT | 57632–16384 | — |
| render_window_frames | INT | 124 | — |
| context_frames | COMBO | 22 | 3 options: 5, 22, 39 |
| global_prompt | STRING | Used when Prompt Relay is disabled. With Relay, leave empty or copy the Plan global prompt exactly. | |
| segment_prompts_json | STRING | Prompt overrides must be empty when Prompt Relay owns the timeline. | |
| prompt_relay_mode | COMBO | disabled | disabled is exact bypass; report_only compiles/projects Relay without attention bias; apply_exp enables the projected route. |
| query_chunk_rows | INT | 25632–2048 | — |
| eav_mode | COMBO | disabled | Stock20 only. report_only audits CFI/g without modifying attention; apply_exp enables target-video FETA gain. |
| eav_tau | FLOAT | 4.00-32–32 | — |
| eav_start_video_progress | FLOAT | 0.150–0.99 | — |
| eav_end_video_progress | FLOAT | 0.900.01–1 | — |
| eav_max_workspace_mib | INT | 324–512 | — |
| eav_g_hard_limit | FLOAT | 1.501–3 | — |
| minimum_free_vram_mib | INT | 5120–65536 | Rechecked before every segment; this is a start floor, not a peak guarantee. |
| base_seed | INT | 1234567890–18446744073709550000 | — |
| seed_policy | COMBO | increment | 3 options: increment, fixed, hash_chain_segment |
| task_type | COMBO | auto | 7 options: auto, T2VA, I2VA, FL2VA, L2VA, Ref2VA, +1 |
| context_audio | COMBO | video_and_audio | 2 options: video_and_audio, video_only |
| audio_mode | COMBO | native | 4 options: lock_source, remix_source, reference_only, native |
| audio_denoise_strength | FLOAT | 0.350–1 | — |
| add_source_as_reference | BOOLEAN | true | — |
| prompt_primary_audio_ordinal | INT | 00–9 | — |
| strict_prompt_tags | BOOLEAN | true | — |
| ref_image_size | COMBO | match | 2 options: match, max |
| reference_video_policy | COMBO | official_2_to_15s | 2 options: official_2_to_15s, model_minimum |
| first_frame_reuse | COMBO | segment0_only | 2 options: segment0_only, persistent_identity_reference |
| persistent_identity_strategy | COMBO | single_reference | 2 options: single_reference, scene_plus_identity |
| persistent_identity_interval | INT | 11–32 | — |
| resume_existing | BOOLEAN | true | — |
| filename_prefix | STRING | H3_In_Node_Effects_Long_Video | — |
| audio_seam_policy | COMBO | cosine_bridge | 2 options: cosine_bridge, none |
| bridge_ms | FLOAT | 5.00–50 | — |
| bit_depth | COMBO | 8 | 2 options: 8, 10 |
| crf | INT | 180–51 | — |
| prompt_relay_planopt | H3_T8_PROMPT_RELAY_PLAN | — | |
| drive_audioopt | AUDIO | — | |
| final_audioopt | AUDIO | — | |
| first_frameopt | IMAGE | — | |
| last_frameopt | IMAGE | — | |
| persistent_identity_imageopt | IMAGE | — | |
| ref_imagesopt | COMFY_AUTOGROW_V3 | — | |
| ref_videosopt | COMFY_AUTOGROW_V3 | — | |
| ref_video_audiosopt | COMFY_AUTOGROW_V3 | — | |
| ref_audiosopt | COMFY_AUTOGROW_V3 | — | |
| source_motionopt | H3_T8_DANCE_MOTION | Dance RGB motion source. Read a different source interval per segment; generated continuity is separate. | |
| color_matchopt | BOOLEAN | true | Match each continuation to the accepted RGB tail; bounded color correction only, not geometry repair. |
| video_context_modeopt | COMBO | reference_only | EXP: constrain high-pass overlap to the accepted final tail. Ramp mode releases three latent cells at 0.25/0.5/0.75. Audio unchanged; inspect the full continuation. |
| low_context_sourceopt | COMBO | independent_low_x0 | Accepted picture: re-encode the previous accepted movie tail for LOW video guidance only. Adds a short VAE encode, no sampling steps. New chain_id when switching. Example reviewed at 0.4MP/8s/22 context/4+4. |
| color_match_modeopt | COMBO | bounded_spatial_v2 | Temporal mode suppresses short RGB flicker in the first12 continuation frames. Motion Color EXP additionally corrects confident, bracketed local color outliers; requires OpenCV. No frame blending, geometry or audio changes; new chain_id required. |
| semantic_bridgeopt | T8_SEMANTIC_BRIDGE | — | |
| semantic_bridge_pass1opt | T8_SEMANTIC_BRIDGE | — | |
| semantic_bridge_pass2opt | T8_SEMANTIC_BRIDGE | — |
Outputs (6)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| video_path | STRING | — |
| manifest_path | STRING | — |
| completed_segments | INT | — |
| status | STRING | — |
| report_json | STRING | — |