Nodes/comfyui-minimax-h3-audio-T8/FastH3 V2 · Dual MODEL4+Upscale+4 Loop (T8 EXP)
ComfyUI Node

FastH3 V2 · Dual MODEL4+Upscale+4 Loop (T8 EXP)

The 4+4 loop that tries to hide its own seam

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
FastH3 V2 · Dual MODEL4+Upscale+4 Loop (T8 EXP)
  • model_pass1
  • model_pass2
  • clip
  • video_vae
  • audio_vae
  • prompt_relay_plan
  • drive_audio
  • final_audio
  • first_frame
  • last_frame
  • persistent_identity_image
  • ref_images
  • ref_videos
  • ref_video_audios
  • ref_audios
  • source_motion
  • semantic_bridge
  • semantic_bridge_pass1
  • semantic_bridge_pass2
  • video
  • video_path
  • manifest_path
  • completed_segments
  • status
  • report_json
◄profiletrained_vsa_exp►
◄low_width512►
◄low_height288►
◄upscaler_model▾►
◄second_audio_sourceauto►
◄second_audio_strength0.00►
◄chain_idfasth3_v2_dual_4plus4_exp►
◄total_duration_seconds8.00►
◄width1024►
◄height576►
◄render_window_frames124►
◄context_frames22►
◄global_prompt►
◄segment_prompts_json►
◄prompt_relay_modedisabled►
◄query_chunk_rows256►
◄eav_modedisabled►
◄eav_tau4.00►
◄eav_start_video_progress0.15►
◄eav_end_video_progress0.90►
◄eav_max_workspace_mib32►
◄eav_g_hard_limit1.50►
◄minimum_free_vram_mib512►
◄base_seed123456789►
◄seed_policyincrement►
◄task_typeauto►
◄context_audiovideo_and_audio►
◄audio_modenative►
◄audio_denoise_strength0.35►
◄add_source_as_referencetrue►
◄prompt_primary_audio_ordinal0►
◄strict_prompt_tagstrue►
◄ref_image_sizematch►
◄reference_video_policyofficial_2_to_15s►
◄first_frame_reusesegment0_only►
◄persistent_identity_strategysingle_reference►
◄persistent_identity_interval1►
◄resume_existingtrue►
◄filename_prefixH3_In_Node_Effects_Long_Video►
◄audio_seam_policycosine_bridge►
◄bridge_ms5.0►
◄bit_depth8►
◄crf18►
◄color_matchtrue►
◄video_context_modereference_only►
◄low_context_sourceindependent_low_x0►
◄color_match_modebounded_spatial_v2►

What it is

This one node is a whole rendering pipeline. Pass 1 runs four steps at low resolution on one FastH3 V2 student. The latent result goes through the learned 3D latent upscaler. Pass 2 runs the remaining four steps at final resolution on a second student. Segments are generated serially inside the node, overlapped for continuity, then stitched and saved - with a manifest, so an interrupted run can pick up where it stopped.

If you only want one 8-second clip, use the recipe node instead. Reach for this when you want the cheap-first-then-refine shape and a resumable multi-segment render: the low pass is where the composition gets decided, and refine at final resolution is where detail gets spent. Chaining clip-to-clip by hand is the standard long-video trick everywhere (Wan workflows live on it), and it's exactly where identity drift and visible joins come from. This node is T8's attempt to do that loop with overlap and colour matching built in rather than in your head.

It's an output node, so it saves for you: video, video_path, manifest_path, completed_segments, status and report_json.

The 4+4 split, exactly

It's built on the pack's dual-model long-video node but pre-bound to the trained V2 profile - coarse 4, refine 4, and shifts 10/3 on both passes, with those six widgets removed so you can't drift back onto the old 12/3 grid. The first pass takes rungs 0 through 4, the second takes 4 through 8: no sigma reset, same ladder, same clocks. The first four steps' audio isn't finished, so second_audio_source: auto keeps the second half doing joint audio-video rather than freezing coarse audio as your final track.

And the geometry people miss: the render window is 124 frames with 22 frames of context overlap. For an 8-second render the join lands around 5.17 seconds, not a tidy halfway point. Watch and listen across that moment, not just the last frame.

Inputs you actually set

profile is trained_vsa_exp (default) or dense_compat_exp. Only the dense profile can carry Prompt Relay - prompt_relay_mode: apply_exp requires it explicitly, and the author's point is that timeline bias is never silently dropped. model_pass1 and model_pass2 are two bare full students; per-pass LoRAs go on their own branch. Do not wire the recipe node's latent-bound MODEL output here. low_width/low_height (512×288) set the first pass, width/height (1024×576) the final output, and upscaler_model picks the learned upscaler from models/latent_upscale_models. Match your reference image's aspect rather than stretching - the reviewed combo is 256×384 into 512×768 for a 2:3 still. total_duration_seconds starts at 8, and minimum_free_vram_mib (512) is rechecked before every segment; it's a start floor, not a peak guarantee.

Then the cache knobs: chain_id plus resume_existing. Caching binds to content - model, LoRA, components, code, recipe, initialisation, first-pass identity - not filenames. Change anything and you start a new chain; never drag an old chain's stage files across.

Seams and colour are the honest part

T8's own notes describe the journey: low_context_source: accepted_picture_low_context_v1 re-encodes the accepted tail of the previous segment to guide the next low pass (a short VAE encode, no extra sampling steps), which visibly improved background continuity while leaving a slight colour jump. color_match_mode: bounded_spatial_temporal_exp stabilises the first 12 frames of a continuation; the motion-colour variant fixes confident local colour outliers on top. Slight seam tinting is still listed as a known limitation, and switching either option requires a new chain_id. Leave color_match on.

Install and the traps

Same pack, same install - Manager, search MiniMax H3 Audio T8, full restart, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

You need the FastH3 student in models/diffusion_models, the latent upscaler in models/latent_upscale_models, a current ComfyUI, and ffmpeg on PATH for saving. KJNodes is not required for the dual 4+4 workflow - v1.82.0 added one specifically so it isn't.

Traps worth naming. Don't put the LowVRAM node's head_chunks=4 and ChunkFFN's chunks=2 on both branches expecting a free win: in the first V2 probe that combo measured slower and with higher occupancy (79.24s sampling versus 43.97s for the default h1/c1). Two students plus the Qwen encoder in residence eats a lot of system RAM; 16 GB of VRAM doesn't imply the rest of the machine copes. Don't stack old EMA/Turbo acceleration LoRAs, and don't migrate old workflows - nothing carries over by design. And the licence caveat applies here as everywhere in this pack: H3 and its derivatives are geofenced out of the EU, UK, Korea and the US (panel).

Last note on expectations: the accepted 8-second loops are specific samples the author's user base reviewed, not a promise about your prompts, and the pack is essentially invisible on Reddit - zero threads mention it. You're early, and docs/FAST_H3_V2_EXP.md in the repo is the actual documentation.

CategoryT8/MiniMax H3/Long Video/Experimental

Inputs (66)

NameTypeDefaultDescription
profileCOMBOtrained_vsa_exp2 options: trained_vsa_exp, dense_compat_exp
model_pass1MODEL—
model_pass2MODEL—
low_widthINT51232–16384—
low_heightINT28832–16384—
upscaler_modelCOMBO0 options:
second_audio_sourceCOMBOauto4 options: auto, legacy_policy, first_pass, highres_template
second_audio_strengthFLOAT0.000–1—
clipCLIPNative MiniMax H3 Qwen3-VL CLIP.
video_vaeVAE—
audio_vaeVAE—
chain_idSTRINGfasth3_v2_dual_4plus4_exp—
total_duration_secondsFLOAT8.00—
widthINT102432–16384—
heightINT57632–16384—
render_window_framesINT124—
context_framesCOMBO223 options: 5, 22, 39
global_promptSTRINGUsed when Prompt Relay is disabled. With Relay, leave empty or copy the Plan global prompt exactly.
segment_prompts_jsonSTRINGPrompt overrides must be empty when Prompt Relay owns the timeline.
prompt_relay_modeCOMBOdisableddisabled is exact bypass; report_only compiles/projects Relay without attention bias; apply_exp enables the projected route.
query_chunk_rowsINT25632–2048—
eav_modeCOMBOdisabledStock20 only. report_only audits CFI/g without modifying attention; apply_exp enables target-video FETA gain.
eav_tauFLOAT4.00-32–32—
eav_start_video_progressFLOAT0.150–0.99—
eav_end_video_progressFLOAT0.900.01–1—
eav_max_workspace_mibINT324–512—
eav_g_hard_limitFLOAT1.501–3—
minimum_free_vram_mibINT5120–65536Rechecked before every segment; this is a start floor, not a peak guarantee.
base_seedINT1234567890–18446744073709550000—
seed_policyCOMBOincrement3 options: increment, fixed, hash_chain_segment
task_typeCOMBOauto7 options: auto, T2VA, I2VA, FL2VA, L2VA, Ref2VA, +1
context_audioCOMBOvideo_and_audio2 options: video_and_audio, video_only
audio_modeCOMBOnative4 options: lock_source, remix_source, reference_only, native
audio_denoise_strengthFLOAT0.350–1—
add_source_as_referenceBOOLEANtrue—
prompt_primary_audio_ordinalINT00–9—
strict_prompt_tagsBOOLEANtrue—
ref_image_sizeCOMBOmatch2 options: match, max
reference_video_policyCOMBOofficial_2_to_15s2 options: official_2_to_15s, model_minimum
first_frame_reuseCOMBOsegment0_only2 options: segment0_only, persistent_identity_reference
persistent_identity_strategyCOMBOsingle_reference2 options: single_reference, scene_plus_identity
persistent_identity_intervalINT11–32—
resume_existingBOOLEANtrue—
filename_prefixSTRINGH3_In_Node_Effects_Long_Video—
audio_seam_policyCOMBOcosine_bridge2 options: cosine_bridge, none
bridge_msFLOAT5.00–50—
bit_depthCOMBO82 options: 8, 10
crfINT180–51—
prompt_relay_planoptH3_T8_PROMPT_RELAY_PLAN—
drive_audiooptAUDIO—
final_audiooptAUDIO—
first_frameoptIMAGE—
last_frameoptIMAGE—
persistent_identity_imageoptIMAGE—
ref_imagesoptCOMFY_AUTOGROW_V3—
ref_videosoptCOMFY_AUTOGROW_V3—
ref_video_audiosoptCOMFY_AUTOGROW_V3—
ref_audiosoptCOMFY_AUTOGROW_V3—
source_motionoptH3_T8_DANCE_MOTIONDance RGB motion source. Read a different source interval per segment; generated continuity is separate.
color_matchoptBOOLEANtrueMatch each continuation to the accepted RGB tail; bounded color correction only, not geometry repair.
video_context_modeoptCOMBOreference_onlyEXP: constrain high-pass overlap to the accepted final tail. Ramp mode releases three latent cells at 0.25/0.5/0.75. Audio unchanged; inspect the full continuation.
low_context_sourceoptCOMBOindependent_low_x0Accepted picture: re-encode the previous accepted movie tail for LOW video guidance only. Adds a short VAE encode, no sampling steps. New chain_id when switching. Example reviewed at 0.4MP/8s/22 context/4+4.
color_match_modeoptCOMBObounded_spatial_v2Temporal mode suppresses short RGB flicker in the first12 continuation frames. Motion Color EXP additionally corrects confident, bracketed local color outliers; requires OpenCV. No frame blending, geometry or audio changes; new chain_id required.
semantic_bridgeoptT8_SEMANTIC_BRIDGE—
semantic_bridge_pass1optT8_SEMANTIC_BRIDGE—
semantic_bridge_pass2optT8_SEMANTIC_BRIDGE—

Outputs (6)

NameTypeDescription
videoVIDEO—
video_pathSTRING—
manifest_pathSTRING—
completed_segmentsINT—
statusSTRING—
report_jsonSTRING—