Nodes/comfyui-minimax-h3-audio-T8/FastH3 V2 · Long Video Relay Conditions + Origin (T8 EXP)
ComfyUI Node

FastH3 V2 · Long Video Relay Conditions + Origin (T8 EXP)

The relay conditioner that leaves a receipt — FastH3 V2 condition provenance

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
FastH3 V2 · Long Video Relay Conditions + Origin (T8 EXP)
  • model
  • clip
  • video_vae
  • audio_vae
  • context
  • prompt_relay_plan
  • drive_audio
  • final_audio
  • first_frame
  • last_frame
  • ref_images
  • ref_videos
  • ref_video_audios
  • ref_audios
  • persistent_identity_image
  • semantic_bridge
  • model
  • positive
  • av_latent
  • mux_audio
  • conditioned_prompt
  • media_map_json
  • report_json
  • condition_receipt
◄segment_index0►
◄context_frames0►
◄context_audiovideo_and_audio►
◄width1056►
◄height608►
◄length124►
◄task_typeauto►
◄audio_modenative►
◄audio_denoise_strength0.35►
◄add_source_as_referencetrue►
◄prompt_primary_audio_ordinal0►
◄strict_prompt_tagstrue►
◄ref_image_sizematch►
◄reference_video_policyofficial_2_to_15s►
◄execution_modereport_only►
◄query_chunk_rows256►
◄first_frame_reusesegment0_only►
◄persistent_identity_strategysingle_reference►
◄persistent_identity_interval1►

What it is

A clone of the pack's native Long Video Relay conditioner, with one addition: a typed origin receipt describing what the call actually consumed and produced - the CLIP, both VAEs, the first frame, the projected prompt plan, and the conditions that came out the other end.

Why bother, when the native node does the same conditioning? Because on H3's long-video path, the conditioner is where all the interesting decisions live: how many context frames carry over, whether audio context comes along, which reference media get attached, how the segmented prompt maps onto the timeline. When a second segment comes out wrong, or a resume doesn't match its parent, this is the call you need to be able to point at. The receipt is how a later verification node can say "yes, this is that call" rather than "these two graphs look similar."

The node's own scope note matters: this alone does not certify a HIGH handoff or an accepted parent. It's one link in the provenance chain, not the chain.

How it works

It calls the unchanged native Long Video Relay conditioner once - same payload repair, same attention routing - and wraps the result with a receipt binding the media identities, the projected plan and the output conditions into a canonical hash. Nothing about the conditioning mathematics changes; this is provenance, not a different algorithm.

That "unchanged" is the important word. Relay routing on H3 is the thing that makes segmented long video hang together, and the pack has been bitten before by compatibility layers that quietly altered the numbers. Here the claim is deliberately narrow: same conditioner, plus a receipt.

The setting that matters most

execution_mode - default report_only. The tooltip's advice is the whole workflow: run it in report-only first, read the report, and only once you're happy switch explicitly to apply_exp. Report-only is not a no-op you can ignore; it's the dry run that tells you what the real call would have done.

The rest of the required list is the normal H3 conditioning surface: model (a native H3 MODEL with no other patches on it, per the tooltip), clip (the native Qwen3-VL encoder), video_vae, audio_vae, context, prompt_relay_plan (which must come from the same segment's Long Video Window node), segment_index, context_frames, context_audio, width/height, length, task_type, audio_mode, audio_denoise_strength, add_source_as_reference, prompt_primary_audio_ordinal, strict_prompt_tags, ref_image_size, reference_video_policy, and query_chunk_rows.

Optional inputs include drive_audio, final_audio, first_frame, last_frame, the reference image/video/audio lists, first_frame_reuse, the persistent-identity trio, and semantic_bridge.

Outputs: model, positive, av_latent, mux_audio, conditioned_prompt, media_map_json, report_json, and the new condition_receipt. Wire positive and av_latent into your V2 stage; carry the receipt forward to the origin-aware attestation node.

Install

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Fully quit, restart, refresh. The pack ships no pip dependencies beyond what ComfyUI already provides, so there is nothing to install and nothing that can clobber your Torch stack. Model-side you'll need H3 in models/diffusion_models, Qwen in models/text_encoders, both VAEs in models/vae - and a recent ComfyUI with native H3 support underneath.

One thing unrelated to wiring but worth knowing before you invest a weekend in this: H3's open weights are under the MiniMax H3 Community License, which excludes the US, EU, UK and South Korea from its Applicable Territory. The hosted Hailuo API is globally available; the local weights are not, for users in those regions. That's a licence question, not a technical one, but it's the most consequential fact about running H3 locally.

Where it bites

prompt_relay_plan wants a plan projected for this segment, not a global plan straight off the Relay window logic - the tooltip is blunt that it has to come from the same segment's Long Video Window node. Pair it with the accepted-timeline projection node when you're continuing an accepted parent, or you'll get a plan built for the wrong frame clock.

And if Relay is active on a V2 stage, you need the dense compatibility profile: the trained sparse kernel has no per-query temporal bias adapter yet. Turning on Relay while running trained VSA and expecting the routing to have applied is a very easy way to be quietly wrong.

CategoryT8/MiniMax H3/Modular Sampling/Continuation Experimental

Inputs (35)

NameTypeDefaultDescription
modelMODEL未接其他补丁的原生 MiniMax H3 MODEL。
clipCLIP原生 MiniMax H3 Qwen3-VL CLIP。
video_vaeVAE—
audio_vaeVAE—
contextH3_T8_CONTEXT—
prompt_relay_planH3_T8_PROMPT_RELAY_PLAN必须来自同段 Long Video Window 节点。
segment_indexINT00–99999—
context_framesINT00–39—
context_audioCOMBOvideo_and_audio2 options: video_and_audio, video_only
widthINT105632–16384—
heightINT60832–16384—
lengthINT124—
task_typeCOMBOauto7 options: auto, T2VA, I2VA, FL2VA, L2VA, Ref2VA, +1
audio_modeCOMBOnative4 options: native, lock_source, remix_source, reference_only
audio_denoise_strengthFLOAT0.350–1—
add_source_as_referenceBOOLEANtrue—
prompt_primary_audio_ordinalINT00–9—
strict_prompt_tagsBOOLEANtrue—
ref_image_sizeCOMBOmatch2 options: match, max
reference_video_policyCOMBOofficial_2_to_15s2 options: official_2_to_15s, model_minimum
execution_modeCOMBOreport_only先 report_only;确认报告后再显式切到 apply_exp。
query_chunk_rowsINT25632–2048—
drive_audiooptAUDIO—
final_audiooptAUDIO—
first_frameoptIMAGE—
last_frameoptIMAGE—
ref_imagesoptCOMFY_AUTOGROW_V3—
ref_videosoptCOMFY_AUTOGROW_V3—
ref_video_audiosoptCOMFY_AUTOGROW_V3—
ref_audiosoptCOMFY_AUTOGROW_V3—
first_frame_reuseoptCOMBOsegment0_only2 options: segment0_only, persistent_identity_reference
persistent_identity_imageoptIMAGE—
persistent_identity_strategyoptCOMBOsingle_reference2 options: single_reference, scene_plus_identity
persistent_identity_intervaloptINT11–32—
semantic_bridgeoptT8_SEMANTIC_BRIDGE—

Outputs (8)

NameTypeDescription
modelMODEL—
positiveCONDITIONING—
av_latentLATENT—
mux_audioAUDIO—
conditioned_promptSTRING—
media_map_jsonSTRING—
report_jsonSTRING—
condition_receiptT8_FAST_H3_V2_CONDITION_RECEIPT—