Nodes/comfyui-minimax-h3-audio-T8/H3 Continuation · ONE Phase Conditions (T8 EXP)
ComfyUI Node

H3 Continuation · ONE Phase Conditions (T8 EXP)

The continuation node where you actually write the prompt — twice

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Continuation · ONE Phase Conditions (T8 EXP)
  • contexts
  • clip
  • video_vae
  • audio_vae
  • drive_audio
  • final_audio
  • first_frame
  • last_frame
  • ref_images
  • ref_videos
  • ref_video_audios
  • ref_audios
  • persistent_identity_image
  • semantic_bridge
  • prepared_phase
  • positive
  • negative
  • source_av
  • mux_audio
  • conditioned_prompt
  • media_map_json
  • report_json
◄phaselow►
◄context_audiovideo_and_audio►
◄prompt—►
◄length124►
◄task_typeauto►
◄audio_modenative►
◄audio_denoise_strength0.35►
◄add_source_as_referencetrue►
◄prompt_primary_audio_ordinal0►
◄strict_prompt_tagstrue►
◄ref_image_sizematch►
◄reference_video_policyofficial_2_to_15s►
◄first_frame_reusesegment0_only►
◄persistent_identity_strategysingle_reference►
◄persistent_identity_interval1►

In most long-video setups there's one prompt box and a lot of hope. This pack's continuation family asks you to build conditions separately for the two phases, and after one run you understand why: LOW is generating cheap motion at half resolution, HIGH is finishing the thing at full resolution with the completed prefix locked. They're different jobs, and giving them the same prompt is a compromise, not a default.

H3 Continuation · ONE Phase Conditions is where that prompt lives. Despite the "ONE" in the name it's not a single node in your graph - you place one per phase, one for low and one for high, each with its own phase socket value, its own prompt and its own references.

What you get out of it

Required inputs: contexts (from Prepare Accepted Contexts), phase, clip, video_vae, audio_vae, context_audio, prompt, length, and the task- and audio-side options - task_type, audio_mode, audio_denoise_strength, add_source_as_reference, prompt_primary_audio_ordinal, strict_prompt_tags, ref_image_size, reference_video_policy. Optional: drive_audio, final_audio, first_frame, last_frame, ref_images, ref_videos, ref_video_audios, ref_audios, the first_frame_reuse / persistent_identity_* group, and semantic_bridge.

Outputs: prepared_phase (the typed phase handoff the LOW stage, HIGH handoff and HIGH stage all want), positive, negative, source_av, mux_audio, conditioned_prompt, media_map_json and report_json.

Two of those deserve comment.

context_audio defaults to video_and_audio. The author's tooltip: video_and_audio continues the generated AV latent; video_only keeps motion context but leaves audio to the selected native/source mode. If you're chasing continuity of room tone, the default is what you want; if you want to control the audio separately, that's the switch.

negative is emitted identical to positive. That's the native CFG-1 shape for this route, not a bug - don't expect to steer with the negative socket.

The knobs that matter

  • length defaults to 124 with a step of 17, matching H3's frame grid. Segments land on that grid; don't fight it.
  • first_frame is "exact frame 0 for the first segment" per its own tooltip - later segments normally ignore it, except when you've opted into persistent_identity_reference as a compatibility fallback.
  • first_frame_reuse defaults to segment0_only, which the author calls the legacy behaviour. Switching it to persistent_identity_reference adds one non-timeline image reference on continuation segments: it adds reference rows and VRAM, and the tooltip is explicit that it is "not identity lock".
  • persistent_identity_image is ignored on segment 0 and ignored entirely unless first_frame_reuse is set as above. persistent_identity_strategy (single-reference vs scene-plus-identity) and persistent_identity_interval (inject on every segment, or every other) are the drift-versus-motion dials - and the author calls the interval control an experiment, not an adaptive drift detector. Believe the output, not the dial.
  • audio_denoise_strength at 0.35 by default is the native audio refinement strength. Changing it changes the audio only in the modes that actually regenerate audio.

Install

Manager → search MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Fully quit and restart ComfyUI, then refresh the page. Recent Core with native H3 support is mandatory. No extra pip packages - the empty requirements.txt is a deliberate guarantee that installing this pack can't swap out your Torch/CUDA. Model side: H3 main model in models/diffusion_models, Qwen in models/text_encoders, video and audio VAE in models/vae, learned 3D upscaler from the author's HuggingFace (t8star). Example graphs: examples/workflows/45-progressive-continuation-split.

Where it goes wrong

Using one conditioning node for both phases. The pack's own description says use independent LOW and HIGH nodes, prompts and references. Sharing one node means the HIGH phase inherits LOW's prompt, and you've thrown away the main reason this family exists.

Mismatched source audio between phases. Both phases must retain matching source audio, or the AV continuation contract breaks.

Relay wired the wrong way. For Relay, you connect the projected compiled_prompt into this node first, then apply the continuation Relay node - not the other way round.

Fighting the 17-frame grid. length values off the grid produce misalignment with the plan and the context clock.

Nodes missing or sockets mismatched. Check custom_nodes for a stale second copy of the pack. Rename it with a .disabled suffix (leading underscore won't disable a node pack), restart, and verify python_module in /object_info.

Assuming saved parents mean saved recipes. Selecting an accepted parent authenticates that parent clip - not that today's edited LOW/HIGH recipe is the one that produced it. The pack treats those as two different claims, and so should you.

CategoryT8/MiniMax H3/Modular Sampling/Continuation Experimental

Inputs (29)

NameTypeDefaultDescription
contextsT8_CONTINUATION_STAGE_CONTEXTS—
phaseCOMBOlow2 options: low, high
clipCLIPNative MiniMax H3 Qwen3-VL CLIP.
video_vaeVAEMiniMax H3 video VAE.
audio_vaeVAEMiniMax H3 audio VAE.
context_audioCOMBOvideo_and_audiovideo_and_audio continues the generated AV latent. video_only keeps motion context but leaves audio to the selected native/source mode.
promptSTRING—
lengthINT124—
task_typeCOMBOauto7 options: auto, T2VA, I2VA, FL2VA, L2VA, Ref2VA, +1
audio_modeCOMBOnative4 options: lock_source, remix_source, reference_only, native
audio_denoise_strengthFLOAT0.350–1—
add_source_as_referenceBOOLEANtrue—
prompt_primary_audio_ordinalINT00–9—
strict_prompt_tagsBOOLEANtrue—
ref_image_sizeCOMBOmatch2 options: match, max
reference_video_policyCOMBOofficial_2_to_15s2 options: official_2_to_15s, model_minimum
drive_audiooptAUDIO—
final_audiooptAUDIO—
first_frameoptIMAGEExact frame 0 for the first segment. Later segments normally ignore it; persistent_identity_reference may reuse it as a compatibility fallback.
last_frameoptIMAGE—
ref_imagesoptCOMFY_AUTOGROW_V3—
ref_videosoptCOMFY_AUTOGROW_V3—
ref_video_audiosoptCOMFY_AUTOGROW_V3—
ref_audiosoptCOMFY_AUTOGROW_V3—
first_frame_reuseoptCOMBOsegment0_onlysegment0_only preserves legacy behavior. persistent_identity_reference adds one non-timeline image reference on continuation segments. A connected persistent_identity_image is preferred; otherwise first_frame is reused. This remains experimental, adds reference rows/VRAM, and is not identity lock.
persistent_identity_imageoptIMAGEOptional continuation-only identity crop. Prefer one clear face or upper-body image. It is ignored on segment 0 and unless first_frame_reuse is set to persistent_identity_reference; first_frame still owns exact frame 0.
persistent_identity_strategyoptCOMBOsingle_referencesingle_reference uses persistent_identity_image when connected, otherwise first_frame. scene_plus_identity supplies both images as separate references; it costs more reference rows/VRAM and remains a gated experiment.
persistent_identity_intervaloptINT11–32Continuation injection interval. 1 preserves the existing every-segment behavior; 2 injects on continuation segments 1, 3, 5, ... and lets the intermediate segments use motion context only. This is an Experimental identity-versus-motion control, not an adaptive drift detector.
semantic_bridgeoptT8_SEMANTIC_BRIDGE—

Outputs (8)

NameTypeDescription
prepared_phaseT8_CONTINUATION_PREPARED_PHASE—
positiveCONDITIONING—
negativeCONDITIONING—
source_avLATENT—
mux_audioAUDIO—
conditioned_promptSTRING—
media_map_jsonSTRING—
report_jsonSTRING—