H3 Continuation · ONE Phase Conditions (T8 EXP)
The continuation node where you actually write the prompt — twice
- contexts
- clip
- video_vae
- audio_vae
- drive_audio
- final_audio
- first_frame
- last_frame
- ref_images
- ref_videos
- ref_video_audios
- ref_audios
- persistent_identity_image
- semantic_bridge
- prepared_phase
- positive
- negative
- source_av
- mux_audio
- conditioned_prompt
- media_map_json
- report_json
In most long-video setups there's one prompt box and a lot of hope. This pack's continuation family asks you to build conditions separately for the two phases, and after one run you understand why: LOW is generating cheap motion at half resolution, HIGH is finishing the thing at full resolution with the completed prefix locked. They're different jobs, and giving them the same prompt is a compromise, not a default.
H3 Continuation · ONE Phase Conditions is where that prompt lives. Despite the "ONE" in the name it's not a single node in your graph - you place one per phase, one for low and one for high, each with its own phase socket value, its own prompt and its own references.
What you get out of it
Required inputs: contexts (from Prepare Accepted Contexts), phase, clip, video_vae, audio_vae, context_audio, prompt, length, and the task- and audio-side options - task_type, audio_mode, audio_denoise_strength, add_source_as_reference, prompt_primary_audio_ordinal, strict_prompt_tags, ref_image_size, reference_video_policy. Optional: drive_audio, final_audio, first_frame, last_frame, ref_images, ref_videos, ref_video_audios, ref_audios, the first_frame_reuse / persistent_identity_* group, and semantic_bridge.
Outputs: prepared_phase (the typed phase handoff the LOW stage, HIGH handoff and HIGH stage all want), positive, negative, source_av, mux_audio, conditioned_prompt, media_map_json and report_json.
Two of those deserve comment.
context_audio defaults to video_and_audio. The author's tooltip: video_and_audio continues the generated AV latent; video_only keeps motion context but leaves audio to the selected native/source mode. If you're chasing continuity of room tone, the default is what you want; if you want to control the audio separately, that's the switch.
negative is emitted identical to positive. That's the native CFG-1 shape for this route, not a bug - don't expect to steer with the negative socket.
The knobs that matter
lengthdefaults to 124 with a step of 17, matching H3's frame grid. Segments land on that grid; don't fight it.first_frameis "exact frame 0 for the first segment" per its own tooltip - later segments normally ignore it, except when you've opted intopersistent_identity_referenceas a compatibility fallback.first_frame_reusedefaults tosegment0_only, which the author calls the legacy behaviour. Switching it topersistent_identity_referenceadds one non-timeline image reference on continuation segments: it adds reference rows and VRAM, and the tooltip is explicit that it is "not identity lock".persistent_identity_imageis ignored on segment 0 and ignored entirely unlessfirst_frame_reuseis set as above.persistent_identity_strategy(single-reference vs scene-plus-identity) andpersistent_identity_interval(inject on every segment, or every other) are the drift-versus-motion dials - and the author calls the interval control an experiment, not an adaptive drift detector. Believe the output, not the dial.audio_denoise_strengthat 0.35 by default is the native audio refinement strength. Changing it changes the audio only in the modes that actually regenerate audio.
Install
Manager → search MiniMax H3 Audio T8, or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Fully quit and restart ComfyUI, then refresh the page. Recent Core with native H3 support is mandatory. No extra pip packages - the empty requirements.txt is a deliberate guarantee that installing this pack can't swap out your Torch/CUDA. Model side: H3 main model in models/diffusion_models, Qwen in models/text_encoders, video and audio VAE in models/vae, learned 3D upscaler from the author's HuggingFace (t8star). Example graphs: examples/workflows/45-progressive-continuation-split.
Where it goes wrong
Using one conditioning node for both phases. The pack's own description says use independent LOW and HIGH nodes, prompts and references. Sharing one node means the HIGH phase inherits LOW's prompt, and you've thrown away the main reason this family exists.
Mismatched source audio between phases. Both phases must retain matching source audio, or the AV continuation contract breaks.
Relay wired the wrong way. For Relay, you connect the projected compiled_prompt into this node first, then apply the continuation Relay node - not the other way round.
Fighting the 17-frame grid. length values off the grid produce misalignment with the plan and the context clock.
Nodes missing or sockets mismatched. Check custom_nodes for a stale second copy of the pack. Rename it with a .disabled suffix (leading underscore won't disable a node pack), restart, and verify python_module in /object_info.
Assuming saved parents mean saved recipes. Selecting an accepted parent authenticates that parent clip - not that today's edited LOW/HIGH recipe is the one that produced it. The pack treats those as two different claims, and so should you.
Inputs (29)
| Name | Type | Default | Description |
|---|---|---|---|
| contexts | T8_CONTINUATION_STAGE_CONTEXTS | — | |
| phase | COMBO | low | 2 options: low, high |
| clip | CLIP | Native MiniMax H3 Qwen3-VL CLIP. | |
| video_vae | VAE | MiniMax H3 video VAE. | |
| audio_vae | VAE | MiniMax H3 audio VAE. | |
| context_audio | COMBO | video_and_audio | video_and_audio continues the generated AV latent. video_only keeps motion context but leaves audio to the selected native/source mode. |
| prompt | STRING | — | |
| length | INT | 124 | — |
| task_type | COMBO | auto | 7 options: auto, T2VA, I2VA, FL2VA, L2VA, Ref2VA, +1 |
| audio_mode | COMBO | native | 4 options: lock_source, remix_source, reference_only, native |
| audio_denoise_strength | FLOAT | 0.350–1 | — |
| add_source_as_reference | BOOLEAN | true | — |
| prompt_primary_audio_ordinal | INT | 00–9 | — |
| strict_prompt_tags | BOOLEAN | true | — |
| ref_image_size | COMBO | match | 2 options: match, max |
| reference_video_policy | COMBO | official_2_to_15s | 2 options: official_2_to_15s, model_minimum |
| drive_audioopt | AUDIO | — | |
| final_audioopt | AUDIO | — | |
| first_frameopt | IMAGE | Exact frame 0 for the first segment. Later segments normally ignore it; persistent_identity_reference may reuse it as a compatibility fallback. | |
| last_frameopt | IMAGE | — | |
| ref_imagesopt | COMFY_AUTOGROW_V3 | — | |
| ref_videosopt | COMFY_AUTOGROW_V3 | — | |
| ref_video_audiosopt | COMFY_AUTOGROW_V3 | — | |
| ref_audiosopt | COMFY_AUTOGROW_V3 | — | |
| first_frame_reuseopt | COMBO | segment0_only | segment0_only preserves legacy behavior. persistent_identity_reference adds one non-timeline image reference on continuation segments. A connected persistent_identity_image is preferred; otherwise first_frame is reused. This remains experimental, adds reference rows/VRAM, and is not identity lock. |
| persistent_identity_imageopt | IMAGE | Optional continuation-only identity crop. Prefer one clear face or upper-body image. It is ignored on segment 0 and unless first_frame_reuse is set to persistent_identity_reference; first_frame still owns exact frame 0. | |
| persistent_identity_strategyopt | COMBO | single_reference | single_reference uses persistent_identity_image when connected, otherwise first_frame. scene_plus_identity supplies both images as separate references; it costs more reference rows/VRAM and remains a gated experiment. |
| persistent_identity_intervalopt | INT | 11–32 | Continuation injection interval. 1 preserves the existing every-segment behavior; 2 injects on continuation segments 1, 3, 5, ... and lets the intermediate segments use motion context only. This is an Experimental identity-versus-motion control, not an adaptive drift detector. |
| semantic_bridgeopt | T8_SEMANTIC_BRIDGE | — |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| prepared_phase | T8_CONTINUATION_PREPARED_PHASE | — |
| positive | CONDITIONING | — |
| negative | CONDITIONING | — |
| source_av | LATENT | — |
| mux_audio | AUDIO | — |
| conditioned_prompt | STRING | — |
| media_map_json | STRING | — |
| report_json | STRING | — |