Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 Long Relay · 窗口文本策略 (T8 EXP)
ComfyUI Node

MiniMax H3 Long Relay · 窗口文本策略 (T8 EXP)

Stop Re-Sending the Scene You Already Shot

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 Long Relay · 窗口文本策略 (T8 EXP)
  • prompt_relay_plan
  • prompt_relay_plan
  • report_json
◄text_policypreserve_all►

What it is

Prompt Relay is this pack's answer to the problem that a single prompt can't describe a four-shot sequence. You write one global prompt for the people, style, camera continuity and overall sound, then attach timed local events - this happens at 0-2s, that at 2-5s - and the pack biases the target video queries toward the local text for the right slice of the timeline. It's conceptually the same job LTX Director does with its timeline and Relay integration: make "this, then that, at this moment" expressible in a graph that otherwise only takes one prompt.

Long video is where that gets awkward. The renderer projects the plan window by window, and the classic all-key behavior means every window sees all the event text - including the events that already finished and the ones that haven't started. That's the paper-faithful shape, and it's also how you get a window that keeps re-describing a scene you moved past six segments ago.

This node is the switch that decides which text each delivery window actually gets sent to CLIP. It changes nothing about the sampler, the sigmas, the accepted AV context or the events themselves - it's a text policy attached to the plan.

How it works

Wire it between the global Plan / Query Route and the long-video inner loop. Two things are worth knowing about the mechanism:

  • The policy is stored on the global plan, and the node refuses to touch a plan that's already been projected. Feed it a projected window and you get Configure the global Plan before projecting a window. That error message is the entire wiring guide.
  • Changing the policy invalidates and recomputes the plan hash, because the plan is the identity downstream stages cache against. preserve_all is a true no-op - it drops the key entirely, and if the policy already matches, the input object passes straight through.

There are three policies. preserve_all is the old behavior and stays the default. accepted_window_text_exp keeps only the local text that intersects the delivered window; finished, future, or known-context-only events are no longer pushed into CLIP for that window. Crossing-event text stays, and its original sigma is untouched. dialogue_start_owner_exp adds a second rule on top: tagged <d>…</d> dialogue is kept only in the delivery window that owns the event's start, so a line isn't re-fed to every later window. Cross-segment visual description and AV context aren't deleted in either mode.

The inputs

  • prompt_relay_plan (H3_T8_PROMPT_RELAY_PLAN) - the global plan, before projection. Type a node ID or a projected window here and you'll get an error rather than a subtle quality regression.
  • text_policy - the combo described above. Default preserve_all.

Outputs are prompt_relay_plan (same type, wired on into the long-video loop) and report_json, which records the chosen policy, the new plan hash, and an explicit experimental: true for anything that isn't preserve_all. That last field matters when you're comparing runs later.

Installing

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Restart fully, refresh the browser, or install through Manager by searching MiniMax H3 Audio T8. The pack has no pip dependencies of its own; EXP features check their own optional requirements when you use them. Model-side prerequisite is the same as any H3 workflow - main model, Qwen text encoder and the H3 VAEs placed per the pack's install doc.

Where people get burned

Dialogue tags are validated strictly. The dialogue policy requires exact paired <d>...</d> tokens; nested, unpaired or unclosed tags raise. And the global prompt can't hold <d> at all - once-only dialogue has to live in its timed local event, ideally entirely inside one physical segment, because a line split across a boundary is exactly the case this policy can't save you from.

It's an experiment, and the author says so twice. This is not the paper's all-key Relay, and it is not a no-repeat guarantee. It also only affects long-video window projection - if you're running ordinary Prompt Relay, this node changes nothing.

Don't judge it on one clip. Run the same seed, same model, same resolution and same events with preserve_all versus the policy you're testing, and listen to the audio too. Text conditioning changes here can reach the joint AV transformer indirectly even though the audio logits don't get the bias directly.

Combining policies needs a new chain id. Keep the plan's provenance honest rather than editing a plan in place and wondering why a cached stage returned the old result.

CategoryT8/MiniMax H3/Conditioning/Experimental

Inputs (2)

NameTypeDefaultDescription
prompt_relay_planH3_T8_PROMPT_RELAY_PLAN接全局Plan/Query Route后,再接长视频内循环;不要接已投影窗口。
text_policyCOMBOpreserve_allpreserve_all保持旧图。accepted_window_text_exp只保留与当前交付窗口相交的局部文本;已结束/未来/只在已知上下文的事件不重复送入CLIP。跨段事件仍保留原sigma。dialogue_start_owner_exp另外只让事件起始交付窗口保留<d>台词;跨段视觉描述及AV上下文不删。Global不可放<d>,未标记的对白无法识别。一次性台词尽量完整放入一个物理段;不保证逐字/不重复,组合需新chain_id。

Outputs (2)

NameTypeDescription
prompt_relay_planH3_T8_PROMPT_RELAY_PLAN—
report_jsonSTRING—