MiniMax H3 TST Temporal Query / 时序Query修正 (EXP/T8)
A research paper's temporal fix as an optional switch
- model
- sigmas
- model
- configuration_json
TST stands for Temporal State Transport, and it comes from a paper rather than from a workflow thread. The pitch is that a video model's attention treats time badly in specific, measurable ways, and that you can nudge the video queries with a signed correction derived from how "spread out" attention is per frame. The pack's version is one MODEL patch node, default off, and it's the most openly provisional thing in the pack - the author's notes say full-weight quality acceptance isn't finished. Read this as a lever for experiments, not a quality preset.
How the mechanism actually works
Here's the version the author committed to, because the paper, the paper's own repo, and this node are three different formulas and they say so out loud. Call it a frame-mean proxy:
- Pick the target video tokens out of the real packed layout - not "everything at the end" - and, per head, spatially average the post-RoPE Q and K of each frame.
- Build an F×F softmax operator over frames.
- Compute a signed imbalance
Tas the normalized row entropy minus the spectral entropy. - Multiply the target video queries by
exp(tau_eff · T).
The paper's version operates on the query temperature before attention; the upstream codebase currently multiplies the attention output by a gain instead. This node does neither of those. If you read the paper and then diff the node, that's the gap. Non-target queries and all keys/values are left alone - though since H3's network is joint audio-video, changing video queries can still reach the sound indirectly. Nobody promises a locked audio track.
Inputs and output
model- your H3 branch after its LoRA chain and single attention backend.sigmas- the complete native video sigma table, and it must match the sampler. This is the field people get wrong: HIGH uses its own interval of the full table and must not restart from zero. Hand it a table with only the last four steps and you've broken the clock.mode-disabled(default) installs no patch at all, not even a wrapper.report_onlymeasures without modifying queries.apply_expactually corrects them.tau(default 0.2, 0–2) - correction strength.0still computes diagnostics, so it is not the same asdisabled. Bigger isn't clearer; you fix the material and A/B it, or you're just nudging a knob.max_workspace_mib(default 256) - an explicit budget for the temporary tensors this node allocates. Exceed it and you get a refusal, not a crash. It's a workspace estimate, not a VRAM ceiling or an OOM guarantee.
Outputs are model and configuration_json. The config report only ever says configured_not_executed - this node configures, it doesn't sample. Real per-run evidence shows up in the console under [T8 TST], and inside the progressive node's own report.
Where to put it in the graph
LoRA chain → one attention backend → optional Relay/EAV config → TST → native Euler sampler
CFG must be 1. Native Euler with a continuous interval of the full table is the only sampler contract accepted - Heun and anything else that evaluates the model more than once per step needs a different clock and gets rejected. The fixed internal order matters too: complete Q/K and target layout, then TST's query correction, then Relay's time bias and chunking, then Sage/KJ/Sol or an explicit fallback, then EAV's gain. Don't run two TST nodes on one branch, and don't combine this node with the progressive sampler's internal tst_mode setting - that's two configurations fighting over one branch.
Install
Manager, search MiniMax H3 Audio T8, full restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
No extra packages - the pack's requirements file is intentionally empty so it can never replace ComfyUI's Torch/CUDA stack.
Should you turn it on?
Probably not yet, and that's the interesting part. disabled and report_only are the useful settings today: they let you measure what TST would do on your footage without changing a single pixel - the pack tested frame-for-frame identity against native output in the disabled case. apply_exp is where the uncertainty lives, and the author says plainly that tensor-score improvements are not video-quality acceptance.
Two smaller caveats. The Sol attention backend steps aside when it meets Relay time bias or unequal query lengths, so a run where Sol falls back isn't evidence of sparse acceleration. And don't chain Sage and a Sol replacer together; one backend per branch is the rule across this entire pack.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| sigmas | SIGMAS | 完整原生视频sigma表;HIGH只使用其对应区间,不能传只有后4步的表冒充完整8步。 | |
| mode | COMBO | disabled | disabled不安装补丁;report_only测量不改Query;apply_exp实际修正。 |
| tau | FLOAT | 0.200–2 | 修正强度;0仍计算诊断。增大不等于更清晰,需固定素材对照。 |
| max_workspace_mib | INT | 2561–32768 | 显式临时tensor预算,超过会拒绝;不是整卡显存上限或无OOM保证。 |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| configuration_json | STRING | — |