Nodes/comfyui-minimax-h3-audio-T8/T8 音频来源/缓存说明(只读)
ComfyUI Node

T8 音频来源/缓存说明(只读)

Is that your recording or generated audio? This node answers it

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
T8 音频来源/缓存说明(只读)
    • 说明
    • report_json
    ◄report_json►

    H3's audio has several possible origins in one graph - the recording you fed in, audio the model generated, a reference track used only for timbre, a remix - and the difference is invisible on the canvas. Two workflows can look identical and deliver completely different soundtracks. MiniMaxH3AudioSourceExplanationT8 is the node that turns the report JSON your conditions and samplers already emit into a sentence a human can read.

    It is a diagnostics node. It changes nothing.

    What you feed it, what you get

    One input: report_json. It is not a prompt box. The tooltip is explicit - connect an existing conditioning/sampling/upscale report to it, don't type audio settings into it. In practice you take report_json from an H3 Audio Conditioning node, a sampler, or an upscale node and wire it here.

    Two outputs: 说明 (the human-readable explanation) and the same report_json passed through so you can chain onward. It's marked as an output node, so the explanation renders in the node's own preview panel - no separate save step, just read it on the canvas.

    How it decides

    It parses JSON and walks it, looking for a known set of keys rather than guessing at prose. Only exact machine fields are trusted; the code comment about this is basically a scar, because a diagnostics node that free-associates is worse than none.

    What it can explain:

    • audio_mode - lock_source means the source recording is driving the picture, and it is not a voice clone generating new lines. native means the model generated the audio and any connected reference track is not the delivered sound. reference_only means the source audio is timbre reference only, not preserved. remix_source means the source participates in a remix, with no promise the original survives.
    • Cache state - booleans like low_reused, high_reused, cache_hit, reused are translated into "reused validated result" or "missed that cache", which is how you find out that your second run didn't actually re-sample the LOW stage.
    • Delivery policy - keys like final_audio_policy, delivery_audio_source, audio_policy and audio_source are mapped to plain statements about where the delivered track comes from.
    • Second-pass input policy - under audio_policy, effective_source distinguishes legacy_policy (native joint continuation), first_pass (from pass one - and to know whether it was actually locked you must also read effective_strength, not the source alone), and highres_template.

    Anything it doesn't recognise it flags as unrecognised rather than inventing meaning: an unknown delivery value becomes "unrecognised source claim, not automatically interpreted as voice cloning or preserved audio". If the report contains nothing it recognises, it says so and stays unknown - and states that unknown is not failure.

    The error you will hit

    The most common mistake is pasting text into the box. The node accepts two formats: JSON, and the plain-text key=value report that H3 Audio Conditioning intentionally emits (task=, audio_mode=, source_audio_tag=, frames=, canvas=). Anything else raises an error whose text (translated from the Chinese) is: connect an existing report_json / Conditioning report - this box is not a prompt and not an audio setting. If you see it, your chain is wired wrong, not broken.

    Installing

    ComfyUI Manager → search MiniMax H3 Audio T8, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
    

    Restart ComfyUI completely, then refresh the browser; new Python doesn't load on a frontend refresh. The pack's requirements.txt adds no packages - this node in particular is pure JSON parsing and has no model, GPU or FFmpeg requirement, so it works on a cold queue while you're debugging.

    Why bother

    Because "the audio is wrong" is the hardest bug class in an AV model to localise, and this is the cheapest instrument in the box. The pack's own audio notes are careful about a fact worth internalising: a report declaring an audio policy is not proof of the delivered track. This node tells you which claim is in the graph so you know which claim to go verify downstream - it does not itself sample, decode, or re-time anything, and it never touches you cache identity.

    CategoryT8/MiniMax H3/Diagnostics

    Inputs (1)

    NameTypeDefaultDescription
    report_jsonSTRING连接已有条件/采样/放大报告,不是提示词,不改变声音。

    Outputs (2)

    NameTypeDescription
    说明STRING—
    report_jsonSTRING—