Nodes/comfyui-minimax-h3-audio-T8/H3 Semantic Bridge / 模型与设置 (T8 EXP)
ComfyUI Node

H3 Semantic Bridge / 模型与设置 (T8 EXP)

A 5120→5120 conditioning adapter, not a LoRA and not a speedup

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Semantic Bridge / 模型与设置 (T8 EXP)
    • semantic_bridge
    • report_json
    ◄model_name▾►
    ◄enabledtrue►
    ◄alpha0.10►
    ◄magnitude_matchper_token►
    ◄token_scopeall_tokens►
    ◄deviceauto►
    ◄chunk_tokens256►

    MiniMax H3's whole trick is that picture and audio come out of one model and one context - text, image, video and audio land in the same token sequence, so dialogue and room tone are generated alongside the frame instead of bolted on afterwards. (The weights also carry a licence that excludes the US, EU, UK and South Korea, so "open" has an asterisk.)

    The Semantic Bridge nodes sit inside that design. They take H3's raw text conditioning embeddings - the [B, T, 5120] rows the native encoder produces - and push them through a small learned adapter before the sampler sees them. That's the entire mechanism: a 5120→512→512→5120 SiLU network, six tensors, a few milliseconds of math. It is not a LoRA, it is not a new base model, it does not make anything faster, and it does not load a teacher model.

    This node is the settings half of a pair. MiniMaxH3SemanticBridgeConfigT8 produces the config object; MiniMaxH3SemanticBridgeApplyT8 applies it to a CONDITIONING. You need both.

    The two models, and why you don't stack them

    There are two independently trained weights and they do different jobs. The original (speach1sdef178/MiniMax-H3-Semantic-Bridge, FP16) is the general semantic bridge. BUNNY ActionLogic (JOKER141/BUNNY_H3_Conditioning_Bridge, FP32) is a separately trained action-semantic adapter - not a conversion of the first one. The author's framing is that the original is general-purpose semantics while BUNNY leans towards action, and the docs are emphatic that you don't chain them. One stage, one bridge.

    T8mars hosts lossless wrapper copies at t8star/Semantic-Bridge-Comfy that keep the original tensor values and dtypes and just add provenance and per-tensor checksums. Download into ComfyUI/models/semantic_bridge/, keeping the t8_compat subdirectory:

    hf download t8star/Semantic-Bridge-Comfy --local-dir ComfyUI/models
    

    Nothing auto-downloads, and the node will not reach out to a teacher model or fetch anything at sample time. Drop the author's own safetensors in the same folder if you prefer; the model_name combo enumerates everything under models/semantic_bridge recursively (including extra_model_paths roots) and shows the real path when two files share a relative name, so a duplicate filename doesn't silently pick one.

    The settings that change your result

    alpha is the only one that moves the needle for most people, and it defaults to a deliberately timid 0.10. Zero strength, or enabled off, is a complete bypass - the node doesn't even read the model file and adds no receipt metadata, which makes it a free A/B switch. Resist the urge to "fix" a bad generation by turning alpha up; the pack's own guidance is to do a same-seed comparison first.

    magnitude_match (default per_token) controls how the adapter's output magnitude is normalised back to the original conditioning. per_token matches each token row; global uses statistics over the whole conditioning item; none skips it. These are genuinely different behaviours, not equivalent spellings of the same thing.

    token_scope (default all_tokens) reproduces the original author's behaviour by touching every row. text_only_preserve_reference only modifies the rows tagged as text by H3's native token tags and leaves reference rows alone - and the node's own tooltip says it still can't guarantee singing survives. That's the honest version of a setting people will otherwise treat as a magic switch.

    The remaining ones are housekeeping: device (auto follows the conditioning tensor), chunk_tokens (256, bounds the adapter's temporary workspace - it does not change clip length and is not segmented generation), and model_name, which is validated up front so a missing file gives a clean error instead of a half-run.

    Outputs are the T8_SEMANTIC_BRIDGE object and a report_json string. Wire the first into Apply's semantic_bridge input; the second is diagnostic - model identity, config, application count - not a quality score.

    Install and the trap doors

    The pack installs the normal ways - ComfyUI Manager, search MiniMax H3 Audio T8, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
    

    Then fully restart ComfyUI, because new Python classes don't appear from a browser refresh. The pack's requirements.txt is intentionally empty, so installing it can't replace ComfyUI's Torch/CUDA stack; the bridge itself needs nothing but a recent ComfyUI with native H3 support.

    Where this bites: reference audio and singing. The author reports degraded results in reference-audio experiments, and original audio passing through untouched is not evidence that generated audio is fine - listen to the whole second segment of a two-segment run before deciding. Second, changing model, alpha, scope or the sampling recipe invalidates the bridge-enabled cache contract; start a new chain id instead of reusing results from the old settings. Third, if you see the bridge applied twice, pull the extra Apply node out and re-encode from the native conditioning node rather than layering another pass on top.

    If you're running H3 through Prompt Relay, don't put an Apply node after the Relay - Relay has its own internal socket for this. Same story for the long-video in-node loops, which expose their own per-pass slots.

    CategoryT8/MiniMax H3/Semantic Bridge

    Inputs (7)

    NameTypeDefaultDescription
    model_nameCOMBO原版/BUNNY:https://huggingface.co/t8star/Semantic-Bridge-Comfy ;T8动漫战斗:https://huggingface.co/t8star/semantic_bridge_T8-comic-combat 。放在 models/semantic_bridge/t8_compat;不是LoRA或H3主模型。
    enabledBOOLEANtrue—
    alphaFLOAT0.100–1—
    magnitude_matchCOMBOper_token3 options: per_token, global, none
    token_scopeCOMBOall_tokensall_tokens复现原作者;text_only仅修改原生tag=1行,仍不能保证歌声。
    deviceCOMBOauto3 options: auto, cpu, cuda
    chunk_tokensINT2560–655360 = whole sequence; required by some trained Transformer bridges.

    Outputs (2)

    NameTypeDescription
    semantic_bridgeT8_SEMANTIC_BRIDGE—
    report_jsonSTRING—