H3 Audio Perception Producer
The node that tells you it can't hear you
- media
- result
You wired a reference clip into an audio perception node expecting a transcript. What you got back is a record that says, in effect, "nobody listened to this." That's not a bug - it's the entire design of this node, and once you get that, it's one of the more honest pieces of plumbing in the H3 pack.
What it's actually for
MiniMax H3 is an omni-modal model: audio is generated jointly with the picture, not bolted on afterwards (MiniMax H3 panel). So if you're feeding a reference video in, its soundtrack matters, and it would be very convenient if something could just describe that soundtrack for you. The Audio Perception Producer is the socket where that would happen.
In this release, it mostly doesn't. MiniMax H3 Studio's own capability manifest classifies it as unsupported, and the node's docstring is blunt about what it needs: a host-configured short English CPU ASR, using profile whisper_large_v3_cpu_en_v1 over the comfyui_native route. "Host-configured" is the key phrase - you can't set that up from the graph.
How it works
The node has two behaviours and picks between them at execution time, not in the UI.
If profile_id is the pinned profile name, the route matches, and cancel_requested is off, it calls a host-registered perception service and returns real transcript candidates under result. Those candidates are deliberately marked uncertain - the tooltip language is "transcript candidates remain uncertain," which is a fair description of ASR output you're about to bake into a prompt.
Anywhere else - which for almost everyone means the default unqualified - it short-circuits into a declared-unavailable result: a typed envelope that says perception did not happen, why not, and which route and profile you asked for. No invention, no plausible-sounding caption. Worth appreciating: most "AI enrichment" nodes in this ecosystem would have handed you a guess.
Inputs and outputs
Only one output, result (H3_AUDIO_PRODUCER_RESULT), which feeds downstream producers that consume perception results.
The inputs a beginner cares about:
- media - required, and it does not take a raw AUDIO socket. It takes the output of H3 Media Admission Producer, so you admit the asset first and then describe it. One admitted asset per producer node.
- profile_id - the string naming the qualified perception profile. Default
unqualified, which is the honest "I'm not pretending" state. - route and device -
comfyui_native/autoare the values the pinned profile accepts. The audio profile runs on CPU orauto; pickingcudaormpswith a pinned profile gets refused withunsupported_route_or_devicerather than silently falling back. - cancel_requested - a boolean. Setting it true doesn't cancel anything running; it's a typed non-execution request that returns a cancellation record instead of running.
Installing and setting up
The pack install is the same for every node here. It isn't in the Comfy Registry yet, so:
cd ComfyUI/custom_nodes
git clone https://github.com/rookiestar28/ComfyUI-MiniMaxH3-Studio.git
# restart ComfyUI
Keep both the repo's root __init__.py and the comfyui_h3_context/ folder. The package declares zero Python dependencies and never touches your models - the ffmpeg download mentioned in the README only happens when you use the sidebar's editor features.
Enabling actual perception is a host-operator job, not a widget: you register a perception host in Python with a python_executable, a private temporary_root, and a Whisper weight plus processor directory. It's documented in docs/HOST_INTEGRATION.md. Ordinary users skip it.
Where people get burned
If you're seeing failures here, it's usually one of three things. perception_unconfigured means exactly what it says - the profile exists as a name but no host service is registered for it. perception_profile_unavailable means you typed a profile id that isn't the qualified one, and the pack refuses to run a profile it can't prove works. And invalid_cancel_requested is the node being strict about types, which is a theme with this author - the same person's earlier pack, ComfyUI-OpenClaw, is a security-first automation pack built on explicit admin boundaries.
The practical advice: for 95% of workflows, don't try to make this node work. Type the dialogue you want into H3 Hard Constraint Producer and treat the audio as commanded rather than perceived. This node exists so that when the pipeline says "I don't know what this audio contains," that uncertainty is a typed fact on the wire instead of a hole someone fills in later.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| media | H3_MEDIA_PRODUCER_RESULT | — | |
| route | COMBO | comfyui_native | 3 options: comfyui_native, ollama, specialist |
| profile_id | STRING | unqualified | — |
| device | COMBO | auto | 4 options: auto, cpu, cuda, mps |
| cancel_requested | BOOLEAN | false | — |
| local_service_consentopt | BOOLEAN | false | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| result | H3_AUDIO_PRODUCER_RESULT | — |