H3 Media Admission Producer
How a file becomes a reference the pipeline will trust
- image
- video
- audio
- media
If you're going to hand a video model a reference and then act on what the pipeline knows about that reference, someone has to write down what the reference is first. That's admission: one node, one media value, one typed envelope that says what the asset is, how big it is, how long it runs, and what role it plays.
The pack's reference workflows use one admission node per asset. Three references means three of these.
What it does
It admits a single host-owned media value and emits a locator-free typed envelope. "Locator-free" is worth translating: the envelope carries your asset_id and declared properties, but not the file path or the runtime payload. The pixels, video frames or samples stay on their own connection, and the metadata travels separately through the producer chain. Nothing in the report can leak where your file lives.
It accepts exactly one optional socket - image, video or audio - and the required media_kind widget has to match what you plugged in. Declare image and wire an audio object and admission refuses rather than coercing.
Inputs and outputs
The widgets you'll actually set:
- media_kind -
image,video, oraudio. Match it to your socket. - asset_id - a stable public identifier like
asset-1. This is what the rest of the pipeline calls your asset. - reference_role - what the asset is for:
input,reference,first_frame,last_frame, orpaired_audio. This is the H3 conditioning role, and it's the field that decides whether your image becomes a first frame, a last frame, or an identity reference. - reference_order - 0..63, the order among grouped media. Roles and order are never guessed, so this is how you control which reference wins.
- width_pixels, height_pixels, duration_seconds, sample_rate_hz, channel_count - declared properties. Zero means "not applicable."
- declared_source_fingerprint - a lowercase SHA-256 you type in. The tooltip is refreshingly candid: "Unverified caller-declared lowercase SHA-256; not trusted-loader provenance." The pipeline records it as a claim, not as verified fact.
One output: media (H3_MEDIA_PRODUCER_RESULT). It feeds the perception producers, H3 Evidence Fusion Producer, H3 Cross Reference Producer and the full-reference timeline producer.
Why the separation exists
H3 is an omni-modal model - text, image, video and audio as one context, with audio generated jointly rather than in a second pass (MiniMax H3 panel). A reference set can therefore contain several kinds of thing, each with a different job. The pack's rule is that media always stays on its own connections and metadata travels as typed values, which is why you'll see a graph where LoadImage goes to one socket and a structured envelope goes to another.
Practically, this is what lets the reference roles be explicit instead of inferred. An image in slot 0 declared as first_frame and an image in slot 1 declared as reference are different requests, and the difference is recorded at admission rather than guessed from which socket happens to be occupied.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/rookiestar28/ComfyUI-MiniMaxH3-Studio.git
# restart ComfyUI
The pack isn't in the Comfy Registry yet, so this manual clone is the install. No Python dependencies (dependencies = []), no downloads at install, prebuilt browser extension. Python 3.10+, and H3 weights only if you intend to generate - building and validating a reference pipeline needs none.
Where people get burned
The classic mistake is reusing one admission node for several assets. It admits exactly one value; a second asset needs a second node with its own asset_id, reference_role and reference_order.
Then there's the audio-only reference, which the pack doesn't offer at all. The README's known limitations say it plainly: audio can't be the only reference - add an image or a video alongside it. That's a model-side constraint reflected in the UI rather than an arbitrary rule, and you'll hit it as a blocked role selection rather than a crash.
And on the fingerprint: declared_source_fingerprint is not verification. Downstream reports label it caller_declared_unverified, so if you were hoping this would prove the loader didn't swap your file, it won't. It's a note to your future self about which version of an asset a run used - genuinely useful when you have six near-identical reference clips, as long as you don't mistake it for a chain of custody.
Inputs (13)
| Name | Type | Default | Description |
|---|---|---|---|
| media_kind | COMBO | image | 3 options: image, video, audio |
| asset_id | STRING | asset-1 | — |
| declared_source_fingerprint | STRING | Unverified caller-declared lowercase SHA-256; not trusted-loader provenance. | |
| width_pixels | INT | 640–32768 | — |
| height_pixels | INT | 640–32768 | — |
| duration_seconds | FLOAT | 0.000–86400 | — |
| sample_rate_hz | INT | 00–768000 | — |
| channel_count | INT | 00–64 | — |
| reference_role | COMBO | reference | 5 options: input, reference, first_frame, last_frame, paired_audio |
| reference_order | INT | 00–63 | — |
| imageopt | IMAGE | Host-owned image value. | |
| videoopt | VIDEO | Host-owned video value. | |
| audioopt | AUDIO | Host-owned audio value. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| media | H3_MEDIA_PRODUCER_RESULT | — |