Nodes/ComfyUI-MiniMaxH3-Studio/H3 Media Admission Producer
ComfyUI Node

H3 Media Admission Producer

How a file becomes a reference the pipeline will trust

By rookiestar28·Created 2 months ago·Updated a day ago· 79
H3 Media Admission Producer
  • image
  • video
  • audio
  • media
◄media_kindimage►
◄asset_idasset-1►
◄declared_source_fingerprint►
◄width_pixels64►
◄height_pixels64►
◄duration_seconds0.00►
◄sample_rate_hz0►
◄channel_count0►
◄reference_rolereference►
◄reference_order0►

If you're going to hand a video model a reference and then act on what the pipeline knows about that reference, someone has to write down what the reference is first. That's admission: one node, one media value, one typed envelope that says what the asset is, how big it is, how long it runs, and what role it plays.

The pack's reference workflows use one admission node per asset. Three references means three of these.

What it does

It admits a single host-owned media value and emits a locator-free typed envelope. "Locator-free" is worth translating: the envelope carries your asset_id and declared properties, but not the file path or the runtime payload. The pixels, video frames or samples stay on their own connection, and the metadata travels separately through the producer chain. Nothing in the report can leak where your file lives.

It accepts exactly one optional socket - image, video or audio - and the required media_kind widget has to match what you plugged in. Declare image and wire an audio object and admission refuses rather than coercing.

Inputs and outputs

The widgets you'll actually set:

  • media_kind - image, video, or audio. Match it to your socket.
  • asset_id - a stable public identifier like asset-1. This is what the rest of the pipeline calls your asset.
  • reference_role - what the asset is for: input, reference, first_frame, last_frame, or paired_audio. This is the H3 conditioning role, and it's the field that decides whether your image becomes a first frame, a last frame, or an identity reference.
  • reference_order - 0..63, the order among grouped media. Roles and order are never guessed, so this is how you control which reference wins.
  • width_pixels, height_pixels, duration_seconds, sample_rate_hz, channel_count - declared properties. Zero means "not applicable."
  • declared_source_fingerprint - a lowercase SHA-256 you type in. The tooltip is refreshingly candid: "Unverified caller-declared lowercase SHA-256; not trusted-loader provenance." The pipeline records it as a claim, not as verified fact.

One output: media (H3_MEDIA_PRODUCER_RESULT). It feeds the perception producers, H3 Evidence Fusion Producer, H3 Cross Reference Producer and the full-reference timeline producer.

Why the separation exists

H3 is an omni-modal model - text, image, video and audio as one context, with audio generated jointly rather than in a second pass (MiniMax H3 panel). A reference set can therefore contain several kinds of thing, each with a different job. The pack's rule is that media always stays on its own connections and metadata travels as typed values, which is why you'll see a graph where LoadImage goes to one socket and a structured envelope goes to another.

Practically, this is what lets the reference roles be explicit instead of inferred. An image in slot 0 declared as first_frame and an image in slot 1 declared as reference are different requests, and the difference is recorded at admission rather than guessed from which socket happens to be occupied.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/rookiestar28/ComfyUI-MiniMaxH3-Studio.git
# restart ComfyUI

The pack isn't in the Comfy Registry yet, so this manual clone is the install. No Python dependencies (dependencies = []), no downloads at install, prebuilt browser extension. Python 3.10+, and H3 weights only if you intend to generate - building and validating a reference pipeline needs none.

Where people get burned

The classic mistake is reusing one admission node for several assets. It admits exactly one value; a second asset needs a second node with its own asset_id, reference_role and reference_order.

Then there's the audio-only reference, which the pack doesn't offer at all. The README's known limitations say it plainly: audio can't be the only reference - add an image or a video alongside it. That's a model-side constraint reflected in the UI rather than an arbitrary rule, and you'll hit it as a blocked role selection rather than a crash.

And on the fingerprint: declared_source_fingerprint is not verification. Downstream reports label it caller_declared_unverified, so if you were hoping this would prove the loader didn't swap your file, it won't. It's a note to your future self about which version of an asset a run used - genuinely useful when you have six near-identical reference clips, as long as you don't mistake it for a chain of custody.

Categoryh3_context/perception

Inputs (13)

NameTypeDefaultDescription
media_kindCOMBOimage3 options: image, video, audio
asset_idSTRINGasset-1—
declared_source_fingerprintSTRINGUnverified caller-declared lowercase SHA-256; not trusted-loader provenance.
width_pixelsINT640–32768—
height_pixelsINT640–32768—
duration_secondsFLOAT0.000–86400—
sample_rate_hzINT00–768000—
channel_countINT00–64—
reference_roleCOMBOreference5 options: input, reference, first_frame, last_frame, paired_audio
reference_orderINT00–63—
imageoptIMAGEHost-owned image value.
videooptVIDEOHost-owned video value.
audiooptAUDIOHost-owned audio value.

Outputs (1)

NameTypeDescription
mediaH3_MEDIA_PRODUCER_RESULT—