Nodes/ComfyUI-MiniMaxH3-Studio/H3 Cross Reference Producer
ComfyUI Node

H3 Cross Reference Producer

Saying 'same person' without letting the model guess

By rookiestar28·Created 2 months ago·Updated a day ago· 79
H3 Cross Reference Producer
  • reference_registry
  • media
  • cross_reference_graph
  • producer_report
◄resolutionsource_only►
◄entity_kindsubject►
◄candidate_a►
◄candidate_b►

The hard part of a reference-set video was never getting a picture of a person into the model. It's the model deciding, on no evidence, that the face in image one and the voice in audio two belong to the same character - and then committing to that for fifteen seconds of generated footage. Keeping identity straight across a reference set is the same multi-subject attribution problem that wrecks captioners (character-consistency.md), and this node is the pack's answer: link what's provable, and label the rest ambiguous.

What it does

It builds a cross-reference graph: source-owned identities, derived from admitted media plus the canonical reference registry. source-owned is the operative word - the graph is anchored to assets you admitted, each with a stable asset_id, so an identity claim can be traced back to the exact asset it came from instead of floating free in the prompt.

The interesting discipline is what it refuses to do. In the default source_only resolution the node makes no identity claims at all - it records each admitted source as an observation bound to its own asset. If you switch resolution to ambiguous, it starts emitting an explicit proposal linking candidates, tagged with an AMBIGUOUS uncertainty: "caller retained unresolved identity candidates without inference." So even the act of saying "these might be the same person" ships with the doubt attached, rather than being promoted into a fact.

Inputs and outputs

Required:

  • reference_registry - the canonical registry from H3 Reference Registry. This is what owns the slots and roles; the producer doesn't re-decide what's a subject and what's a first frame.
  • media - an H3_MEDIA_PRODUCER_RESULT from H3 Media Admission Producer. One admitted asset.

Optional, and this is where the node's real behaviour lives:

  • resolution - source_only (default) or ambiguous. source_only produces a graph of unlinked observations.
  • entity_kind - what kind of thing you're claiming is shared: subject, voice, object or scene. Requires the ambiguous path.
  • candidate_a and candidate_b - the two candidate identities you're declining to resolve. Both are plain strings.

Outputs are cross_reference_graph (H3_CROSS_REFERENCE_GRAPH) and producer_report (H3_DOWNSTREAM_PRODUCER_REPORT). Wire the graph and its report together into H3 Full Reference Timeline Producer or H3 Local Reconstruction Acceptance - those nodes assert that a report actually belongs to the value it's paired with, so splitting the pair is how you get a stack of contract errors.

The report is the honest part

The producer report for this stage comes back with a disposition of PARTIAL, a reason code of ADMITTED_SOURCE_PROJECTED, and a fingerprint claim marked as caller-declared, unverified. That last one is worth pausing on. The node admits it has no way to verify the SHA-256 you typed into the media admission step. It's not pretending the label is provenance. In a pipeline that's about to fork identity across image, video and audio, "partial, unverified" is a much better thing to have on the wire than a confident green tick.

Install

Not in the Comfy Registry yet:

cd ComfyUI/custom_nodes
git clone https://github.com/rookiestar28/ComfyUI-MiniMaxH3-Studio.git
# restart ComfyUI

No dependencies to install (dependencies = []), no models downloaded at install, prebuilt frontend extension. Python 3.10+. The workflows/m15_02_downstream_*.json files in the repo are API-format graphs that wire this node next to the other producers - a decent reference for what "the same pair of outputs everywhere" looks like in practice.

Where it bites

Two things catch people.

First, the node requires a REF2VA-shaped pipeline underneath it. The graph it builds is scoped to a reference-set request, so an image-to-video request with a first frame wired in won't produce the cross-references you were expecting; the media admission reference_role is doing that work instead.

Second, if you set resolution to ambiguous, you must actually name both candidates. Empty candidate_a/candidate_b gives you an ambiguous proposal with no candidates in it, which is a receipt that says nothing. The useful pattern is: keep it on source_only until you have a real reason to claim two assets share an entity, then name the candidates and let the uncertainty ride along.

Categoryh3_context/assembly

Inputs (6)

NameTypeDefaultDescription
reference_registryH3_REFERENCE_REGISTRY—
mediaH3_MEDIA_PRODUCER_RESULT—
resolutionoptCOMBOsource_only2 options: source_only, ambiguous
entity_kindoptCOMBOsubject4 options: subject, voice, object, scene
candidate_aoptSTRING—
candidate_boptSTRING—

Outputs (2)

NameTypeDescription
cross_reference_graphH3_CROSS_REFERENCE_GRAPH—
producer_reportH3_DOWNSTREAM_PRODUCER_REPORT—