Nodes/MiniMax H3 Edit/Add H3 Edit Reference
ComfyUI Node

Add H3 Edit Reference

Piling Up <Picture 2>, <Picture 3>, <Picture 4> In Order

By ethanfel·Created about a month ago·Updated 19 days ago· 30
Add H3 Edit Reference
  • image
  • previous_references
  • references
  • info
◄transportsemantic (Qwen only)►
◄semantic_resolution1024►
◄native_reference_sizematch output area►

The encoder's reference_image socket holds exactly one guide, and H3 prompting is all about numbered references - <Picture 2> is the glasses, <Picture 3> is the jacket. Add H3 Edit Reference is how you get from two pictures to five, in the order you actually meant.

It does one small thing well: it takes an image (or a whole IMAGE batch), stamps a transport policy on each entry, and hands you back a stack you can chain into another one of these nodes, and then into reference_stack on Text Encode H3 Edit / Generate.

What it actually does

The stack isn't a fixed-size array and there's no "number of references" widget. Each node appends its input onto whatever the previous builder passed in through previous_references, so you build the list by wiring builders in a chain. The last one's references output goes to the encoder.

Two things are worth knowing before you wire it:

A batch counts as separate references. If your image input is a batch of three, you get three entries in the stack, in batch order. That's deliberate - it's how people feed a folder of angles without five LoadImage nodes.

The node doesn't decide the numbering. source_image on the encoder is always <Picture 1>, a connected direct reference_image is <Picture 2> (unless its mode is none (source only)), and then stack entries follow in chain order. The info output spells out exactly which stack positions you just added, which is the fastest way to confirm you're about to describe <Picture 4> correctly in your prompt. Connect it to a text preview node while you're building; it's free.

The inputs you'll actually touch

transport is the one that matters. semantic (Qwen only) sends the image through the Qwen3-VL encoder as visual tokens and stops there - no VAE latent, no extra memory, and it's the right choice for transferring an idea: a pair of glasses, a material, a palette, a category of object. native (Qwen + VAE ref) also VAE-encodes the image into a minimax_refs block, which gives a much stronger low-level match but costs VRAM and, per the author's own README, stays experimental when it's mixed into an edit graph that already uses a frame-zero keyframe.

semantic_resolution (256–3584, step 32, default 1024) is an equivalent-square pixel budget for the Qwen path - aspect ratio is preserved, so you're setting how many vision tokens the encoder spends, not a literal width. Bump it if the model is missing fine detail on a small subject; leave it at 1024 otherwise.

native_reference_size picks the VAE resize policy when the transport is native: match output area or up to 2048px short edge.

The output side is just references and info.

Install

ComfyUI Manager → search the pack title, or:

cd ComfyUI/custom_nodes
git clone https://github.com/ethanfel/ComfyUI-MiniMax-H3-Edit

Restart, and the nodes land under MiniMax H3/Edit. The pack's requirements.txt is literally a comment saying nothing beyond ComfyUI's own runtime is needed, so there's no pip step to fight - this is one of the rare packs that installs clean into an existing environment.

What you do need is the model stack: a current ComfyUI with native MiniMax H3 support, the H3 Qwen3-VL text encoder, the H3 video VAE, and an FL2VA or REF2VA diffusion model depending on the route you're taking. Those are big downloads, and the H3 weights are territory-restricted by their licence, which is worth knowing before you queue a 40-odd-gigabyte download.

Where people get burned

  • Ordering surprises. Adding a guide changes the ordinal of everything after it. If you insert a new builder in the middle of a chain, your prompt's <Picture 3> now points at something else. The info output exists for exactly this.
  • Mixing transports. Native entries in an edit graph that has a strong frame-zero anchor are flagged experimental in the source, and for good reason: Qwen-only pictures have no matching VAE reference block, so a mixed stack presents the checkpoint with a task it wasn't trained for. Semantic-only stacks are the safe default.
  • Reference count is not free. Qwen context, conditioning time, RAM and VRAM all scale with entries. Four or five well-chosen references behave better than twelve lazy ones.
  • Redundant sizing. If you already set semantic_resolution on H3 Edit Options, that value is used unless you enable overrides. Setting a different number on each builder while overrides are off is a no-op you'll spend twenty minutes blaming on the model.
CategoryMiniMax H3/Edit

Inputs (5)

NameTypeDefaultDescription
imageIMAGEReference image(s) appended in batch order as the next <Picture N> entries.
transportCOMBOsemantic (Qwen only)Semantic skips the guide VAE; native also creates a minimax_refs latent.
semantic_resolutionINT1024256–3584Equivalent-square Qwen budget used when transport is semantic.
native_reference_sizeCOMBOmatch output areaVAE resize policy used when transport is native.
previous_referencesoptH3EDIT_REFERENCE_STACKConnect the previous Add H3 Edit Reference node to keep appending in order.

Outputs (2)

NameTypeDescription
referencesH3EDIT_REFERENCE_STACK—
infoSTRING—