Nodes/MiniMax H3 Edit/Text Encode H3 Edit / Generate
ComfyUI Node

Text Encode H3 Edit / Generate

One Encoder, Three Very Different Jobs for Picture 1

By ethanfel·Created about a month ago·Updated 19 days ago· 30
Text Encode H3 Edit / Generate
  • clip
  • vae
  • source_image
  • reference_image
  • reference_stack
  • options
  • positive
  • latent
  • fitted_source
  • encoded_prompt
  • info
◄promptAdd the glasses from <Picture 2> to the woman.►
◄primary_image_roleedit | strong scene anchor (FL2VA)►
◄reference_modesemantic (Qwen only)►
◄width768►
◄height1344►
◄source_fitcrop center►
◄prompt_modeedit instruction►
◄semantic_resolution1024►
◄native_reference_sizematch output area►
◄quality_profilerecommended | 5-frame context -> 1 image►
◄coverage_views12►
◄coverage_arc_degrees360►
◄coverage_directionclockwise / camera right►
◄coverage_hold_frames5►
◄coverage_loop_closuretrue►
◄compiled_prompt—►

This is the brain of the pack. It's a text encoder in the sense that it tokenizes your prompt, but what it really does is decide what your input image means to MiniMax H3 - and that single decision changes which checkpoint you need loaded.

MiniMax H3 is a 33B omni-modal video model that ComfyUI supports natively. Instruction-editing style work is what the community actually does with these models in 2026 - you have an image you like, you want one thing changed, and you don't want the model re-inventing the composition. H3 Edit does that on the video model's own conditioning contract rather than through an API.

The switch that changes everything

primary_image_role has three values and they are not cosmetic:

  • edit | strong scene anchor (FL2VA) - <Picture 1> goes through Qwen and the H3 VAE, and is attached as a frame-zero keyframe. The scene, geometry and composition are locked. This is the one to use for "add these glasses to this woman."
  • generate | semantic Picture 1 (FL2VA) - the keyframe is removed entirely. Picture 1 is a design reference: Qwen visual tokens, nothing pixel-anchored. The scene gets created from scratch.
  • generate | native Picture 1 (REF2VA) - same as above but Picture 1 also becomes a minimax_refs block, which needs REF2VA weights instead of FL2VA.

Both generation routes drop minimax_keyframes from the conditioning and switch the prompt compiler to a creation contract. Your prompt stops being an edit instruction and starts being an art direction brief. People miss this and then wonder why their carefully anchored edit came back as a different room.

Alongside it, reference_mode sets the transport for the optional direct reference_image - the convenient <Picture 2> socket. semantic (Qwen only) transfers an idea, native (Qwen + VAE ref) adds a VAE-encoded reference block for stronger low-level matching, none (source only) disables the socket even if something is plugged into it. For more than one guide, chain Add H3 Edit Reference nodes into reference_stack.

There's one more source of whole-graph ambiguity worth knowing: prompt_mode. The directed and character-sheet compilers rewrite your short instruction into a timed H3 prompt with motion phases and preservation rules. If you want your text to go through untouched, that's use prompt verbatim - but the Qwen picture blocks are still prepended in order, so <Picture N> tags still work.

Inputs, outputs, and the wiring

Required: clip (the H3 Qwen3-VL encoder), vae (the H3 video VAE), source_image, prompt, primary_image_role, reference_mode, width and height (defaults 768x1344, multiples of 32), source_fit (crop center, contain / pad, stretch - how anchored Picture 1 is fitted to the output canvas), prompt_mode, semantic_resolution and native_reference_size.

Optional: reference_image, quality_profile, reference_stack, the five coverage controls, options (connect H3 Edit Options here - it hides all the legacy task widgets), and compiled_prompt.

compiled_prompt is the one for tinkerers. Connect a STRING from anywhere - an external prompt compiler, another pack's editor - and it replaces only the built-in prompt construction. Your prompt widget stays untouched, and the encoder still prepares visual references, Qwen picture ordering, native reference blocks, keyframes, latent timing and decoder metadata. Disconnect it and you're instantly back to the built-in compiler with your original prompt intact.

Outputs: positive (CONDITIONING), latent (the H3 latent with your selected temporal profile), fitted_source (Picture 1 after preparation - handy to preview so you can see what the model actually got), encoded_prompt (the exact compiled text after the visual token blocks - read this the first time you use each mode, it teaches you more than any guide), and info, which reports transport, Qwen size, VAE usage and a checkpoint hint. That hint is worth wiring to a text preview node: it tells you whether you should be on FL2VA or REF2VA.

The encoder also stamps the decoder metadata onto the latent - frame count, capture windows, character-sheet profile. That's why the decoders throw precise errors when handed a latent that didn't come from the matching profile.

Install and model setup

cd ComfyUI/custom_nodes
git clone https://github.com/ethanfel/ComfyUI-MiniMax-H3-Edit

Restart. No third-party Python packages - the pack's requirements file deliberately installs nothing beyond ComfyUI's own runtime. You do need a current ComfyUI build with native MiniMax H3 support plus the Qwen3-VL text encoder, the H3 video VAE, and at least one of the FL2VA/REF2VA diffusion models.

Sampling stays a normal native H3 graph: ModelSamplingMiniMaxH3, BasicGuider, RandomNoise, a sampler and scheduler, SamplerCustomAdvanced. The included example_workflows/H3_Edit_Mixed_References.json is a working single-image edit you can swap your own images into.

Where people get burned

  • Wrong checkpoint for the role. Native Picture 1 or native guides want REF2VA. Running those against FL2VA weights is a "works but behaves oddly" failure, not a crash - swap the diffusion model in your loader.
  • Expecting the anchor to survive a generation role. It doesn't. That's the feature.
  • Mixed semantic/native guides. The packed layout can carry it, but the released weights were trained for different task presentations. Semantic-only is the reliable path; treat mixed stacks as experiments.
CategoryMiniMax H3/Edit

Inputs (22)

NameTypeDefaultDescription
clipCLIPMiniMax H3 Qwen3-VL text/vision encoder.
vaeVAEMiniMax H3 video VAE used by anchor and native-reference modes.
source_imageIMAGEPicture 1: the photo to edit in anchor mode, or the first semantic/native reference in generation mode.
promptSTRINGAdd the glasses from <Picture 2> to the woman.The requested edit or new image. Use explicit <Picture N> roles for references.
primary_image_roleCOMBOedit | strong scene anchor (FL2VA)Edit creates a frame-zero VAE keyframe. Generate removes that keyframe and treats Picture 1 as either a semantic FL2VA or native REF2VA reference.
reference_modeCOMBOsemantic (Qwen only)Semantic sends Picture 2 only through Qwen. Native also VAE-encodes it into minimax_refs. None disables Picture 2 even if its socket remains connected.
widthINT76832–16384—
heightINT134432–16384—
source_fitCOMBOcrop centerHow Picture 1 is fitted to the output canvas before native VAE anchoring.
prompt_modeCOMBOedit instructionChoose ordinary edit/generation, an anchored directed transformation, frozen-scene camera coverage, or verbatim text. Scene coverage requires a matching 124/243/362-frame profile.
semantic_resolutionINT1024256–3584Equivalent-square Qwen pixel budget for semantic direct or Picture 1 generation refs. Aspect ratio is preserved and no VAE latent is allocated.
native_reference_sizeCOMBOmatch output areaResize policy for native direct or Picture 1 generation-reference VAE conditioning.
reference_imageoptIMAGEOptional next Picture guide. Choose semantic or native transport above.
quality_profileoptCOMBOrecommended | 5-frame context -> 1 imageH3 is video-trained. Recommended matches Studio's short 5-frame context, then the decoder returns one stable frame. Character-sheet profiles create calibrated 73/124/171-frame sequences for the dedicated sheet decoder. Scene-coverage profiles create a trained-range camera path for the dedicated coverage decoder. True 1-frame mode is often poor quality.
reference_stackoptH3EDIT_REFERENCE_STACKOptional ordered guides from chained Add H3 Edit Reference nodes. These follow the direct reference_image and become the next <Picture N> entries.
coverage_viewsoptINT122–24Number of unique scene viewpoints and output images for frozen scene coverage.
coverage_arc_degreesoptFLOAT36015–360Physical camera arc from the source/generated opening viewpoint.
coverage_directionoptCOMBOclockwise / camera rightDirection in which the physical camera travels around the declared orbit center.
coverage_hold_framesoptINT51–9Requested static frames around each capture; automatically reduced if views are close.
coverage_loop_closureoptBOOLEANtrueFor a 360-degree anchored room, internally reuse the one source image as the final keyframe. Ignored for partial arcs and reference-generated rooms.
optionsoptH3EDIT_OPTIONSRecommended: connect H3 Edit Options to choose one coherent task preset and hide all legacy task-specific widgets on this encoder.
compiled_promptoptSTRINGOptional complete H3 prompt from an external compiler profile. When connected, this bypasses only the built-in prompt compiler; the prompt widget remains unchanged, and reference images, Qwen ordering, latent timing, and decoder metadata are still prepared by this encoder. The connected text must already be a complete H3 prompt.

Outputs (5)

NameTypeDescription
positiveCONDITIONINGPositive conditioning with the selected Picture 1 role and every ordered guide transport.
latentLATENTAn H3 latent with the selected still, character-sheet, or frozen-scene temporal profile.
fitted_sourceIMAGEPicture 1 after preparation for its selected anchor/reference role.
encoded_promptSTRINGThe exact prompt sent after the visual token blocks.
infoSTRINGReference transport, Qwen size, VAE usage, and checkpoint guidance.