Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 Five-View Decode (EXP/T8)
ComfyUI Node

MiniMax H3 Five-View Decode (EXP/T8)

Why you can't just VAE-decode a character sheet

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 Five-View Decode (EXP/T8)
  • five_view_latent
  • video_vae
  • views_batch
  • horizontal_sheet
  • report_json

The five-view sheet from the previous node is a latent with five image slots packed where the H3 video VAE expects a time axis. Hand that to a normal VAE decode and you don't get five pictures - you get whatever the decoder makes of a sequence it was never given. MiniMaxH3FiveViewDecodeEXPT8 is the two-input node that turns that latent into images properly.

What "properly" means

H3's VAE is a video VAE. It expects clips of a legal temporal length - in practice two latent tokens. So the decode node takes each of the five slots, duplicates that single latent token into a legal two-token clip, decodes that clip on its own, and keeps the first decoded pixel frame. Five separate decodes, five stills.

That's slower than one batched decode and it's the honest way to do it: the alternative is lying to the VAE about the structure of the latent. It also explains the marker check - the node rejects any latent lacking the dedicated five-view protocol marker, which is a deliberate guardrail rather than an inconvenience. If you feed it ordinary video latents it will stop you, and that's the desired behaviour.

Inputs and outputs

Two inputs, both non-negotiable: five_view_latent and video_vae - the H3 video VAE (the FP16 one in models/vae).

Three outputs:

  • views_batch - five images as a batch, ready for a Save Image node's batch handling or an image-comparison node
  • horizontal_sheet - the five views stitched left-to-right into one strip. At size: 512 that's a 2560×512 sheet; this is the one you feed onward as a reference asset.
  • report_json - the receipt, including what got decoded and at what geometry.

Saving the views_batch plus the strip is the practical move: the batch is what you keep as individual references, the strip is what you look at to decide whether the turnaround is coherent enough to keep.

Install

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Restart ComfyUI fully, refresh the browser. No pipeline dependencies of note - the pack's requirements.txt installs nothing, and this node leans entirely on the video VAE ComfyUI loads for you. You do need the H3 video VAE in models/vae; there's no plumbing here that fetches it.

Where it bites

"No marker" errors are almost always a sampler problem upstream. If you replaced stock SamplerCustomAdvanced with a custom sampler that rebuilds the latent dict instead of copying it, the marker can be dropped on the way through. The pack verified against installed Core source that the stock sampler preserves it. So when the decode node refuses your latent, check what sampled it before you suspect the decode.

Don't decode it twice. Passing the same latent through both this node and a normal VAE decode wastes a lot of time and produces garbage from one of them. The sheet path is this node, full stop.

Five decodes means five chances to OOM on a small card. Each decode is small, but you're doing them in sequence with the model already resident, and the conditioning side of this workflow is the heavy half. If you're tight on VRAM, decode after the model has been unloaded, not during.

The strip is a reference, not a standard. A sheet with five related angles is exactly what the community's character-consistency workflows have been using turnaround sheets for - as a stable reference to feed downstream, not as a finished asset. Inspect it first; the pack is explicit that its own validated sample showed viewpoint progression without guaranteeing a textbook turnaround.

CategoryT8/MiniMax H3/Still/Experimental

Inputs (2)

NameTypeDefaultDescription
five_view_latentLATENT—
video_vaeVAE—

Outputs (3)

NameTypeDescription
views_batchIMAGE—
horizontal_sheetIMAGE—
report_jsonSTRING—