Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 Five-View Character Sheet (EXP/T8)
ComfyUI Node

MiniMax H3 Five-View Character Sheet (EXP/T8)

H3's turnaround sheet, sampled in a single pass

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 Five-View Character Sheet (EXP/T8)
  • clip
  • video_vae
  • ref_image
  • positive
  • five_view_latent
  • conditioned_prompt
  • report_json
◄promptThe camera orbits the subject of <Picture 1> clockwise; five coordinated views, consistent identity and clothing.►
◄size512►

Turnaround sheets have been the community's answer to character consistency for years: don't fight for consistency across five separate generations, generate all the views in one context and use the sheet as ground truth afterwards. MiniMaxH3FiveViewConditioningEXPT8 moves that idea into MiniMax H3 - one reference image in, five jointly-sampled views out.

The important mechanical detail is that H3's five views are latent image slots, not video frames. They're denoised together in one pass, which is what buys you shared identity, and they're decoded separately afterwards. That's why it needs a partner node to get pictures out (see the Five-View Decode article - this node does not give you images).

What it actually builds

Give it a single reference - ref_image is exactly one image, and it becomes <Picture 1> in native MiniMax reference conditioning. The node does an aspect-preserving, down-only resize, then builds a video latent shaped for five slots (24 channels, 5 positions) plus an audio latent sized for two tokens. size is the square size of each image, default 512, not the total strip width.

The prompt field comes pre-filled with something useful rather than empty, which tells you the intent: a coordinated orbit around the subject with consistent identity and clothing. Keep that structure - the views have to be related for the sheet to be worth anything.

Outputs are positive (CONDITIONING, which goes to your sampler), five_view_latent (LATENT - this goes to the decode node, not to a normal VAE decode), conditioned_prompt and report_json.

The prerequisite everyone misses

This is not a stock H3 trick. It needs:

  1. An H3 Ref2VA pruned base, plus the Qwen3-VL CLIP and the H3 video VAE.
  2. A dedicated turnaround LoRA loaded with the stock LoraLoaderModelOnly at strength 1.0 - minimax_h3_five_view_512_s1500.safetensors from the matlod turnaround repo, dropped into models/loras. The pack does not ship it.
  3. Stock res_multistep / simple sampling for 28 steps, with the pack's sampling setup supplying the native H3 AV sigmas and sampler.

The example workflow is examples/workflows/03-image-video-edit/2026-09-21_H3_Five_View_Character_Sheet_EXP.json in the pack. Swap the placeholder image for one you have the rights to.

Install

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Full quit and restart, then refresh the browser. The pack's requirements.txt installs no packages deliberately, so nothing here will touch your Torch build. Model layout for this node: H3 base in models/diffusion_models, Qwen3-VL in models/text_encoders, H3 video VAE in models/vae, turnaround LoRA in models/loras.

Where it bites

Turn the LoRA off when you go back to video. The author documents degraded motion when the turnaround LoRA is left connected during normal video generation. It's a sheet-making adapter, not a general accelerator; treat it as a switch you flip on for this graph and off for the next.

Don't pass the five-view latent to a normal decoder. The decode node refuses latents without the protocol marker, and going through the ordinary Still or video decoder is not equivalent. If a custom sampler copies the latent and drops the marker, decode will reject it - stock SamplerCustomAdvanced preserves it.

VRAM is not small here. The author's own isolated run on an RTX 4060 Ti 16 GB sampled 15,574 MiB used / 536 MiB free at one point during sampling. That's one observation, not a peak measurement, but it means "it's just stills" is the wrong intuition - five slots denoised together plus a LoRA stack is heavy.

A sheet is a candidate, not a fact. The validated test produced visible viewpoint and pose progression, but the docs are explicit that it doesn't guarantee a proper standard-angle turnaround, and that on the public test image - a wide scene with a small subject - the subject stayed small. Feed it a decent close-up reference. Then look at the sheet before adding its views to a project's shared reference library.

CategoryT8/MiniMax H3/Still/Experimental

Inputs (5)

NameTypeDefaultDescription
clipCLIPNative MiniMax H3 Qwen3-VL CLIP.
video_vaeVAEMiniMax H3 video VAE.
promptSTRINGThe camera orbits the subject of <Picture 1> clockwise; five coordinated views, consistent identity and clothing.—
ref_imageIMAGEExactly one reference image, <Picture 1>.
sizeINT512512–2048Square size of each of five images, not the total strip size.

Outputs (4)

NameTypeDescription
positiveCONDITIONING—
five_view_latentLATENT—
conditioned_promptSTRING—
report_jsonSTRING—