Nodes/MiniMax H3 Edit/Decode H3 Scene Coverage
ComfyUI Node

Decode H3 Scene Coverage

Walk a Camera Around a Room You Photographed Once

By ethanfel·Created about a month ago·Updated 19 days ago· 30
Decode H3 Scene Coverage
  • samples
  • vae
  • contact_sheet
  • selected_views
  • all_frames
  • info
◄columns4►
◄gutter_px6►
◄gutter_colorblack►

The character-sheet trick - orbit the camera, keep the settled frames - generalises past people. Point it at a rigid room or a location and you get something a photographer would charge for: 6 to 24 consistent viewpoints of one space, from a single source photograph.

That's the pitch, and it's more useful than it sounds. Multi-angle background plates, a location set your later edits can reference without re-describing the room, a storyboard pass for a scene you can't revisit. The alternative is reshooting, or six images that are six different rooms.

How the pair works

scene coverage | canonical camera path in H3 Edit Options hands the encoder a frozen-scene camera compiler plus a 124-frame path, 12 views, a 360° clockwise arc and five-frame holds per capture. The compiler writes exact timed camera waypoints and, critically, static capture windows into the latent - each view has a start and end frame that bracket the moment the camera has settled.

The decoder then does the other half: it decodes the path, and for every hold window it runs the pack's frame-quality scoring (sharpness, contrast, exposure, temporal stability) inside that window and keeps the best frame. You never pick a frame; you pick how many viewpoints you want.

This node refuses to run on the wrong latent. If the samples you feed it didn't come from a scene-coverage profile, there are no capture windows and it tells you so by name instead of guessing. Same for a malformed window or a path shorter than the encoded frames.

Inputs and outputs

Required: samples, vae, columns (1–8, default 4 - contact sheet columns), gutter_px (0–256, default 6) and gutter_color (black, neutral gray, white).

Outputs: contact_sheet (every viewpoint in one image), selected_views (the viewpoints as an IMAGE batch in camera-path order - this is the one you save when you want them individually), all_frames (the entire decoded path, for when a hold landed awkwardly and you'd rather grab a neighbour), and info, which reports capture targets, chosen frames, angles, windows, loop mode and sheet dimensions. Read info once per new configuration; it's the only place the chosen frames and angles are visible.

Nothing about the path is set here. Views, arc, direction, hold length and loop closure live on H3 Edit Options (or the encoder's coverage widgets if you're not using the options node), and they have to be decided before sampling, because they're compiled into the prompt and the timing.

Choosing settings that work

| Profile | Duration at 24 fps | Sensible view count | |---|---:|---| | 124 frames | 5.17 s | 6–8 views | | 243 frames | 10.13 s | 8–12 views | | 362 frames | 15.08 s | 12–16 views |

Up to 24 captures are permitted, but more views means shorter travel and hold windows, and a short hold gives the decoder less to choose from. 362 frames is the safer pick for dense coverage; 243 is where the H3_Art_Room_Precise_Hard_Cuts.json example lives, with an editable camera plan connected through compiled_prompt.

Two other modes skip the orbit. cinematic hard cuts composes eight static shots around one named target, the prompt explicitly forbidding any camera travel. room + object study is the all-semantic route: 362 frames, 16 hard-cut stills, six room-establishing views then ten contextual views of one named object, with every connected picture coerced to semantic transport even if it carries a stale native setting.

And a reality check from the author's own README: one photograph contains no ground truth for surfaces the camera never saw, so those areas are a conservative reconstruction. Real alternate angles genuinely help - which is what Add H3 Edit Reference is for.

Install

Manager → search the pack title, or:

cd ComfyUI/custom_nodes
git clone https://github.com/ethanfel/ComfyUI-MiniMax-H3-Edit

Restart. Zero third-party dependencies by design; you do need ComfyUI's native MiniMax H3 support plus the Qwen3-VL encoder, the H3 video VAE and an FL2VA or REF2VA diffusion model.

For a 360° anchored room with loop closure on, you only load the source image once - the encoder internally presents it as both the first and final keyframe. For a partial arc, only the first frame is anchored.

Where people get burned

  • Decoding a still latent with this node. It needs the scene metadata. If your graph is in still | edit or generate mode, you want Decode H3 Edit to One Image instead.
  • Changing the view count after sampling. The capture windows are baked into the latent and the prompt during encoding. Raise coverage_views afterwards and you're just decoding a graph that doesn't match.
  • Long paths on tight VRAM. 362 frames of H3 latent plus a full video VAE decode in one go is the heaviest thing this pack does.
  • Expecting the model to invent a second room. Alternate-angle references constrain occluded walls, openings, fixtures and materials in one shared coordinate system. Semantic Qwen-only angles are the recommended first try; native alternate views mixed with FL2VA keyframes are still experimental territory.
CategoryMiniMax H3/Edit

Inputs (5)

NameTypeDefaultDescription
samplesLATENTSampled latent from frozen scene coverage mode.
vaeVAEMiniMax H3 video VAE.
columnsINT41–8Contact-sheet columns.
gutter_pxINT60–256—
gutter_colorCOMBOblack3 options: black, neutral gray, white

Outputs (4)

NameTypeDescription
contact_sheetIMAGEOne contact sheet containing every selected scene viewpoint.
selected_viewsIMAGEThe selected viewpoints as an IMAGE batch in camera-path order.
all_framesIMAGEEvery decoded camera-path frame for inspection or alternate picks.
infoSTRINGCapture targets, chosen frames, angles, windows, loop mode, and sheet dimensions.