Extensions/MiniMax H3 LongMedia
ComfyUI Extension

MiniMax H3 LongMedia

Long-form MiniMax H3 video/audio generation for ComfyUI with unified clip planning, continuity, lip-sync, and adaptive low-VRAM execution.

By vizart-vj·Created about a month ago·Updated 11 days ago· 234
vizart-vj/ComfyUI-MiniMax-H3-LongMedia
Nodes34
On cloudLocal install
CategoryMiniMax H3/LongMedia/LongMedia, MiniMax H3/LongMedia/Streams
Stars234
Updated11 days ago

Nodes (34)

MiniMax H3 • Low-VRAM Attention Chunking (internal)

The low-VRAM attention knob ComfyUI hides from you

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • Encode Audio Stream

Get your soundtrack into H3's latent space, properly

MiniMax H3/LongMedia/Streams
MiniMaxH3LatentLabBlockMemoryTracer

Find out which H3 transformer block ate your VRAM

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • First-Step Memory Profiler (internal)

See exactly what your first denoise step costs in VRAM

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • AV Latent Info

A health check for H3's weird two-stream latent

MiniMax H3/LongMedia/Utility
MiniMax H3 • LipSync Latent Setup

The honest way to do lip-sync in latent space

MiniMax H3/LongMedia/Utility
MiniMax H3 • Long Media Decode

Where a long H3 run finally becomes video you can watch

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • Long Media Next Segment

Handing H3 the previous clip so the next one matches

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • Long Media Sampler

The sampler that turns one H3 prompt into a whole long clip

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • Long Media Setup

The one node that decides what your H3 video will be

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • Merge AV Latents

Blend a new video stream into an old clip without touching its audio

MiniMax H3/LongMedia/Streams
MiniMax H3 • Low-VRAM MLP Chunking (internal)

The other half of H3's low-VRAM story (the MLP half)

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • Pack AV Streams

The tiny node that makes H3's two streams one

MiniMax H3/LongMedia/Streams
MiniMax H3 • Prepare Continuation

Grow a clip by opening the next one with the last one's ending

MiniMax H3/LongMedia/Continuation
MiniMaxH3LatentLabProtectRefineAV

The guard that stops your refine pass from wrecking the seam

MiniMax H3/LongMedia/LongMedia
MiniMaxH3LatentLabRefineSigmas

Cut your sigma schedule in two so the refine tail doesn't fight the base pass

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • Replace Audio Stream (deprecated)

A deprecated node that still does its job (if you have old graphs)

MiniMax H3/LongMedia/Streams
MiniMax H3 • Replace Stream

Swap one half of an H3 clip without regenerating the other

MiniMax H3/LongMedia/Streams
MiniMax H3 • Replace Video Stream (deprecated)

The deprecated video-swap node old graphs still load with

MiniMax H3/LongMedia/Streams
MiniMax H3 • Runtime Continuation Guider

The guider that finally knows what the previous clip looked like

MiniMax H3/LongMedia/LongMedia
MiniMaxH3LatentLabSeededDisableNoise

The no-noise pass that still remembers the seed

MiniMax H3/LongMedia/Internal
MiniMax H3 • Split AV Streams

Splitting H3's joint video+audio latent into two editable streams

MiniMax H3/LongMedia/Streams
MiniMax H3 • Stitch Continuation

Joining H3 continuation segments without a visible seam

MiniMax H3/LongMedia/Continuation
MiniMax H3 • Stream Denoise Controls

Keep the audio, regenerate the picture — H3 streams on separate dials

MiniMax H3/LongMedia/Streams
MiniMaxH3LatentLabUltraPinnedMemoryGate

Turning off pinned memory so the huge H3 model can page itself

MiniMax H3/LongMedia/LongMedia
MiniMaxH3LatentLabUltraPinnedMemoryRestore

Putting ComfyUI's pinned-memory setting back where you found it

MiniMax H3/LongMedia/LongMedia
MiniMaxH3LatentLabUnifiedRuntimeSampler

The engine that runs every H3 segment in a single sampling lifecycle

MiniMax H3/LongMedia/Internal
MiniMax H3 • Encode Video Stream

Turning frames into an H3 video stream the model can actually eat

MiniMax H3/LongMedia/Streams
MiniMax H3 • Video Inpaint

Inpainting inside H3's video latent, not on the pixels

MiniMax H3/LongMedia/Streams
MiniMax H3 • VRAM Cache Cleanup (internal)

The post-run VRAM flush that tells you how much it clawed back

MiniMax H3/LongMedia/LongMedia
MiniMax H3 • VRAM Pressure Guard (internal)

A SAMPLER wrapper that flushes VRAM before you run out, not after

MiniMax H3/LongMedia/LongMedia
MiniMax H3 LongMedia Cameras

Tell MiniMax H3 how to move the camera without the camera ending up in the shot

MiniMax H3/LongMedia/LongMedia
MiniMax H3 LongMedia Planner

A prompt, duration and seed per clip

MiniMax H3/LongMedia/LongMedia
MiniMax H3 LongMedia Video Reconstructor

Restore long, low-quality video with MiniMax H3 — without VRAM scaling to the runtime

MiniMax H3/LongMedia/LongMedia
Readme

ComfyUI-MiniMax-H3-LongMedia

Production-oriented ComfyUI nodes for MiniMax H3 long-form video/audio generation, native reference editing, MultiClip planning, camera direction, fixed segmentation, lip-sync/redubbing, latent hi-res refinement, and adaptive low-VRAM execution.

screenshot

Current release: 0.5.40

Main Nodes

  • MiniMax H3 • Long Media Setup
  • MiniMax H3 • Long Media Planner
  • MiniMax H3 • Long Media Cameras
  • MiniMax H3 • Long Media Sampler
  • MiniMax H3 • Long Media Decode
  • MiniMax H3 • Long Media Video Reconstructor

Legacy internal MiniMaxH3LatentLab... class identifiers remain registered for workflow compatibility.

Installation

Install into:

ComfyUI/custom_nodes/ComfyUI-MiniMax-H3-LongMedia

or install the package from the Comfy Registry.

Package identity:

GitHub:             vizart-vj/ComfyUI-MiniMax-H3-LongMedia
Comfy PublisherId:  noise

Restart ComfyUI after installation/update.

Current Setup Model

New workflows use independent semantic controls:

control_mode
h3_mode
timeline_mode
duration_source
audio_mode

This is the main change in how the project should be understood compared with the public 0.4.40 documentation.

See Operating Modes.

H3 Conditioning

t2va
fl2va
ref2va
hybrid
video_ref_edit

Timeline

single
segmented
multiclip

Duration Ownership

auto
video
audio
manual
longest_input

duration_source controls timeline length only. It does not remove audio references or change final-audio policy.

video_ref_edit

Typical source-character replacement:

video_1 = source video frames
image_1 = replacement identity
audio_1 = source soundtrack or new dub

A Video input is an IMAGE batch and never contains soundtrack data.

For preserve-style modes, Video1+Audio1 can be presented as a native paired source-performance reference while Audio1 also owns the target timing/output waveform.

For audio_mode=lip_sync, Audio1 is intentionally independent from Video1's original facial performance so completely new dialogue or singing can drive the replacement character.

Audio2/Audio3 remain prompt-addressable references for music, percussion, bass, ambience, or other semantic timing. In video_ref_edit, they are conditioning references only; Audio1 remains the sole preserved/passthrough source soundtrack.

See Audio Modes and video_ref_edit.

MultiClip + Cameras

Recommended connection:

Long Media Planner
        ↓ clip_plan
Long Media Cameras
        ↓ clip_plan
Long Media Setup

Use:

timeline_mode = multiclip

Planner owns diegetic scene/action prompts, durations, names, and optional seeds. Cameras owns framing, rig/lens, movement, speed, spatial relation, entity continuity, and transitions.

See:

Segmented Long Form

Use:

timeline_mode = segmented

for one continuous semantic movie split into fixed-duration internal units for VRAM/stability. It is not a storyboard scheduler.

See Fixed Segmentation Prompting.

Two-Stage H3 / Latent Hi-Res

The Long Media Sampler can:

  1. generate a low-resolution Stage-1 denoised x0;
  2. learned-upscale the video latent only;
  3. preserve the audio latent;
  4. rebuild target-grid conditioning;
  5. optionally run an independent same-seed fresh-noise high-resolution H3 pass.

Without Latent Hi-Res, the Refiner remains a continuous zero-noise low-sigma tail.

See Two-Stage Sampling, Latent Hi-Res and Refiner.

Sampler / Memory

Production starting point:

sampler_mode   = auto
memory_mode    = auto
attention_mode = auto

Keep ComfyUI Dynamic VRAM enabled.

Current memory-safety work includes exact Comfy Kitchen query streaming for structurally impossible fused-QKV workloads, repeat-run memory isolation, guarded native INT8 VBAR prefetch on constrained GPUs, and RAM-pressure-aware pinned host memory.

See Sampler, VRAM and Performance Guide.

FastH3 / FastVideo VSA

LongMedia includes isolated compatibility paths for supported H3ddle/PulpCut FastH3 VSA and Kijai FastVideo VSA packages.

These paths use strict structural detection and reset their runtime state when switching back to ordinary H3 checkpoints.

Loop Closure

Loop Closure returns the generated tail toward the opening macro-state in latent/H3 space. It is independent from conditioning/timeline mode and does not use an RGB crossfade as its primary mechanism.

Documentation

Start at docs/README.md.

Important guides:

Example Workflows

  • workflows/MiniMax-H3-LongMedia-SAFE-1080p-15s.json
  • workflows/MiniMax-H3-LongMedia-LatentUpscale-Detailer.json

Release workflows contain neutral media placeholders rather than user media or local preview paths.

See workflows/README.md.

Third-Party Code

This repository contains adapted third-party components under their respective licenses, including:

  • Saganaki22/ComfyUI-sol-attn-derived code under Apache-2.0;
  • MiniMax H3 latent-upscaler-derived code under Apache-2.0.

See THIRD_PARTY_NOTICE.md and THIRD_PARTY_APACHE_2_0.txt.

Release History

0.5.40 is the release consolidation after public v0.4.40.

Historical release notes remain in docs/ for compatibility/reference. They may use the old workflow_mode terminology.

License

See LICENSE.