Nodes/H3 Continuum/H3 Continuum Sampler V3.5
ComfyUI Node

H3 Continuum Sampler V3.5

Chunked long-form MiniMax H3 in one node

By ukr8b3g-cmyk·Created 2 months ago·Updated about 15 hours ago· 113
H3 Continuum Sampler V3.5
  • model
  • clip
  • video_vae
  • sampler
  • sigmas
  • first_frame
  • last_frame
  • reference_image_1
  • reference_image_2
  • reference_image_3
  • reference_video_1
  • driving_audio
  • audio_vae
  • reference_audio_1
  • reference_audio_vae
  • video_latents
  • audio_latents
  • assembly_plan
  • status
  • driving_audio
  • refine_context
◄sequence_prompt—►
◄prompt_modeAuto►
◄chunks3►
◄chunk_seconds5.0►
◄width1344►
◄height768►
◄continuityBalanced — 22 frames►
◄base_seed0►
◄audio_continuitytrue►
◄diagnosticsBasic►
◄reroll_from_chunkAuto►
◄reroll_nonce0►
◄strict_compatibilitytrue►
◄debugfalse►
◄show_previewtrue►
◄run_storageOff►
◄run_name►
◄reference_sizeMatch Output►
◄project_id►
◄video_reference_sizeEfficient - 0.4 MP►

This is the node everything else in the pack hangs off. H3 Continuum Sampler V3.5 is how you generate long-form MiniMax H3 video in ComfyUI: instead of one bounded pass, it splits the run into sequential chunks - 5 seconds per chunk is the validated default - carries visual and audio context across each boundary, and lets you keep a supplied video or audio track in the loop. If you've been frustrated that single-pass H3 tops out around fifteen seconds, this is the workaround that actually holds continuity.

MiniMax H3 is the 33B omni-modal model that generates video with native stereo audio, and it got day-zero ComfyUI support - but its native single-pass path is short. Continuum's bet is chunking with preserved context: at every boundary it carries a "continuity" window of prior video (the default is Balanced - 22 frames; Auto is conservative, Fast trims to 5, Strong's 39 frames is marked Experimental), so the next chunk starts from where the last one ended rather than from scratch. For 5-second FL2VA runs it also does what it calls Terminal Merge: the final two logical chunks become one physical sampling-and-decode unit so Last Frame handling matches Core behavior. Outputs come back per physical group, which is why several of the V3.5 nodes return lists.

The inputs a beginner actually sets

The sampler has a lot of widgets. Here's what matters:

  • model, clip, video_vae, sampler, sigmas - the standard H3 stack. The VAE is used only to encode image conditioning; Continuum never decodes with it.
  • sequence_prompt - one Text (Multiline) for the whole run, plus prompt_mode (Auto / Fixed / List / Timeline).
  • chunks (default 3) and chunk_seconds (default 5) - your total length is roughly chunks × chunk_seconds.
  • width / height - multiples of 32; defaults are 1344×768.
  • continuity, base_seed, audio_continuity (on passes prior audio context into continuation chunks).
  • reroll_from_chunk + run_storage (Save + Auto Resume) - for reusing completed chunks and regenerating only the part you don't like.

Everything else is optional connectivity: first_frame, last_frame, reference_image_1 through reference_image_3 (persistent identity refs), reference_video_1 (that's the Video Guide Frames socket - feed it a video loader's IMAGE frame batch, not a still), driving_audio plus audio_vae (your audio track, preserved as the final audio), and reference_audio_1 plus reference_audio_vae (H3 conditioning only, does not replace the output audio). Leave all image/audio inputs disconnected and you're doing text-to-video.

Outputs

video_latents and audio_latents (lists), assembly_plan (the map everything downstream reads), status, driving_audio (your preserved track, if connected), and - new in V3.5 - refine_context, which is what makes Hi-Res Fix and Second Pass context-aware. Wire the first three into Core VAE Decode + Assemble + Seam, or the last one into the refine nodes.

Installing it

ComfyUI Manager, search "H3 Continuum", or:

cd ComfyUI/custom_nodes
git clone https://github.com/ukr8b3g-cmyk/ComfyUI-H3-Continuum.git

Restart after install. No pip deps beyond ComfyUI's bundled torch. The H3 weights (~42.5 GB) install the normal way, and be aware of the MiniMax H3 Community License's territory restrictions. For First/Last + Reference in one run you want a hybrid-capable model, e.g. the B2049 variant the README links.

Where people get burned

  • Timeline prompts: the time header must be on its own line. [0-5s] then the description on the next line - [0-5s] Describe the scene isn't recognized and Auto falls back to Fixed, silently reusing one prompt for every chunk. Text before the first header becomes a global preamble applied to all chunks.
  • Don't rely on "continue from the previous chunk." Continuum carries context, but describing the next action explicitly gives H3 a much better shot at not drifting or cutting.
  • Video Guide Frames is a guide, not a copy, and it expects 24 fps input (set force_rate to 24 on the loader). It won't reproduce every frame.
  • Driving Audio vs Reference Audio: driving audio replaces the final output's audio; reference audio only conditions it. Connect the wrong one and you'll wonder why your generated sound is still there (or isn't).
  • Run Storage needs a fixed seed for reproducible resume, and changing models, references, prompts, or resolution creates a new revision, so earlier chunks stop being reusable.
  • No universal memory guarantee: the README notes two-chunk 800×800 runs pass on a 16 GB card, but long high-resolution sequences can still exceed system RAM during decode and assembly.
CategoryMiniMax H3/Continuum

Inputs (35)

NameTypeDefaultDescription
modelMODEL—
clipCLIP—
video_vaeVAEUsed only to encode image conditioning. T2VA does not use it; Continuum never decodes with it.
samplerSAMPLER—
sigmasSIGMAS—
sequence_promptSTRINGConnect one Text (Multiline) for the complete sequence.
prompt_modeCOMBOAutoAuto accepts Fixed, list-separated, and timeline prompt styles.
chunksINT31–16Number of sequential Continuum chunks to generate.
chunk_secondsFLOAT5.04–30Target duration shared by every chunk. 5–15 seconds is the recommended and validated range. Values above 15 seconds are supported, but VRAM use and processing time can increase substantially, especially at high resolution.
widthINT134432–16384Output width. Use a multiple of 32.
heightINT76832–16384Output height. Use a multiple of 32.
continuityCOMBOBalanced — 22 framesAmount of prior video context retained at each chunk boundary.
base_seedINT00–18446744073709550000Base seed used to derive deterministic per-chunk seeds.
audio_continuityBOOLEANtrueOn passes prior audio context into continuation chunks. Turn it off only to isolate or replace generated audio.
diagnosticsCOMBOBasic3 options: Basic, Detailed Report, Off
reroll_from_chunkCOMBOAutoAuto resumes the longest compatible saved prefix. Choosing a chunk reuses earlier chunks and regenerates that chunk and everything after it.
reroll_nonceINT00–4294967295Change only when regenerating an explicit chunk and you want a new variation with otherwise identical settings.
strict_compatibilityBOOLEANtrue—
debugBOOLEANfalse—
show_previewBOOLEANtrue—
run_storageCOMBOOffAtomically save raw AV chunks and resume a compatible saved run.
run_nameSTRINGEnter a stable name for this saved run. Compatible chunks are selected automatically.
reference_sizeCOMBOMatch OutputMatch Output is the practical default; Max Identity preserves more reference detail.
project_idSTRINGOptional. Leave blank to derive a stable ID from this sampler node. Run Name remains the explicit override.
video_reference_sizeCOMBOEfficient - 0.4 MPEfficient limits Video Guide Frames to about 0.4 MP; Balanced uses about 0.6 MP; Match Output uses the output pixel area. Source aspect ratio is preserved and smaller sources are not enlarged.
first_frameoptIMAGEOptional. Leave all image inputs disconnected for T2VA.
last_frameoptIMAGE—
reference_image_1optIMAGE—
reference_image_2optIMAGE—
reference_image_3optIMAGE—
reference_video_1optIMAGEOptional video guide. Connect the IMAGE frame batch from a video loader; source-video audio is not included. Frames are interpreted at 24 fps and applied to every chunk. Non-native frame counts are padded by repeating the final frame to the next H3 17k+5 count, up to the one-chunk limit.
driving_audiooptAUDIOOptional original audio timeline. It is used as native H3 guide conditioning and selected unchanged for final output.
audio_vaeoptVAERequired only when Driving Audio is connected. Uses the same Audio VAE encode path as ComfyUI Core MiniMax H3 Add Guide.
reference_audio_1optAUDIOOptional standalone audio reference for H3 conditioning. It is not the audio track of Video Guide Frames. Unlike Driving Audio, it does not replace the generated final audio.
reference_audio_vaeoptVAERequired only when Reference Audio is connected. It encodes the reference for H3 conditioning; generated audio remains the output.

Outputs (6)

NameTypeDescription
video_latentsLATENT—
audio_latentsLATENT—
assembly_planH3_CONTINUUM_ASSEMBLY_PLAN—
statusSTRING—
driving_audioAUDIO—
refine_contextH3_CONTINUUM_REFINE_CONTEXT—