RAVEN Streaming Sampler
MiniMax H3's 'streaming' sampler doesn't stream in real time — and that's fine
- model
- positive
- latent
- video_vae
- audio_vae
- LATENT
- IMAGE
- AUDIO
Set expectations up front, because the name will mislead you: "streaming" here means incremental delivery, not real-time generation. The author measured a 192-frame (8-second) clip taking about 301 seconds of sampling, with the first preview fragment arriving 51 seconds in. It's not keeping up with playback - nobody claims it is. What you do get is a video-and-audio preview that streams into the node's own widget chunk by chunk while H3 is still working, plus final IMAGE and AUDIO outputs that never go through a whole-clip VAE decode. That last part is the memory trick that lets this run at all.
This node pairs with RAVEN Model Loader from the same pack. Together they run RAVEN's chunk-major rollout over MiniMax H3 - the open-weight 33B omni-modal model whose native stereo audio made it the first credible open answer to Veo (watch the license, though: the H3 Community License excludes the US, EU, UK and South Korea).
How it works
A stock sampler denoises the whole clip at once, sigma by sigma. RAVEN is chunk-major: chunk 0 is the text prefill, then each media chunk starts from fresh noise, is carried to completion in a fixed number of consistency steps, and is written into the KV cache before the next chunk begins - conditioned on everything before it. Because of that loop, this is not a stock sampler: there's no sampler or scheduler selector, and no CFG or negative prompt. One conditioning branch only. A stock bidirectional H3 model is rejected outright - it needs the chunk-causal DiT the RAVEN loader produces.
Inputs that matter
model- theMODELfromRAVEN Model Loader(extra official LoRAs stacked after it are fine).positive- the conditioning fromMiniMaxH3ImageToVideoused in T2VA form (no first/last frame). It has noCLIPor prompt input; text encoding happened upstream.latent- the empty AV latent from that same node. This is where you set width, height and frames - the sampler deliberately has no canvas inputs. A non-empty latent is refused.video_vaeandaudio_vae- the H3 video VAE (24 channels) and audio VAE (32-channel stereo @ 32 kHz).
The two you'll actually tune are seed and steps. Leave steps at 4 - the RAVEN preview adapter was distilled for it, and more is not a free quality win. video_shift (12) and audio_shift (3) run independent sigma grids; the defaults are fine to start. sink (2) and window (2) control how many KV-cache chunks stay pinned as attention sinks vs. kept as recent context. And kv_cache_storage is the memory lever: default cpu_pinned keeps the chunk cache in page-locked host RAM (~0.56 GiB of VRAM at 192 frames instead of ~28 GiB on the card). It changes where bytes live, not what's computed.
Outputs and the frame math
Three outputs: LATENT (the finished AV latent), IMAGE (every frame, in order), and AUDIO (stereo waveform with whole-clip loudness normalization applied once at the end). All are built incrementally per chunk - there's no whole-clip decode, which is exactly what avoids the OOM that killed the naive path.
Frames must satisfy 17k + 5 with k ≥ 1: minimum 22, hard maximum 362. Width and height need to be multiples of 32 and width × height ≤ 1376 × 768. A 22-frame clip technically works but streams almost nothing before the end - the fixed audio tail eats it.
Install and gotchas
Same install as its sibling: ComfyUI Manager (search MiniMax H3 RAVEN Streaming) or
cd ComfyUI/custom_nodes
git clone https://github.com/YanzuoLu/ComfyUI-MiniMax-H3-RAVEN-Streaming.git
then restart. No extra pip installs in a stock ComfyUI; grab the five model files the README lists (DiT, RAVEN LoRA, Qwen3-VL text encoder, two VAEs) and drag in the bundled minimax_h3_raven_streaming_t2va template.
Where people hit walls: a frame count that isn't 17k+5 (fails loud - good); feeding keyframe or reference conditioning (refused with an explicit error, an honest implementation limit, not a quality bug); and RAM - measured host RSS peaked near 130 GB at 192 frames, so 128 GB boxes will swap or die. It's single-GPU only, the preview is best-effort (a failure never touches LATENT/IMAGE/AUDIO), and the preview's audio is pre-normalization, so what you hear won't be exactly the delivered file. Consider it a progress bar with sound.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | A MODEL from RAVEN Model Loader (optionally with official LoRAs stacked after it). A stock bidirectional H3 model is rejected: the loop needs the chunk-causal DiT. | |
| positive | CONDITIONING | The positive CONDITIONING from MiniMax H3 Image to Video used in T2VA form. There is no negative input and no CFG: the chunk-major loop runs one conditioning branch, so a second one would be silently ignored. Keyframe (fl2va) and reference (ref2va) extras are refused with an explicit error: this sampler has not implemented or verified the causal packed layout for condition rows, so refusing beats dropping them silently. That is an implementation limit here, not a statement about the RAVEN LoRA. | |
| latent | LATENT | The empty AV latent from the same node (or Empty MiniMax H3 AV Latent). It defines the frame count and canvas; this node deliberately has no width/height/frames inputs. A non-empty latent is refused - every chunk starts from its own fresh noise. | |
| video_vae | VAE | The MiniMax H3 video VAE (24 latent channels). | |
| audio_vae | VAE | The MiniMax H3 audio VAE (32 channels, stereo, 32 kHz). | |
| seed | INT | 00–18446744073709550000 | Seeds a private generator; the rollout never touches global RNG, so the same seed is the same clip. |
| steps | INT | 41–100 | Consistency NFEs per chunk. RAVEN's published preview trial is 4; more steps is not a free quality win, the schedule was distilled for this budget. |
| video_shift | FLOAT | 12.000.01–100 | Shift of the video stream's trailing sigma grid. |
| audio_shift | FLOAT | 3.000.01–100 | Shift of the audio stream's own trailing sigma grid. The two streams run independent grids, not one remapped grid. |
| sink | INT | 21–64 | Attention-sink cache chunks pinned from the start. Chunk 0 is the text prefill, so 2 means text + the first media chunk. |
| window | INT | 20–64 | Most recent cache chunks kept besides the sinks. 0 keeps only the sinks. |
| kv_cache_storage | COMBO | cpu_pinned | Where the retained chunk KV cache lives. 'cpu_pinned' (default) keeps it in page-locked host memory and copies one layer's retained rows back per block: about 0.56 GiB of VRAM at 192 frames instead of the ~28 GiB the whole cache costs on the card. 'cpu' is the same without page-locking (slower copies, no pinned-memory pressure). 'gpu' keeps it resident and is only for cards with room to spare. This changes where bytes live, not what is computed. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| LATENT | LATENT | The finished AV latent, same nested (video, audio) structure as the input. |
| IMAGE | IMAGE | Every frame, in order, as written by the incremental video collector while the rollout ran. The clip is not decoded again at the end. |
| AUDIO | AUDIO | The stereo waveform, as written by the incremental (overlap-save) audio collector while the rollout ran, with the official whole-clip loudness normalisation applied once at finalize. |