RAVEN Model Loader
The RAVEN loader is where MiniMax H3 streaming begins — and it has no off switch
- MODEL
RAVEN is what makes MiniMax H3 feel like a live generator instead of a render job. MiniMax's official ComfyUI support samples the whole clip in one shot: build an empty latent, encode the prompt, run a stock sampler, decode the full thing at the end. RAVEN is a third-party LoRA (from the mvp-lab group on HuggingFace) that changes the sampling regime itself - the clip rolls out chunk by chunk, each chunk finished in a handful of consistency steps and committed into a growing KV cache before the next one starts. That's the "streaming" part, and the pack's sampler runs it. RAVEN Model Loader is the only thing that builds the model for that loop.
The name undersells the job. ComfyUI already has loaders that will happily load the H3 DiT. What this node adds is the mandatory RAVEN LoRA, applied the one way RAVEN tolerates - and it's a one-way door. There is no strength knob, no "off" setting, no checkbox to load the model plain. The RAVEN residual is fixed at 1.0 because the adapter is what the 4-NFE consistency schedule was distilled with; "off" isn't a state this node can be put in. If you want a plain H3 model, use the official loader instead.
How it works
The loader walks Comfy's normal diffusion-model loading chain and injects a chunk-causal DiT class in place of the stock one, then attaches the RAVEN LoRA as an FP32 activation residual - computed on activations and added to the base output, never fused as B @ A into the BF16 weights. That detail is what lets the whole thing survive partial CPU offload, where the base weights and the adapter can sit on different devices. The result is a stock static ModelPatcher, so ordinary Comfy offload accounting applies, and you can chain official LoraLoaderModelOnly nodes after it for extra LoRAs.
The three inputs
unet_name- pickminimax_h3_fl2va_bf16.safetensors, the full, non-pruned BF16 DiT. The pruned / adaln-curve checkpoint is rejected with an error: it has notime_embedderfor the LoRA's 266-module mapping to attach to. This is a hard check, not a subtle silent corruption.lora_name- the mandatory RAVEN LoRA,minimax_h3_raven_streaming_lora_4nfe_preview.safetensors. About 5 GB of FP32 A/B tensors (532 tensors, rank 128). Without it, no model. Nothing else satisfies this input.weight_dtype-default(lets Comfy pick BF16 for H3) orbf16. Don't touchfp32unless you're sure the box can hold it: it doubles the base weights to 132 GB+ on top of the FP32 residual, which means heavy offload churn or an OOM on anything realistic. There's no FP8/INT8 option on purpose - quantized linear fusion would silently skip the residual hook and give you a subtly broken model.
The single output is a standard MODEL, which feeds the pack's RAVEN Streaming Sampler - or any other official H3 workflow, if you want to experiment. Nothing here locks the model to one generation mode.
Install and the downloads
Install the pack via ComfyUI Manager (search MiniMax H3 RAVEN Streaming or comfyui-minimax-h3-raven-streaming), or clone it:
cd ComfyUI/custom_nodes
git clone https://github.com/YanzuoLu/ComfyUI-MiniMax-H3-RAVEN-Streaming.git
Then restart ComfyUI. The pack adds only one dependency - PyAV (av>=16.0.0), which ComfyUI already pins - and downloads no models. You grab those yourself into the standard folders:
models/diffusion_models/minimax_h3_fl2va_bf16.safetensors- Comfy-Org/MiniMax-H3models/loras/minimax_h3_raven_streaming_lora_4nfe_preview.safetensors- mvp-lab/MiniMax-H3-RAVEN-Streaming-LoRA- Plus the H3 text encoder and the video + audio VAEs (see the README) for the sampler's other inputs.
A bundled example workflow, minimax_h3_raven_streaming_t2va, appears under Workflow → Browse Templates → Extensions once the pack is installed - drag it in and re-pick the filenames if yours differ.
Where people get burned
Two things trip people up. First, the checkpoint: grab the pruned variant of the H3 DiT and the loader refuses it outright - that's by design, not a bug. Second, memory: this is a 33B model with a 5 GB FP32 adapter bolted on. The author validated against a simulated 24 GiB VRAM envelope, but host RAM is the real wall - measured peak RSS was ~130 GB at 192 frames, so plan for 160 GB available, not the older "128 GB is fine" advice. And remember the weights carry the MiniMax H3 Community License, which geofences the US, EU, UK and South Korea out of using them at all - check your region before downloading.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| unet_name | COMBO | Official full, non-pruned BF16 MiniMax-H3 DiT. Pruned/adaln-curve checkpoints are rejected. The RAVEN adapter is trained against the full BF16 weights; the pruned/adaln-curve checkpoint has no time_embedder for its 266-module mapping to attach to. | |
| lora_name | COMBO | Mandatory RAVEN PEFT LoRA (532 tensors / 266 modules, FP32 A/B). Without it this node cannot produce a model. About 5 GB of FP32 A/B tensors, applied as an activation residual (never fused into the BF16 weights) and counted by Comfy's memory accounting from the first model_size() call. | |
| weight_dtype | COMBO | default | 'default' lets comfy.model_management pick (BF16 for H3). 'fp32' doubles the weights to 132 GB+ for the full non-pruned model, on top of the FP32 RAVEN residual - expect heavy offload traffic or OOM unless you know the box can hold it. There is no FP8/INT8 choice on purpose: comfy.ops fuses quantised linears and would silently skip the residual. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | A standard ComfyUI MODEL (stock static ModelPatcher) whose diffusion_model is the chunk-causal RAVEN DiT. Stock LoraLoaderModelOnly can be chained after it, and nothing here restricts it to one generation mode - other official H3 workflows can be explored with it, unverified and unsupported. The RAVEN Streaming Sampler in this pack is the narrower part: it refuses conditioning carrying condition/reference rows it has not implemented. |