Nodes/ComfyUI-VibeVoice/Load VibeVoice Model
ComfyUI Node

Load VibeVoice Model

Point the pack at weights you already have

By wildminder·Created about a year ago·Updated 2 days ago· 599
Load VibeVoice Model
    • VibeVoice Model
    ◄model_fileNo model files found in diffusion_models►
    ◄config_nameAuto-detect►
    ◄attention_modesdpa►
    ◄quantize_llm_4bitfalse►
    ◄dtypeauto►

    The rest of ComfyUI-VibeVoice fetches its models from Hugging Face for you, which is lovely until you've already got the bytes, or the 17-gig original isn't going to fit, or the mirror you found is a GGUF and ComfyUI's own loaders won't touch it. Load VibeVoice Model is the way in: point it at a file, wire its output into the TTS or ASR node, and the dropdown on that node stops mattering.

    Why it exists

    Two reasons, both practical. First, quantization - the 7B and ASR checkpoints are 17.4 GB apiece at full precision, and the community has been asking for GGUF and 8-bit VibeVoice since the day each dropped, precisely because that's the difference between running and not on a consumer card. Second, availability: Microsoft pulled VibeVoice's repo and weights off the internet in September 2025, days after everyone fell in love with it, and put a stripped repo back the next day. Weights that vanish do not always come back on schedule, and a wrapper that only knows how to download from one place is fragile. There's a reason this pack's registry points its 7B entry at a community mirror rather than Microsoft's org.

    How it works

    You drop the weight file into ComfyUI/models/diffusion_models/ - yes, the diffusion folder; that's where audio weights go when a pack wants ComfyUI's existing file-listing machinery. The node then does the loading itself rather than handing the file to ComfyUI, because detect_unet_config() has no idea what VibeVoice's state dict keys mean. It builds the model on CPU, binds the architecture config, audio preprocessor and Qwen2.5 tokenizer from sidecar JSONs or packaged defaults, and hands back a bundle that ComfyUI's VRAM manager then moves to the device in a single transfer.

    Sidecars are optional and take priority over the packaged defaults when present: <weights>.config.json for the architecture, <weights>.preprocessor.json for the audio preprocessor, and a tokenizer.json in the same directory. Skip them and the node falls back to config.json/preprocessor_config.json in the folder, then to its built-in defaults.

    Inputs and output

    model_file is the weight picker. It lists everything it can find in diffusion_models plus ComfyUI-GGUF's unet_gguf folder - GGUF included, since .gguf is filtered out of ComfyUI's normal extension list and gets found by a separate scan. If the dropdown reads No model files found in diffusion_models, that's not an error state, it's a pointer: the file isn't where the node is looking (or you haven't restarted ComfyUI since putting it there).

    config_name defaults to Auto-detect, which reads an embedding fingerprint out of the weights and picks the family for you. Leave it there. The manual options are the four families, and if you choose one that contradicts the file, the node corrects you with a warning instead of failing.

    attention_mode - sdpa by default, with eager for maximum compatibility, flash_attention_2 (only offered when flash-attn is actually installed), and sage. Two traps here, both caught for you but worth knowing: sage with dtype pinned to fp32 is rejected at queue time, and sage is dropped entirely for ASR models because its kernel ignores the padding mask and would return a plausible-looking, wrong transcript.

    quantize_llm_4bit quantizes the Qwen2.5 language model to NF4 through bitsandbytes, leaving the diffusion head at full precision. TTS and realtime models only - ASR ignores it.

    dtype - leave on auto. bf16 and fp32 both produce correct speech; there's no precision you need to hunt for.

    The single output is VibeVoice Model (VIBEVOICE_MODEL). Wire it into the external_model input on VibeVoice TTS or VibeVoice ASR; connected, it overrides that node's model_name dropdown. Type guards stop you mis-wiring - an ASR model into the TTS node fails with an error naming the correct node rather than exploding halfway through loading. A realtime bundle into the TTS node is fine and takes the realtime path.

    Install

    Same pack, same install: Manager → search ComfyUI-VibeVoice, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/wildminder/ComfyUI-VibeVoice.git
    cd ComfyUI-VibeVoice
    pip install -r requirements.txt
    

    For GGUF weights you also need the gguf package - ComfyUI's stock diffusion loader can't parse a .gguf, which is exactly why this node exists:

    pip install gguf
    

    4-bit quantization needs bitsandbytes, which the pack's requirements.txt already pulls in. Restart ComfyUI after dropping new weight files into the folder so the dropdown picks them up.

    When it breaks

    A config mismatch that used to be a wall of red. Getting the family wrong now fails fast with the offending tensor names and a suggested config_name, not a raw size mismatch stack trace. If you see that message, believe the suggestion - it read the weights.

    A flood of [WARNING] Pin error. on Windows. Harmless, and it comes from ComfyUI core pinning partially offloaded weights under RAM pressure. Quant-resident loads (GGUF, fp8) never offload and dodge it; otherwise start ComfyUI with --disable-pinned-memory.

    A 17.4 GB file that won't load. The model is built on CPU before it goes anywhere near your GPU, so you need the system RAM to hold it, not just VRAM. And the 7B TTS weights in particular came off a mirror after Microsoft's pull, so check you actually downloaded the complete checkpoint before blaming the node.

    CategoryWMNodes/sound/tts

    Inputs (5)

    NameTypeDefaultDescription
    model_fileCOMBONo model files found in diffusion_modelsWeight file in the diffusion_models folder. Place the VibeVoice .safetensors (or .bin/.gguf) file there.
    config_nameCOMBOAuto-detectArchitecture config to use. Auto-detect reads the weight file's embedding fingerprint and selects the matching family; an explicit selection that contradicts the weights is auto-corrected with a warning. A sidecar config next to the weight file takes priority over the packaged default fallback.
    attention_modeCOMBOsdpaAttention implementation: Eager (safest), SDPA (balanced), Flash Attention 2 (fastest), Sage (quantized)
    quantize_llm_4bitBOOLEANfalseQuantize the Qwen2.5 LLM to 4-bit NF4 via bitsandbytes.
    dtypeCOMBOautoData type for model precision. 'auto' selects optimal type for device.

    Outputs (1)

    NameTypeDescription
    VibeVoice ModelVIBEVOICE_MODEL—