Nodes/ComfyUI-VibeVoice/VibeVoice ASR
ComfyUI Node

VibeVoice ASR

A local transcript with speaker labels already in it

By wildminder·Created about a year ago·Updated 2 days ago· 599
VibeVoice ASR
  • audio
  • external_model
  • Transcription
  • Segments (JSON)
◄model_nameVibeVoice-ASR-HF►
◄context_info►
◄max_new_tokens32768►
◄temperature0.00►
◄top_p1.00►
◄do_samplefalse►
◄num_beams1►
◄devicecpu►
◄dtypeauto►
◄attention_modesdpa►
◄force_offloadfalse►

Most speech-to-text in ComfyUI is Whisper wearing a costume, and the annoying part was never the words - it's the "who said this, and when". Diarization used to mean a second model, a speaker-embedding step and a bunch of glue. VibeVoice ASR hands you one JSON blob with speaker, text, start and end already in it, running entirely on your own machine.

What it's for

The obvious jobs: captioning your own recordings, turning a two-person interview into a transcript you can read, checking what a TTS model actually said before you commit it to a lip-synced clip. The less obvious one is dataset work. If you are building a voice-cloning set, you want labelled segments, not a wall of prose.

This is the ASR half of Microsoft's VibeVoice family - the same lineage as the 1.5B/7B long-form conversational TTS the community went feral over in August 2025. Microsoft released the ASR model in January 2026, and it landed in an audio layer that ComfyUI never really designed for: every audio tool here is a node pack bolted on afterwards, fighting transformers and torch pins with every other pack in your custom_nodes folder. Keep that in mind - it's the thing that will actually bite you, not the model.

How it works

VibeVoice ASR is a ~7B model: a Qwen2.5-style language backbone that reads audio through VibeVoice's continuous acoustic tokenizer and writes timestamped, speaker-tagged segments back out. The node runs the Hugging Face-native single-pass path for VibeVoice-ASR-HF - chat-template prompt in, parsed segments out - and hands you both the raw text and the structured list. The author claims 50+ languages and up to 60 minutes of audio in a single pass, which is the real selling point versus chunking a long file yourself.

Decoding defaults to greedy: temperature 0, do_sample off. That's correct here - deterministic and repeatable, nothing invented for style.

Inputs and outputs

Four fields do the work:

  • audio - the AUDIO input from Load Audio, or anything else in your graph that emits audio (including a TTS node, which is a handy QC loop).
  • model_name - defaults to VibeVoice-ASR-HF. Old workflows pointing at the retired VibeVoice-ASR name get silently redirected to this checkpoint.
  • context_info - optional hotwords. This one is genuinely worth typing: names, product terms, jargon. The author's own example is Tea Brew, Aiden Host. Whisper-style models guess at proper nouns; this gives them a nudge.
  • max_new_tokens - 32768 by default. Long audio that gets truncated mid-sentence means raise this.

The rest is tuning you can leave alone: temperature/top_p/do_sample/num_beams (note num_beams above 1 forces greedy output anyway - beam search and sampling don't mix here), device, dtype, attention_mode, force_offload, and the optional external_model input.

On attention_mode: sage is excluded from the ASR path deliberately. The ASR processor left-pads batches, and the sage kernel discards the additive mask, so every token attends to the pad columns and your transcript comes out silently wrong - no crash, just garbage. Pick it and the node downgrades to sdpa with a warning. Leave it on sdpa.

Outputs are two strings: Transcription (plain text - wire it into a text-display or save node) and Segments (JSON), an array of {speaker, text, start, end} objects, which is what you feed to anything that wants subtitles.

Install

Manager → Install Custom Nodes → search ComfyUI-VibeVoice → Install → restart. Manually:

cd ComfyUI/custom_nodes
git clone https://github.com/wildminder/ComfyUI-VibeVoice.git
cd ComfyUI-VibeVoice
pip install -r requirements.txt

Then the honest part: the ASR checkpoint is 17.4 GB and downloads on first run into ComfyUI/models/tts/VibeVoice/. The pack's 4-bit toggle doesn't exist on this node - ASR always loads at full precision - so there's no cheap way in. If that doesn't fit, run it on CPU and go make coffee, or wait for community quants and load them through Load VibeVoice Model.

When it breaks

A transformers conflict, nine times out of ten. The ASR model classes don't exist in transformers 4.x, so the node needs 5.3.x or newer and fails at load with a message saying so. The catch is that requirements.txt deliberately ships an uncapped transformers, because pinning it would stomp whatever else you have installed. Wrapper authors have said the same thing from the other side: making the newer ASR model coexist with VibeVoice TTS on an older transformers was the hardest part of the build. You share one Python environment with every custom node you own, and this is the known tax.

A model that won't download. Microsoft deleted VibeVoice's repo and weights in September 2025 and restored them days later, so links and mirrors have moved around. The ASR-HF checkpoint is on Microsoft's Hugging Face org today and the node fetches it with no extra setup, but if a repo 404s, check the mirror situation before you start debugging your install.

License, briefly. VibeVoice ships with an anti-impersonation acceptable-use policy attached to an otherwise MIT license. Worth thirty seconds of reading if you're transcribing other people's audio for anything commercial.

CategoryWMNodes/sound/asr

Inputs (13)

NameTypeDefaultDescription
model_nameCOMBOVibeVoice-ASR-HFSelect the VibeVoice ASR model to use.
audioAUDIOAudio to transcribe. Supports up to 60 minutes of audio.
context_infoSTRINGOptional hotwords or context info to improve transcription accuracy (e.g., 'Tea Brew, Aiden Host').
max_new_tokensINT32768256–131072Maximum number of tokens to generate. Increase for longer audio.
temperatureFLOAT0.000–2Temperature for sampling. 0 = greedy (deterministic).
top_pFLOAT1.000–1Top-p for nucleus sampling. Active only if temperature > 0.
do_sampleBOOLEANfalseEnable sampling for more varied output. Disable for deterministic transcription.
num_beamsINT11–10Number of beams for beam search. 1 = no beam search. Higher values may improve quality but are slower.
deviceCOMBOcpuDevice to run inference on.
dtypeCOMBOautoData type for model precision. 'auto' selects optimal type for device.
attention_modeCOMBOsdpaAttention implementation: Eager (safest), SDPA (balanced), Flash Attention 2 (fastest).
force_offloadBOOLEANfalseForce model to be offloaded from VRAM after transcription.
external_modeloptVIBEVOICE_MODELOptional externally-loaded VibeVoice ASR model (from the 'Load VibeVoice Model' node). When connected, this overrides the model_name dropdown.

Outputs (2)

NameTypeDescription
TranscriptionSTRING—
Segments (JSON)STRING—