Nodes/Ninode Utils/VibeVoice Voice Design
ComfyUI Node

VibeVoice Voice Design

A ComfyUI node in Ninode Utils/Voice Design with 13 inputs and 2 outputs.

By iGavroche·Created 10 months ago·Updated 10 months ago· 1
VibeVoice Voice Design
  • reference_audio
  • voice_id
  • trial_audio
promptA narrator telling a suspenseful story, with a deep and magnetic voice, varying speech pace to create a tense and mysterious atmosphere.
preview_textIt was late at night, and he was alone in the old house. Faint footsteps could be heard outside the window. He held his breath and slowly, slowly, walked toward the creaking door...
model_name
attention_modesdpa
cfg_scale1.30
inference_steps10
seed42
custom_voice_id
quantize_llm_4bitfalse
temperature0.95
top_p0.95
force_offloadfalse
CategoryNinode Utils/Voice Design

Inputs (13)

NameTypeDefaultDescription
promptSTRINGA narrator telling a suspenseful story, with a deep and magnetic voice, varying speech pace to create a tense and mysterious atmosphere.
preview_textSTRINGIt was late at night, and he was alone in the old house. Faint footsteps could be heard outside the window. He held his breath and slowly, slowly, walked toward the creaking door...
model_nameCOMBOSelect the VibeVoice model to use. Official models will be downloaded automatically.
attention_modeCOMBOsdpaAttention implementation: Eager (safest), SDPA (balanced), Flash Attention 2 (fastest)
cfg_scaleFLOAT1.300.1–50Classifier-Free Guidance scale. Higher values increase adherence to the voice prompt.
inference_stepsINT101–500Number of diffusion steps for audio generation.
seedINT420–18446744073709550000Seed for reproducibility. Set to 0 for a random seed on each run.
custom_voice_idoptSTRING
reference_audiooptAUDIOReference audio for voice cloning (optional). If provided, will be used as speaker voice.
quantize_llm_4bitoptBOOLEANfalseQuantize the Qwen2.5 LLM to 4-bit NF4 via bitsandbytes.
temperatureoptFLOAT0.950–2Controls randomness in generation.
top_poptFLOAT0.950–1Nucleus sampling (Top-P).
force_offloadoptBOOLEANfalseForce model to be offloaded from VRAM after generation.

Outputs (2)

NameTypeDescription
voice_idSTRING
trial_audioSTRING