ComfyUI Node

VoxCPM TTS

The one node that does zero-shot speech, voice design, and voice cloning

By wildminder·Created 11 months ago·Updated 3 months ago· 496
VoxCPM TTS
  • voice_config
  • advanced_params
  • Generated Audio
model_nameVoxCPM2
lora_nameNone
textVoxCPM is an innovative TTS model designed to generate highly expressive speech.
voice_design
cfg2.0
steps10
normalize_textfalse
seed-1
force_offloadfalse
devicecpu
dtypeauto

This is the only node in the pack you strictly need. Drop it, type some text, hit run, and you get speech - everything else in ComfyUI-VoxCPM (voice cloning config, advanced params, the training nodes) is optional extra plumbing that plugs into this one node. It's built on VoxCPM, a tokenizer-free TTS model out of OpenBMB/ModelBest: instead of chopping speech into discrete audio tokens like most TTS pipelines, it models speech directly in continuous space on top of a MiniCPM-4 language-model backbone. The practical upshot is expressive, context-aware prosody and genuinely good zero-shot voice cloning from a short sample.

The author, wildminder, isn't new to this - they also wrote the well-known ComfyUI-VibeVoice and ComfyUI-Chatterbox nodes, so this is someone who's been building TTS plumbing for ComfyUI for a while, not a one-off port. Community sentiment backs that up: people who've actually used VoxCPM tend to say it's underrated next to the louder names in the space (Chatterbox, Kokoro, VibeVoice) - more accurate, fewer skipped words on longer passages, and fast enough that the original model claims a real-time factor as low as 0.17 on a consumer RTX 4090.

Two models, four ways to use it

model_name picks the backbone: VoxCPM2 (2B params, 48kHz, 30+ languages, the one with voice design and reference cloning) or VoxCPM1.5 (800M params, 44.1kHz, mostly English/Chinese, the one with LoRA support). Default and the recommended pick is VoxCPM2.

From there this single node covers four modes depending on what you connect: plain zero-shot TTS with just text; voice design (VoxCPM2 only) where voice_design takes a plain-English description like "warm female voice with a slight British accent"; voice cloning via a connected voice_config carrying prompt or reference audio; and "ultimate cloning," which combines a reference clip (for identity) with a prompt clip (for prosody) for the closest match. Note that voice_design only does anything in plain-TTS mode - the moment reference or prompt audio is wired in, the model clones from that instead and ignores the description.

The inputs and outputs that matter

  • text - what gets spoken. Multiline, so paste in a full script.
  • cfg (default 2.0) - guidance scale. Push it up for tighter adherence to the voice description or reference; push it too far and it starts sounding less natural.
  • steps (default 10) - diffusion steps. More can help quality, at the cost of generation time.
  • normalize_text - text normalization for things like numbers and abbreviations. Off by default, and it needs the inflect and wetext Python packages installed to work at all.
  • seed, lora_name (pick a trained LoRA from the dropdown, None by default), force_offload, device, dtype round out the rest.

Two optional inputs are where the modular part of the pack comes in: voice_config (from the VoxCPM Voice Cloning node) and advanced_params (from VoxCPM Advanced Parameters). Leave both unplugged and you get sane zero-shot defaults. The single output is Generated Audio, an AUDIO - wire it straight into a SaveAudio node or anything downstream that expects audio.

Installing it

Easiest path is ComfyUI Manager - search ComfyUI-VoxCPM, install, restart. Manually:

cd ComfyUI/custom_nodes
git clone https://github.com/wildminder/ComfyUI-VoxCPM.git
cd ComfyUI-VoxCPM
pip install -r requirements.txt

The first time you actually run a node, it auto-downloads the selected model into ComfyUI/models/tts/VoxCPM/ - VoxCPM2 is a 2B-parameter model, so budget the disk space and expect a pause on that first generation.

Common issues & troubleshooting

device defaults to cpu in this node, not auto - if generation is crawling and you've got an NVIDIA card sitting idle, that's why. Switch it to cuda explicitly.

normalize_text errors out. It depends on the inflect and wetext packages, which aren't pulled in automatically - pip install inflect wetext into the same Python environment ComfyUI runs on fixes it. This is a genuinely common one; people hit it the moment they flip the toggle on.

First generation after a fresh install fails, second one works. A few users have reported exactly this pattern - if your very first run comes back empty or errors right after setup, just run it again before you start digging for a real bug.

VoxCPM1.5 output turns slightly metallic toward the end of longer sentences. This is a known artifact of the 1.5 model specifically, reported by people who've used it for longer-form narration. Switching to VoxCPM2 or breaking long text into shorter chunks both help.

Categoryaudio/tts

Inputs (13)

NameTypeDefaultDescription
model_nameCOMBOVoxCPM2Select the VoxCPM model to use. VoxCPM2 supports voice design and reference cloning.
lora_nameCOMBONoneSelect a LoRA to apply for style/fine-tuning.
textSTRINGVoxCPM is an innovative TTS model designed to generate highly expressive speech.Text to synthesize.
voice_designSTRINGVoice design instruction for VoxCPM2 (e.g., 'warm female voice', 'deep male voice'). Only used in plain TTS mode (no reference audio). Ignored when reference audio is connected or for VoxCPM1.5.
cfgFLOAT2.00.1–10Guidance scale. Higher values adhere more to the prompt but may sound less natural.
stepsINT101–100Number of diffusion steps. Higher values may improve quality but are slower.
normalize_textBOOLEANfalseEnable text normalization (requires 'inflect' and 'wetext' packages).
seedINT-1-1–18446744073709550000Seed for reproducibility. -1 for random.
force_offloadBOOLEANfalseForce the model to be offloaded from VRAM after generation.
deviceCOMBOcpuDevice to run inference on.
dtypeCOMBOautoData type for model precision. 'auto' selects optimal type for device.
voice_configoptVOICE_CLONING_CONFIGVoice cloning configuration from VoxCPM Voice Cloning node. Provides voice_design, prompt_audio, reference_audio, trim_silence.
advanced_paramsoptADVANCED_PARAMSAdvanced parameters from VoxCPM Advanced Parameters node. Provides min/max_tokens, temperature, sway, use_cfg_zero_star, retry_*.

Outputs (1)

NameTypeDescription
Generated AudioAUDIO