Nodes/ZONOS2 TTS/ZONOS2 Voice Clone
ComfyUI Node

ZONOS2 Voice Clone

Clone a voice from a clip — no transcript, no training, no API

By Saganaki22·Created 3 months ago·Updated 3 months ago· 12
ZONOS2 Voice Clone
  • zonos2_model
  • reference_audio
  • audio
textHello! This is ZONOS2 running natively inside ComfyUI.
clean_speaker_backgroundfalse
accurate_modetrue
max_new_tokens1024
temperature1.15
top_k106
top_p0.00
min_p0.18
repetition_window50
repetition_penalty1.20
repetition_codebooks8
speaking_ratedefault
loudness_lufsdefault
estimated_snrdefault
maximum_pausedefault
estimated_bandlimit_hzdefault
leading_silencedefault
trailing_silence3: 0.25-0.5
seed0

This is the node that makes the pack worth installing. You give ZONOS2 Voice Clone a few seconds of someone talking and some text, and it synthesizes that text in their voice - audio-only zero-shot cloning, no reference transcript, no training run, no API call. Everything stays on your GPU.

It works the way IP-Adapter works for faces: the model doesn't look at the reference waveform at all. Instead a speaker encoder (an ECAPA-TDNN model) squeezes your reference clip into a single 2048-dimensional speaker embedding, and ZONOS2 conditions its generation on that fixed vector. Identity transfers, and because the embedding is a fixed point rather than a growing context, cloning is cheap and fast. The flip side, and the README is refreshingly honest about this: accent, cadence, emotion, and other utterance-level prosody don't reliably survive the trip. You get the voice, not necessarily the performance.

Wiring it up

The node takes three required inputs on top of the model:

  • zonos2_model - from the ZONOS2 Model Loader.
  • text - what the cloned voice says.
  • reference_audio - a native ComfyUI AUDIO noodle. Feed it from a Load Audio node (the pack's example workflow uses LoadAudio). Best results need 5–30 seconds of clear, single-speaker speech; anything past 60 seconds gets clipped with a warning in the CLI, and very short clips get a recommendation warning too.

Two toggles genuinely change behavior:

  • clean_speaker_background (default false) - the official clean/noisy recording flag. Leave it off for ordinary recordings; it's not a denoiser, it describes the source. Enable it only for genuinely studio-clean speech.
  • accurate_mode (default true) - requests stricter adherence to the embedding, improving identity match at the cost of expressiveness. Turn it off for looser, more expressive delivery. It is not an accent-strength slider.

Below those sit the same sampling controls as the generation node (max_new_tokens, temperature, top_k, top_p, min_p, repetition settings, seed) and the same quality-conditioning buckets. temperature 0.8–1.0 plus a fixed seed is the "make it stop wiggling" combo.

The output is audio - mono, 44.1 kHz, native AUDIO - which wires into Save Audio/SaveAudioMP3 or downstream audio nodes. The speaker encoder is loaded lazily, so the first clone run will take a moment longer than plain generation.

Installing

Same pack as the other two nodes. ComfyUI Manager → ZONOS2 TTS, or:

cd ComfyUI/custom_nodes
git clone https://github.com/Saganaki22/Zonos2_TTS-ComfyUI.git
../venv/bin/python Zonos2_TTS-ComfyUI/install.py

Restart. Needs Transformers 5.0–5.12 (4.x won't work), and the loader handles the model downloads. Also: read the model's acceptable-use policy once before you build anything. Cloning a voice you don't have permission to use is the one line this project draws hard.

Troubleshooting

  • Clone sounds weak or wrong → it's usually the reference. Get 5–30 seconds of clean single-speaker speech with minimal silence, set clean_speaker_background to match the actual recording (not on by reflex), and keep accurate_mode on.
  • Voice matches but the accent doesn't → model limitation, not a bug. The embedding is built for identity and discards prosody. Use a reference that clearly demonstrates the accent (ideally in the target language), compare a few fixed seeds, and know that accurate_mode and temperature can't conjure accent info the embedding never carried.
  • CLI warning about a long reference → expected; it clips to the first 60 seconds.
  • CUDA OOM → the FP8 preset from the loader, unload other models, lower max_new_tokens.

For getting one specific person's voice out of a short clip, this is currently the most direct path ComfyUI has - just manage your expectations on how they sound, not whose voice it is.

CategoryZONOS2 TTS

Inputs (21)

NameTypeDefaultDescription
zonos2_modelZONOS2_MODELConnect the zonos2_model output from ZONOS2 Model Loader.
textSTRINGHello! This is ZONOS2 running natively inside ComfyUI.UTF-8 text to synthesize.
reference_audioAUDIOReference voice AUDIO noodle. 5-30 seconds of clear, single-speaker speech is recommended. Use a sample that clearly demonstrates the accent you want, with little silence, music, or reverb. Audio longer than 60 seconds is clipped to the first 60 seconds with a CLI warning.
clean_speaker_backgroundBOOLEANfalseOfficial ZONOS2 clean/noisy conditioning flag. Leave disabled for ordinary recordings or any audible room tone, noise, reverb, or ambience. Enable only for genuinely clean studio-like speech.
accurate_modeBOOLEANtrueEnable the official ZONOS2 accurate-cloning token for stricter speaker-embedding adherence. It can improve identity matching, but it does not guarantee transfer of accent, cadence, emotion, or other prosody. Disable for looser, potentially more expressive conditioning.
max_new_tokensINT102432–6000Maximum DAC-code frames the model may generate. More frames allow longer speech but increase generation time and KV-cache memory. Generation can stop earlier when ZONOS2 emits end-of-audio.
temperatureFLOAT1.150–2Sampling randomness. Lower values are steadier; higher values add variation but can reduce clarity. 0 uses greedy sampling. The official ZONOS2 default is 1.15.
top_kINT1060–1026Keep only the K most likely tokens in each audio codebook before sampling. 0 disables Top-K filtering. The official default is 106.
top_pFLOAT0.000–1Keep the smallest token set whose combined probability reaches this value. 0 disables Top-P filtering. ZONOS2 normally uses Min-P instead.
min_pFLOAT0.180–1Remove tokens whose probability is below this fraction of the most likely token. 0 disables Min-P. The official default is 0.18.
repetition_windowINT500–512Number of recent generated frames checked for repeated audio tokens. 0 disables repetition tracking.
repetition_penaltyFLOAT1.201–2Reduces the probability of recently generated tokens to discourage loops. 1.0 disables the penalty. The official default is 1.2.
repetition_codebooksINT8-1–9Apply repetition penalty to this many codebooks starting from codebook 0. -1 applies it to all 9; 0 disables it for every codebook. The official default is 8.
speaking_rateCOMBOdefaultOptional ZONOS2 speaking-rate conditioning in cleaned UTF-8 bytes per second. Lower ranges generally produce slower speech. Default leaves speaking rate unconditioned.
loudness_lufsCOMBOdefaultOptional target integrated loudness in LUFS. More-negative ranges are quieter; less-negative ranges are louder. Default leaves loudness unconditioned.
estimated_snrCOMBOdefaultOptional estimated signal-to-noise ratio in dB. Higher ranges bias toward cleaner audio; lower ranges can reproduce noisier recording characteristics. Default leaves SNR unconditioned.
maximum_pauseCOMBOdefaultOptional maximum internal pause duration in seconds. Lower ranges favor tighter delivery; higher ranges permit longer pauses. Default leaves maximum pause unconditioned.
estimated_bandlimit_hzCOMBOdefaultOptional estimated recording bandlimit in Hz. Higher ranges bias toward wider-band, brighter audio; lower ranges can sound more bandwidth-limited. Default leaves bandlimit unconditioned.
leading_silenceCOMBOdefaultOptional amount of silence before speech begins. Default leaves leading silence unconditioned.
trailing_silenceCOMBO3: 0.25-0.5Requested silence after speech ends. Bucket 3 (0.25-0.5 seconds) is the official ZONOS2 default. Select default to leave it unconditioned.
seedINT00–9223372036854776000Sampling seed. A positive value makes generation repeatable for identical inputs and settings. 0 uses the current random state.

Outputs (1)

NameTypeDescription
audioAUDIOVoice-cloned mono speech as native ComfyUI AUDIO at 44.1 kHz.