ComfyUI Node
ZONOS2 Voice Clone
Clone a voice from a native ComfyUI AUDIO input without reference text.
ZONOS2 Voice Clone
- zonos2_model
- reference_audio
- audio
◄textHello! This is ZONOS2 running natively inside ComfyUI.►
◄clean_speaker_backgroundfalse►
◄accurate_modetrue►
◄max_new_tokens1024►
◄temperature1.15►
◄top_k106►
◄top_p0.00►
◄min_p0.18►
◄repetition_window50►
◄repetition_penalty1.20►
◄repetition_codebooks8►
◄speaking_ratedefault►
◄loudness_lufsdefault►
◄estimated_snrdefault►
◄maximum_pausedefault►
◄estimated_bandlimit_hzdefault►
◄leading_silencedefault►
◄trailing_silence3: 0.25-0.5►
◄seed0►
CategoryZONOS2 TTS
Inputs (21)
| Name | Type | Default | Description |
|---|---|---|---|
| zonos2_model | ZONOS2_MODEL | Connect the zonos2_model output from ZONOS2 Model Loader. | |
| text | STRING | Hello! This is ZONOS2 running natively inside ComfyUI. | UTF-8 text to synthesize. |
| reference_audio | AUDIO | Reference voice AUDIO noodle. 5-30 seconds of clear, single-speaker speech is recommended. Use a sample that clearly demonstrates the accent you want, with little silence, music, or reverb. Audio longer than 60 seconds is clipped to the first 60 seconds with a CLI warning. | |
| clean_speaker_background | BOOLEAN | false | Official ZONOS2 clean/noisy conditioning flag. Leave disabled for ordinary recordings or any audible room tone, noise, reverb, or ambience. Enable only for genuinely clean studio-like speech. |
| accurate_mode | BOOLEAN | true | Enable the official ZONOS2 accurate-cloning token for stricter speaker-embedding adherence. It can improve identity matching, but it does not guarantee transfer of accent, cadence, emotion, or other prosody. Disable for looser, potentially more expressive conditioning. |
| max_new_tokens | INT | 102432–6000 | Maximum DAC-code frames the model may generate. More frames allow longer speech but increase generation time and KV-cache memory. Generation can stop earlier when ZONOS2 emits end-of-audio. |
| temperature | FLOAT | 1.150–2 | Sampling randomness. Lower values are steadier; higher values add variation but can reduce clarity. 0 uses greedy sampling. The official ZONOS2 default is 1.15. |
| top_k | INT | 1060–1026 | Keep only the K most likely tokens in each audio codebook before sampling. 0 disables Top-K filtering. The official default is 106. |
| top_p | FLOAT | 0.000–1 | Keep the smallest token set whose combined probability reaches this value. 0 disables Top-P filtering. ZONOS2 normally uses Min-P instead. |
| min_p | FLOAT | 0.180–1 | Remove tokens whose probability is below this fraction of the most likely token. 0 disables Min-P. The official default is 0.18. |
| repetition_window | INT | 500–512 | Number of recent generated frames checked for repeated audio tokens. 0 disables repetition tracking. |
| repetition_penalty | FLOAT | 1.201–2 | Reduces the probability of recently generated tokens to discourage loops. 1.0 disables the penalty. The official default is 1.2. |
| repetition_codebooks | INT | 8-1–9 | Apply repetition penalty to this many codebooks starting from codebook 0. -1 applies it to all 9; 0 disables it for every codebook. The official default is 8. |
| speaking_rate | COMBO | default | Optional ZONOS2 speaking-rate conditioning in cleaned UTF-8 bytes per second. Lower ranges generally produce slower speech. Default leaves speaking rate unconditioned. |
| loudness_lufs | COMBO | default | Optional target integrated loudness in LUFS. More-negative ranges are quieter; less-negative ranges are louder. Default leaves loudness unconditioned. |
| estimated_snr | COMBO | default | Optional estimated signal-to-noise ratio in dB. Higher ranges bias toward cleaner audio; lower ranges can reproduce noisier recording characteristics. Default leaves SNR unconditioned. |
| maximum_pause | COMBO | default | Optional maximum internal pause duration in seconds. Lower ranges favor tighter delivery; higher ranges permit longer pauses. Default leaves maximum pause unconditioned. |
| estimated_bandlimit_hz | COMBO | default | Optional estimated recording bandlimit in Hz. Higher ranges bias toward wider-band, brighter audio; lower ranges can sound more bandwidth-limited. Default leaves bandlimit unconditioned. |
| leading_silence | COMBO | default | Optional amount of silence before speech begins. Default leaves leading silence unconditioned. |
| trailing_silence | COMBO | 3: 0.25-0.5 | Requested silence after speech ends. Bucket 3 (0.25-0.5 seconds) is the official ZONOS2 default. Select default to leave it unconditioned. |
| seed | INT | 00–9223372036854776000 | Sampling seed. A positive value makes generation repeatable for identical inputs and settings. 0 uses the current random state. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | Voice-cloned mono speech as native ComfyUI AUDIO at 44.1 kHz. |