ComfyUI Node
ZONOS2 Voice Clone (Emotion)
Clone a voice from a native ComfyUI AUDIO input without reference text.
ZONOS2 Voice Clone (Emotion)
- zonos2_model
- reference_audio
- audio
◄textHello! This is ZONOS2 running natively inside ComfyUI.►
◄clean_speaker_backgroundfalse►
◄accurate_modetrue►
◄max_new_tokens1024►
◄temperature1.15►
◄top_k106►
◄top_p0.00►
◄min_p0.18►
◄repetition_window50►
◄repetition_penalty1.20►
◄repetition_codebooks8►
◄speaking_ratedefault►
◄loudness_lufsdefault►
◄estimated_snrdefault►
◄maximum_pausedefault►
◄estimated_bandlimit_hzdefault►
◄leading_silencedefault►
◄trailing_silence3: 0.25-0.5►
◄seed0►
◄emotionnone►
◄emotion_strength1.00►
◄emotion_valence0.00►
◄emotion_arousal0.00►
◄emotion_cfg_scale1.00►
CategoryZONOS2 TTS (Emotion)
Inputs (26)
| Name | Type | Default | Description |
|---|---|---|---|
| zonos2_model | ZONOS2_MODEL | Connect the zonos2_model output from ZONOS2 Model Loader. | |
| text | STRING | Hello! This is ZONOS2 running natively inside ComfyUI. | UTF-8 text to synthesize. |
| reference_audio | AUDIO | Reference voice AUDIO noodle. 5-30 seconds of clear, single-speaker speech is recommended. Use a sample that clearly demonstrates the accent you want, with little silence, music, or reverb. Audio longer than 60 seconds is clipped to the first 60 seconds with a CLI warning. | |
| clean_speaker_background | BOOLEAN | false | Official ZONOS2 clean/noisy conditioning flag. Leave disabled for ordinary recordings or any audible room tone, noise, reverb, or ambience. Enable only for genuinely clean studio-like speech. |
| accurate_mode | BOOLEAN | true | Enable the official ZONOS2 accurate-cloning token for stricter speaker-embedding adherence. It can improve identity matching, but it does not guarantee transfer of accent, cadence, emotion, or other prosody. Disable for looser, potentially more expressive conditioning. |
| max_new_tokens | INT | 102432–6000 | Maximum DAC-code frames the model may generate. More frames allow longer speech but increase generation time and KV-cache memory. Generation can stop earlier when ZONOS2 emits end-of-audio. |
| temperature | FLOAT | 1.150–2 | Sampling randomness. Lower values are steadier; higher values add variation but can reduce clarity. 0 uses greedy sampling. The official ZONOS2 default is 1.15. |
| top_k | INT | 1060–1026 | Keep only the K most likely tokens in each audio codebook before sampling. 0 disables Top-K filtering. The official default is 106. |
| top_p | FLOAT | 0.000–1 | Keep the smallest token set whose combined probability reaches this value. 0 disables Top-P filtering. ZONOS2 normally uses Min-P instead. |
| min_p | FLOAT | 0.180–1 | Remove tokens whose probability is below this fraction of the most likely token. 0 disables Min-P. The official default is 0.18. |
| repetition_window | INT | 500–512 | Number of recent generated frames checked for repeated audio tokens. 0 disables repetition tracking. |
| repetition_penalty | FLOAT | 1.201–2 | Reduces the probability of recently generated tokens to discourage loops. 1.0 disables the penalty. The official default is 1.2. |
| repetition_codebooks | INT | 8-1–9 | Apply repetition penalty to this many codebooks starting from codebook 0. -1 applies it to all 9; 0 disables it for every codebook. The official default is 8. |
| speaking_rate | COMBO | default | Optional ZONOS2 speaking-rate conditioning in cleaned UTF-8 bytes per second. Lower ranges generally produce slower speech. Default leaves speaking rate unconditioned. |
| loudness_lufs | COMBO | default | Optional target integrated loudness in LUFS. More-negative ranges are quieter; less-negative ranges are louder. Default leaves loudness unconditioned. |
| estimated_snr | COMBO | default | Optional estimated signal-to-noise ratio in dB. Higher ranges bias toward cleaner audio; lower ranges can reproduce noisier recording characteristics. Default leaves SNR unconditioned. |
| maximum_pause | COMBO | default | Optional maximum internal pause duration in seconds. Lower ranges favor tighter delivery; higher ranges permit longer pauses. Default leaves maximum pause unconditioned. |
| estimated_bandlimit_hz | COMBO | default | Optional estimated recording bandlimit in Hz. Higher ranges bias toward wider-band, brighter audio; lower ranges can sound more bandwidth-limited. Default leaves bandlimit unconditioned. |
| leading_silence | COMBO | default | Optional amount of silence before speech begins. Default leaves leading silence unconditioned. |
| trailing_silence | COMBO | 3: 0.25-0.5 | Requested silence after speech ends. Bucket 3 (0.25-0.5 seconds) is the official ZONOS2 default. Select default to leave it unconditioned. |
| seed | INT | 00–9223372036854776000 | Sampling seed. A positive value makes generation repeatable for identical inputs and settings. 0 uses the current random state. |
| emotion | COMBO | none | Official ZONOS2 emotion direction added to the speaker vector after the model projects it. Speaker identity is preserved and only delivery shifts. Select none to leave conditioning untouched, which reproduces the output of a build without emotion support. |
| emotion_strength | FLOAT | 1.000–3 | Multiplier on top of the strengths ZONOS2 calibrated for each direction (3.0 for most, 4.0 for surprised), so 1.0 already gives the intended amount. Raise it for a stronger reading; large values distort the voice. 0 disables emotion. |
| emotion_valence | FLOAT | 0.00-1–1 | Continuous affect axis mixed in alongside the selected emotion. Negative is unpleasant, positive is pleasant, 0 leaves the axis unused. |
| emotion_arousal | FLOAT | 0.00-1–1 | Continuous affect axis mixed in alongside the selected emotion. Negative is calm, positive is excited, 0 leaves the axis unused. |
| emotion_cfg_scale | FLOAT | 1.001–3 | Official ZONOS2 emotion classifier-free guidance. 1.0 disables it. Above 1.0 the model also generates an unguided twin and amplifies the difference, which strengthens the emotion at roughly double the generation time. Upstream recommends 1.5 with accurate_mode disabled. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | Voice-cloned mono speech as native ComfyUI AUDIO at 44.1 kHz. |