Nodes/zonos2-tts-comfyui-emotion/ZONOS2 Voice Clone (Emotion)
ComfyUI Node

ZONOS2 Voice Clone (Emotion)

Clone a voice from a native ComfyUI AUDIO input without reference text.

By shiwano·Created 17 days ago·Updated 6 days ago· 0
ZONOS2 Voice Clone (Emotion)
  • zonos2_model
  • reference_audio
  • audio
textHello! This is ZONOS2 running natively inside ComfyUI.
clean_speaker_backgroundfalse
accurate_modetrue
max_new_tokens1024
temperature1.15
top_k106
top_p0.00
min_p0.18
repetition_window50
repetition_penalty1.20
repetition_codebooks8
speaking_ratedefault
loudness_lufsdefault
estimated_snrdefault
maximum_pausedefault
estimated_bandlimit_hzdefault
leading_silencedefault
trailing_silence3: 0.25-0.5
seed0
emotionnone
emotion_strength1.00
emotion_valence0.00
emotion_arousal0.00
emotion_cfg_scale1.00
CategoryZONOS2 TTS (Emotion)

Inputs (26)

NameTypeDefaultDescription
zonos2_modelZONOS2_MODELConnect the zonos2_model output from ZONOS2 Model Loader.
textSTRINGHello! This is ZONOS2 running natively inside ComfyUI.UTF-8 text to synthesize.
reference_audioAUDIOReference voice AUDIO noodle. 5-30 seconds of clear, single-speaker speech is recommended. Use a sample that clearly demonstrates the accent you want, with little silence, music, or reverb. Audio longer than 60 seconds is clipped to the first 60 seconds with a CLI warning.
clean_speaker_backgroundBOOLEANfalseOfficial ZONOS2 clean/noisy conditioning flag. Leave disabled for ordinary recordings or any audible room tone, noise, reverb, or ambience. Enable only for genuinely clean studio-like speech.
accurate_modeBOOLEANtrueEnable the official ZONOS2 accurate-cloning token for stricter speaker-embedding adherence. It can improve identity matching, but it does not guarantee transfer of accent, cadence, emotion, or other prosody. Disable for looser, potentially more expressive conditioning.
max_new_tokensINT102432–6000Maximum DAC-code frames the model may generate. More frames allow longer speech but increase generation time and KV-cache memory. Generation can stop earlier when ZONOS2 emits end-of-audio.
temperatureFLOAT1.150–2Sampling randomness. Lower values are steadier; higher values add variation but can reduce clarity. 0 uses greedy sampling. The official ZONOS2 default is 1.15.
top_kINT1060–1026Keep only the K most likely tokens in each audio codebook before sampling. 0 disables Top-K filtering. The official default is 106.
top_pFLOAT0.000–1Keep the smallest token set whose combined probability reaches this value. 0 disables Top-P filtering. ZONOS2 normally uses Min-P instead.
min_pFLOAT0.180–1Remove tokens whose probability is below this fraction of the most likely token. 0 disables Min-P. The official default is 0.18.
repetition_windowINT500–512Number of recent generated frames checked for repeated audio tokens. 0 disables repetition tracking.
repetition_penaltyFLOAT1.201–2Reduces the probability of recently generated tokens to discourage loops. 1.0 disables the penalty. The official default is 1.2.
repetition_codebooksINT8-1–9Apply repetition penalty to this many codebooks starting from codebook 0. -1 applies it to all 9; 0 disables it for every codebook. The official default is 8.
speaking_rateCOMBOdefaultOptional ZONOS2 speaking-rate conditioning in cleaned UTF-8 bytes per second. Lower ranges generally produce slower speech. Default leaves speaking rate unconditioned.
loudness_lufsCOMBOdefaultOptional target integrated loudness in LUFS. More-negative ranges are quieter; less-negative ranges are louder. Default leaves loudness unconditioned.
estimated_snrCOMBOdefaultOptional estimated signal-to-noise ratio in dB. Higher ranges bias toward cleaner audio; lower ranges can reproduce noisier recording characteristics. Default leaves SNR unconditioned.
maximum_pauseCOMBOdefaultOptional maximum internal pause duration in seconds. Lower ranges favor tighter delivery; higher ranges permit longer pauses. Default leaves maximum pause unconditioned.
estimated_bandlimit_hzCOMBOdefaultOptional estimated recording bandlimit in Hz. Higher ranges bias toward wider-band, brighter audio; lower ranges can sound more bandwidth-limited. Default leaves bandlimit unconditioned.
leading_silenceCOMBOdefaultOptional amount of silence before speech begins. Default leaves leading silence unconditioned.
trailing_silenceCOMBO3: 0.25-0.5Requested silence after speech ends. Bucket 3 (0.25-0.5 seconds) is the official ZONOS2 default. Select default to leave it unconditioned.
seedINT00–9223372036854776000Sampling seed. A positive value makes generation repeatable for identical inputs and settings. 0 uses the current random state.
emotionCOMBOnoneOfficial ZONOS2 emotion direction added to the speaker vector after the model projects it. Speaker identity is preserved and only delivery shifts. Select none to leave conditioning untouched, which reproduces the output of a build without emotion support.
emotion_strengthFLOAT1.000–3Multiplier on top of the strengths ZONOS2 calibrated for each direction (3.0 for most, 4.0 for surprised), so 1.0 already gives the intended amount. Raise it for a stronger reading; large values distort the voice. 0 disables emotion.
emotion_valenceFLOAT0.00-1–1Continuous affect axis mixed in alongside the selected emotion. Negative is unpleasant, positive is pleasant, 0 leaves the axis unused.
emotion_arousalFLOAT0.00-1–1Continuous affect axis mixed in alongside the selected emotion. Negative is calm, positive is excited, 0 leaves the axis unused.
emotion_cfg_scaleFLOAT1.001–3Official ZONOS2 emotion classifier-free guidance. 1.0 disables it. Above 1.0 the model also generates an unguided twin and amplifies the difference, which strengthens the emotion at roughly double the generation time. Upstream recommends 1.5 with accurate_mode disabled.

Outputs (1)

NameTypeDescription
audioAUDIOVoice-cloned mono speech as native ComfyUI AUDIO at 44.1 kHz.