Nodes/DIGIT Nodes/DIGIT ElevenLabs Speech to Speech
ComfyUI Node

DIGIT ElevenLabs Speech to Speech

Re-voice existing audio — keep the delivery, swap the voice

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT ElevenLabs Speech to Speech
  • audio
  • audio
◄voice_id►
◄modeleleven_multilingual_sts_v2►
◄stability0.50►
◄similarity_boost0.75►
◄seed0►
◄api_key►
◄speed1.00►
◄style0.00►
◄use_speaker_boostfalse►
◄remove_background_noisefalse►
◄output_formatpcm_44100►

Most TTS starts from text. DIGIT ElevenLabs Speech to Speech starts from audio: you feed it a voice recording and a target voice ID, and it re-speaks the words in the new voice while keeping the original delivery - the pacing, the emphasis, the emotional read. That distinction is the whole node. You're not re-synthesizing the line from text; you're re-voicing an existing performance.

The use cases are the obvious ones: you recorded a scratch VO yourself and want a polished voice on top, you have a decent AI voice that's one step away from the client's preferred voice, or you need to swap a voice across an existing track without rewriting the script. It's the "fix it in post" node of the voice world.

How it works

You connect audio (an AUDIO tensor - from the pack's ElevenLabs TTS/STT nodes, or loaded from a file) and give it voice_id, the target voice. The model choice is eleven_multilingual_sts_v2 (default) or eleven_english_sts_v2 if you're working in English and want the single-language path. Then the delivery knobs:

  • stability (0.5) - lower means the model is more willing to get expressive; higher means more consistent.
  • similarity_boost (0.75) - how closely it matches the target voice. Crank it for a stricter clone.
  • speed (0.7–1.3), style (0–0.2), use_speaker_boost, and remove_background_noise for cleaning up a messy source take.
  • seed for reproducible runs; output_format (pcm_44100 default, or MP3/Opus).

api_key is optional - it auto-detects ELEVENLABS_API_KEY or the pack's DIGIT_ELEVENLABS_API_KEY. Output is a single audio tensor.

Installing it

Standard pack install:

cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt

Or ComfyUI Manager → search comfyui-digit → install → restart. Then:

export ELEVENLABS_API_KEY=your_key_here

What trips people up

The mental model people get wrong: STS is not TTS-with-a-reference. It needs actual audio in, and it's preserving that audio's delivery - so if your source take is flat and robotic, the output will be a flat, robotic read in a different voice. Garbage in, garbage out, but the garbage now has a better voice. Also, similarity_boost interacts with stability: pushing both to extremes can make the output feel a bit processed, so the 0.75/0.5 defaults are a decent starting point. And like everything ElevenLabs, it's per-call billed, so treat the seed as your friend when you're iterating - same seed, same output, no wasted retries.

CategoryDIGIT/ElevenLabs

Inputs (12)

NameTypeDefaultDescription
audioAUDIO—
voice_idSTRINGTarget voice ID. Connect from Voice Selector or paste directly.
modelCOMBOeleven_multilingual_sts_v22 options: eleven_multilingual_sts_v2, eleven_english_sts_v2
stabilityFLOAT0.500–1—
similarity_boostFLOAT0.750–1—
seedINT00–4294967295—
api_keyoptSTRINGElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var.
speedoptFLOAT1.000.7–1.3—
styleoptFLOAT0.000–0.2—
use_speaker_boostoptBOOLEANfalse—
remove_background_noiseoptBOOLEANfalse—
output_formatoptCOMBOpcm_441003 options: pcm_44100, mp3_44100_192, opus_48000_192

Outputs (1)

NameTypeDescription
audioAUDIO—