DIGIT ElevenLabs Speech to Speech
Re-voice existing audio — keep the delivery, swap the voice
- audio
- audio
Most TTS starts from text. DIGIT ElevenLabs Speech to Speech starts from audio: you feed it a voice recording and a target voice ID, and it re-speaks the words in the new voice while keeping the original delivery - the pacing, the emphasis, the emotional read. That distinction is the whole node. You're not re-synthesizing the line from text; you're re-voicing an existing performance.
The use cases are the obvious ones: you recorded a scratch VO yourself and want a polished voice on top, you have a decent AI voice that's one step away from the client's preferred voice, or you need to swap a voice across an existing track without rewriting the script. It's the "fix it in post" node of the voice world.
How it works
You connect audio (an AUDIO tensor - from the pack's ElevenLabs TTS/STT nodes, or loaded from a file) and give it voice_id, the target voice. The model choice is eleven_multilingual_sts_v2 (default) or eleven_english_sts_v2 if you're working in English and want the single-language path. Then the delivery knobs:
stability(0.5) - lower means the model is more willing to get expressive; higher means more consistent.similarity_boost(0.75) - how closely it matches the target voice. Crank it for a stricter clone.speed(0.7–1.3),style(0–0.2),use_speaker_boost, andremove_background_noisefor cleaning up a messy source take.seedfor reproducible runs;output_format(pcm_44100default, or MP3/Opus).
api_key is optional - it auto-detects ELEVENLABS_API_KEY or the pack's DIGIT_ELEVENLABS_API_KEY. Output is a single audio tensor.
Installing it
Standard pack install:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Or ComfyUI Manager → search comfyui-digit → install → restart. Then:
export ELEVENLABS_API_KEY=your_key_here
What trips people up
The mental model people get wrong: STS is not TTS-with-a-reference. It needs actual audio in, and it's preserving that audio's delivery - so if your source take is flat and robotic, the output will be a flat, robotic read in a different voice. Garbage in, garbage out, but the garbage now has a better voice. Also, similarity_boost interacts with stability: pushing both to extremes can make the output feel a bit processed, so the 0.75/0.5 defaults are a decent starting point. And like everything ElevenLabs, it's per-call billed, so treat the seed as your friend when you're iterating - same seed, same output, no wasted retries.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| audio | AUDIO | — | |
| voice_id | STRING | Target voice ID. Connect from Voice Selector or paste directly. | |
| model | COMBO | eleven_multilingual_sts_v2 | 2 options: eleven_multilingual_sts_v2, eleven_english_sts_v2 |
| stability | FLOAT | 0.500–1 | — |
| similarity_boost | FLOAT | 0.750–1 | — |
| seed | INT | 00–4294967295 | — |
| api_keyopt | STRING | ElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var. | |
| speedopt | FLOAT | 1.000.7–1.3 | — |
| styleopt | FLOAT | 0.000–0.2 | — |
| use_speaker_boostopt | BOOLEAN | false | — |
| remove_background_noiseopt | BOOLEAN | false | — |
| output_formatopt | COMBO | pcm_44100 | 3 options: pcm_44100, mp3_44100_192, opus_48000_192 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |