DIGIT ElevenLabs Voice Clone
Clone a voice from a short sample — no training run required
- audio1
- audio2
- audio3
- audio4
- voice_id
Voice cloning usually sounds like a project: gather hours of clean audio, train a model, babysit it. ElevenLabs' instant voice clone is the shortcut that made the company famous, and this node puts that shortcut inside ComfyUI. Give it an audio clip of someone talking, and it returns a voice_id you can feed straight into the DIGIT Text to Speech node. One queue, no training, done.
It's a great fit for the "keep one consistent voice across a whole video pipeline" problem. Clone the narrator once, reuse the same voice_id for every line, and your characters stop wandering between takes.
How it works
The node takes ComfyUI AUDIO input, converts it to WAV bytes, and sends it to ElevenLabs' instant voice clone endpoint with your API key. The API does the heavy lifting - no local model, no GPU. You get back a voice_id, which is just a string: a reference to the cloned voice stored on ElevenLabs' side. That's the whole trick. The clone isn't a file you download; it's a handle that TTS, speech-to-speech, and the rest of the ElevenLabs suite can use.
The inputs:
- audio1 - required. Your reference clip. ElevenLabs instant cloning wants a few seconds of clean, single-speaker speech. More audio (up to four clips total) tends to clone better.
- audio2, audio3, audio4 - optional extra samples. Use them if you have them; a couple of different takes of the same person gives the clone more to work from.
- remove_background_noise - default off. Flip it on if your sample has hiss, room tone, or music. Better to clean the source first with the Voice Isolation node, but this is a decent safety net.
- voice_name - optional label. Leave it empty and the API auto-generates one.
- api_key - auto-detected from
ELEVENLABS_API_KEYorDIGIT_ELEVENLABS_API_KEY, so usually you don't touch it.
The single output is voice_id (STRING). Wire that into a Text to Speech node's voice_id input and you're speaking as your clone.
Install
It's one of 53 nodes in the comfyui-digit pack. ComfyUI Manager → search comfyui-digit → install → restart, or:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Restart ComfyUI, then look under DIGIT/ElevenLabs.
Common issues
Set your key first:
export ELEVENLABS_API_KEY=sk_...
Without it you'll get the pack's standard ElevenLabs API key is required error. Beyond that, clone quality is mostly about the source audio. A 3-second clip with someone's dog barking in the background clones badly no matter what the node does. Keep the reference clean and single-voiced, and if you can, let the Voice Isolation node strip the noise before you clone.
The other thing to know: the cloned voice lives on ElevenLabs' servers, tied to your account. Your reference audio leaves the machine to get there - that's inherent to the product, not a bug in this node. Fine for most voice work; a real consideration if you're cloning for a client who cares about where their voice files end up.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| audio1 | AUDIO | — | |
| remove_background_noise | BOOLEAN | false | — |
| api_keyopt | STRING | ElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var. | |
| voice_nameopt | STRING | Name for the cloned voice. Auto-generated if empty. | |
| audio2opt | AUDIO | — | |
| audio3opt | AUDIO | — | |
| audio4opt | AUDIO | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| voice_id | STRING | — |