DIGIT ElevenLabs Voice Isolation
Stripping the noise out of a voice track, without the studio
- audio
- audio
Clean dialogue is the thing that separates a demo that feels like a demo from one that feels finished. This node takes an AUDIO clip - field recording, noisy VO take, video audio with music underneath - and hands it to ElevenLabs to pull the voice out and drop the rest. In, isolated voice out, as a ComfyUI AUDIO tensor. It's one of those nodes you don't need until you need it, and then you need it constantly.
The obvious pipeline is to run this before a Voice Clone or a transcription. Cloning from a noisy sample is a waste of a clone; feeding a speech-to-text model garbage audio guarantees garbage text. Clean first, clone or transcribe second.
How it works
Feed it an AUDIO tensor, it converts it to WAV, sends it to ElevenLabs' voice isolation endpoint, and converts the returned audio back into an AUDIO tensor. That's the entire mechanism - there's no local model, no weights to download, no GPU involved. The isolation happens on ElevenLabs' servers, so the quality ceiling is "whatever ElevenLabs' isolation model can do," which is a lot: it handles music bed, wind, traffic, and general room noise in a way that would take you an afternoon of spectral editing to match.
The inputs are refreshingly short:
- audio - required. The clip to clean.
- api_key - optional because it auto-detects
ELEVENLABS_API_KEY(or the DIGIT-specificDIGIT_ELEVENLABS_API_KEY).
One output, audio - the isolated voice track. Same type as the input, so you can chain it into anything that eats AUDIO: a Voice Clone node, an ElevenLabs Speech to Text node, a video mux, whatever your workflow uses for sound.
Install
Part of the 53-node comfyui-digit pack. ComfyUI Manager → search comfyui-digit → install → restart, or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Restart, and find it under DIGIT/ElevenLabs.
Common issues
Same auth story as every node in this suite:
export ELEVENLABS_API_KEY=sk_...
Forget it and you'll see ElevenLabs API key is required. Otherwise the practical complaints are about expectations. Voice isolation removes background - it doesn't fix a clip that's fundamentally distorted, and if the voice and the music are glued together in the same frequency range you'll hear artifacts on the voice itself. It's a one-click cleanup, not a miracle worker. If the result sounds robotic on a specific clip, that's the source material fighting you, not the node.
Also worth remembering: this is a paid API call per clip, and the whole clip goes to ElevenLabs. If you're isolating dialogue from a confidential client video, that's a data-leaves-the-machine decision you should make with your eyes open - same as any cloud audio tool.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| audio | AUDIO | — | |
| api_keyopt | STRING | ElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |