DIGIT ElevenLabs Text to Speech
DIGIT ElevenLabs TTS inside your graph
- audio
You want narration on a generated video, a voice line for a character, or a text-to-speech pass that doesn't sound like a cheap text-to-speech. ElevenLabs is the bar that every open TTS model gets compared against, and this node is the direct route to it: paste a voice ID, feed it text, and out comes a ComfyUI AUDIO tensor you can wire into a video, a preview, or an audio saver.
The important part is direct. Plenty of ElevenLabs wrappers route through a third-party proxy or a wrapper service that marks up the API and gets to read your text. DIGIT calls api.elevenlabs.io/v1/text-to-speech/<voice_id> straight from the node with your own key, so you pay ElevenLabs' list price and nothing else. No middleman, no extra rate limits beyond your account's.
How it works
The node POSTs your text plus voice settings to ElevenLabs' text-to-speech endpoint and converts the response into an AUDIO tensor. By default that's PCM at 44.1kHz - clean, lossless, and ready for mixing - but you can switch output_format to mp3_44100_192 or opus_48000_192 if you're saving a file and want it smaller.
The inputs that matter most:
- text - the line to speak. Multiline, so paste a paragraph.
- voice_id - connect the Voice Selector node, or paste a raw ElevenLabs voice ID. It's required; the node errors if it's blank.
- model -
eleven_multilingual_v2(the default, great across languages) oreleven_v3(the newer one). - stability and similarity_boost - the classic ElevenLabs sliders. Lower stability gets more emotional variation; higher similarity keeps the voice closer to its reference.
- speed - 0.7 to 1.3, so you can rush or drawl the line.
- seed - ElevenLabs now takes a seed, so you can get the same take back on re-runs.
Beyond that there's language_code (leave empty for auto-detect), style for a touch of expressiveness, and use_speaker_boost for presence on loud mixes. The api_key field is optional in the UI because the node auto-detects ELEVENLABS_API_KEY - and it also checks DIGIT_ELEVENLABS_API_KEY, which is the DIGIT pack's own convention.
The output is a single audio output of type AUDIO. That's what you plug into whatever consumes audio in your workflow.
Install
The pack ships 53 nodes, so you get the whole DIGIT family whether you wanted it or not. Install via ComfyUI Manager (search comfyui-digit) and restart, or:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Restart ComfyUI. The ElevenLabs nodes live under the DIGIT/ElevenLabs category.
Common issues
The classic gotcha is auth. This node needs a real ElevenLabs key - there's no free tier hidden in here. Set it before you launch ComfyUI:
export ELEVENLABS_API_KEY=sk_...
You'll see ElevenLabs API key is required if you haven't, and the fix is just the env var (or pasting the key into the node). The other one people trip on: voice_id must be filled in. The Voice Selector makes this painless, but if you're typing it by hand, a mistyped ID throws a 404-ish error from the API - check that before you blame the node.
One honest caveat about any cloud TTS: your text goes to ElevenLabs' servers. Fine for most narration work, worth remembering if the line is confidential.
This is a per-call service - a few thousand characters of speech won't break the bank, but it isn't free per use like a local model. For that tradeoff you get the best-in-class voice quality, which is exactly why people keep reaching for it.
Inputs (13)
| Name | Type | Default | Description |
|---|---|---|---|
| text | STRING | Text to convert to speech. | |
| voice_id | STRING | Voice ID. Connect from Voice Selector or paste directly. | |
| model | COMBO | eleven_multilingual_v2 | 2 options: eleven_multilingual_v2, eleven_v3 |
| stability | FLOAT | 0.500–1 | — |
| similarity_boost | FLOAT | 0.750–1 | — |
| speed | FLOAT | 1.000.7–1.3 | — |
| seed | INT | 10–2147483647 | — |
| api_keyopt | STRING | ElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var. | |
| language_codeopt | STRING | ISO-639-1/3 language code. Leave empty for auto-detect. | |
| styleopt | FLOAT | 0.000–0.2 | — |
| use_speaker_boostopt | BOOLEAN | false | — |
| apply_text_normalizationopt | COMBO | auto | 3 options: auto, on, off |
| output_formatopt | COMBO | pcm_44100 | 3 options: pcm_44100, mp3_44100_192, opus_48000_192 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |