DIGIT ElevenLabs Speech to Text
Speech to text with speaker labels and word timestamps, inside the graph
- audio
- text
- language_code
- words_json
The DIGIT ElevenLabs Speech to Text node transcribes audio using ElevenLabs' scribe_v2 model - and the reason you'd pick it over a free local transcription tool is the extras: optional speaker diarization, word-level timestamps, and audio-event tagging, all returned as usable graph outputs. It's the transcription node for when "here's the words" isn't quite enough and you want the structure too.
In practice that means: transcribe a scratch VO to feed back into a TTS node as a refined script, turn a recorded interview into text for captioning, or get word-level timing out of a clip for subtitle work. The outputs are plain STRINGs, so they wire anywhere.
How it works
Feed it audio (AUDIO tensor - from the pack's TTS node, or loaded from a file) and it transcribes with scribe_v2. The useful extras live in the optional inputs:
diarize- annotate which speaker is talking, withnum_speakers(0 = auto-detect, up to 32) anddiarization_threshold(0.22) controlling how aggressively speakers are split.timestamps_granularity-word,character, ornone. Word-level is what you want for subtitles.tag_audio_events- annotate sounds like(laughter)in the transcript. Cheaper than a separate sound-design pass for knowing what was happening.language_code- set it for a specific language, or leave empty for auto-detect.temperature(0) - keep it low for transcription; you don't want creativity here.seed- reproducibility.
Outputs: text (the transcript), language_code (what it detected/used), and words_json - the word-level data as JSON when you've asked for timestamps, which is the thing that makes downstream timing work possible.
api_key is optional, auto-detected from ELEVENLABS_API_KEY or the pack's DIGIT_ELEVENLABS_API_KEY.
Installing it
Standard pack install:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Or ComfyUI Manager → search comfyui-digit → install → restart. Then:
export ELEVENLABS_API_KEY=your_key_here
Using it well
Diarization is the feature that makes this more than a toy, and it's worth tuning: num_speakers auto-detects, but if your two-person recording keeps collapsing into one speaker (or inventing a third), set it explicitly. The pairing that gets the most mileage: STT a recorded line, then feed the cleaned transcript into the pack's ElevenLabs TTS or Dialogue node - you've effectively turned an existing performance into a script you can re-voice. And since it's a hosted, per-call service, one practical tip: get the transcript right in one pass by setting language_code when you know it, rather than burning calls on auto-detect that gets it wrong.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| audio | AUDIO | — | |
| model | COMBO | scribe_v2 | 1 options: scribe_v2 |
| seed | INT | 10–2147483647 | — |
| api_keyopt | STRING | ElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var. | |
| language_codeopt | STRING | ISO-639-1/3 language code. Leave empty for auto-detect. | |
| tag_audio_eventsopt | BOOLEAN | false | Annotate sounds like (laughter) in transcript. |
| diarizeopt | BOOLEAN | false | Annotate which speaker is talking. |
| diarization_thresholdopt | FLOAT | 0.220.1–0.4 | — |
| temperatureopt | FLOAT | 0.000–2 | — |
| timestamps_granularityopt | COMBO | word | 3 options: word, character, none |
| num_speakersopt | INT | 00–32 | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| text | STRING | — |
| language_code | STRING | — |
| words_json | STRING | — |