Nodes/DIGIT Nodes/DIGIT ElevenLabs Speech to Text
ComfyUI Node

DIGIT ElevenLabs Speech to Text

Speech to text with speaker labels and word timestamps, inside the graph

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT ElevenLabs Speech to Text
  • audio
  • text
  • language_code
  • words_json
◄modelscribe_v2►
◄seed1►
◄api_key►
◄language_code►
◄tag_audio_eventsfalse►
◄diarizefalse►
◄diarization_threshold0.22►
◄temperature0.00►
◄timestamps_granularityword►
◄num_speakers0►

The DIGIT ElevenLabs Speech to Text node transcribes audio using ElevenLabs' scribe_v2 model - and the reason you'd pick it over a free local transcription tool is the extras: optional speaker diarization, word-level timestamps, and audio-event tagging, all returned as usable graph outputs. It's the transcription node for when "here's the words" isn't quite enough and you want the structure too.

In practice that means: transcribe a scratch VO to feed back into a TTS node as a refined script, turn a recorded interview into text for captioning, or get word-level timing out of a clip for subtitle work. The outputs are plain STRINGs, so they wire anywhere.

How it works

Feed it audio (AUDIO tensor - from the pack's TTS node, or loaded from a file) and it transcribes with scribe_v2. The useful extras live in the optional inputs:

  • diarize - annotate which speaker is talking, with num_speakers (0 = auto-detect, up to 32) and diarization_threshold (0.22) controlling how aggressively speakers are split.
  • timestamps_granularity - word, character, or none. Word-level is what you want for subtitles.
  • tag_audio_events - annotate sounds like (laughter) in the transcript. Cheaper than a separate sound-design pass for knowing what was happening.
  • language_code - set it for a specific language, or leave empty for auto-detect.
  • temperature (0) - keep it low for transcription; you don't want creativity here.
  • seed - reproducibility.

Outputs: text (the transcript), language_code (what it detected/used), and words_json - the word-level data as JSON when you've asked for timestamps, which is the thing that makes downstream timing work possible.

api_key is optional, auto-detected from ELEVENLABS_API_KEY or the pack's DIGIT_ELEVENLABS_API_KEY.

Installing it

Standard pack install:

cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt

Or ComfyUI Manager → search comfyui-digit → install → restart. Then:

export ELEVENLABS_API_KEY=your_key_here

Using it well

Diarization is the feature that makes this more than a toy, and it's worth tuning: num_speakers auto-detects, but if your two-person recording keeps collapsing into one speaker (or inventing a third), set it explicitly. The pairing that gets the most mileage: STT a recorded line, then feed the cleaned transcript into the pack's ElevenLabs TTS or Dialogue node - you've effectively turned an existing performance into a script you can re-voice. And since it's a hosted, per-call service, one practical tip: get the transcript right in one pass by setting language_code when you know it, rather than burning calls on auto-detect that gets it wrong.

CategoryDIGIT/ElevenLabs

Inputs (11)

NameTypeDefaultDescription
audioAUDIO—
modelCOMBOscribe_v21 options: scribe_v2
seedINT10–2147483647—
api_keyoptSTRINGElevenLabs API key. Auto-detected from ELEVENLABS_API_KEY env var.
language_codeoptSTRINGISO-639-1/3 language code. Leave empty for auto-detect.
tag_audio_eventsoptBOOLEANfalseAnnotate sounds like (laughter) in transcript.
diarizeoptBOOLEANfalseAnnotate which speaker is talking.
diarization_thresholdoptFLOAT0.220.1–0.4—
temperatureoptFLOAT0.000–2—
timestamps_granularityoptCOMBOword3 options: word, character, none
num_speakersoptINT00–32—

Outputs (3)

NameTypeDescription
textSTRING—
language_codeSTRING—
words_jsonSTRING—