Nodes/ComfyUI/ElevenLabs Speech to Text
ComfyUI Node Runs on cloud

ElevenLabs Speech to Text

Transcribe audio to text. Supports automatic language detection, speaker diarization, and audio event tagging.

By Comfy-Org·Created 4 years ago·Updated 20 days ago· 121,575
ElevenLabs Speech to Text
  • audio
  • text
  • language_code
  • words_json
model
language_code
num_speakers0
seed1
Categorypartner/audio/ElevenLabs

Inputs (5)

NameTypeDefaultDescription
audioAUDIOAudio to transcribe.
modelCOMBOModel to use for transcription.
language_codeSTRINGISO-639-1 or ISO-639-3 language code (e.g., 'en', 'es', 'fra'). Leave empty for automatic detection.
num_speakersINT00–32Maximum number of speakers to predict. Set to 0 for automatic detection.
seedINT10–2147483647Seed for reproducibility (determinism not guaranteed).

Outputs (3)

NameTypeDescription
textSTRING
language_codeSTRING
words_jsonSTRING