ComfyUI Node
Transcribe Audio
Recognize speech or sung words with the local Whisper large-v3 model. Outputs text and full-audio timing for review, subtitles or lyric captions. No supplied lyrics are required.
Transcribe Audio
- audio
- vocals
- transcript
- timing_json
- srt
- report
◄languageAuto►
◄deviceauto►
Category🌒 Eclipse/ Audio
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| audio | AUDIO | Complete recording. Output timestamps start at this audio's time zero. | |
| language | COMBO | Auto | 101 options: Auto, en, zh, de, es, ru, +95 |
| device | COMBO | auto | 3 options: auto, cpu, cuda |
| vocalsopt | AUDIO | Optional isolated vocals with the same time zero and duration as audio (within 50 ms). |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| transcript | STRING | — |
| timing_json | STRING | — |
| srt | STRING | — |
| report | STRING | — |