Audio Transcribe
Subtitles out of ComfyUI, no separate app required
- transcript
- segments_json
- segment_count
- transcribe_report
What it is
Speech-to-text via Whisper, from inside a ComfyUI graph. Give it an audio or video file path and it returns a transcript, a JSON array of timestamped segments, a segment count and a report string.
The reason to have this in a node graph rather than reaching for a subtitle app: the transcript is data. You can wire it into a prompt, feed the segments into a caption burn-in, time cuts to the speech, or check a generated voiceover against a script. That's a different job from "make me an SRT", and it's the job this node does.
The three backends and how Auto picks
backend defaults to Auto, which resolves in this order:
local_whisperif theopenai-whisperpackage is importable. Runs the model locally, most accurate, offline.openai_apiif a key resolves - the OpenAI Whisper API, which always useswhisper-1regardless ofmodel_size.whisper_cliif thewhispercommand is on your PATH, run as a subprocess.
The critical practical detail: openai-whisper is not in Radiance's requirements.txt. The pack detects it, it doesn't install it. So a stock Radiance install will never take the local path - Auto will fall through to the API if you've set a key, and otherwise to a CLI you probably don't have. If you want local transcription, install it yourself into ComfyUI's environment:
python -m pip install openai-whisper
Whisper's larger models download on first use and the large sizes want real VRAM. base is the default and it's a reasonable starting point; tiny is fine for testing the wiring, and large-v3 is the one you use when accuracy actually matters.
Inputs
audio_filepath is an absolute path to audio or video - it will pull the audio track itself. backend and language (16 codes plus auto, which lets Whisper detect the language) are the other required fields. model_size picks from tiny, base, small, medium, large, large-v2, large-v3.
On the optional side, openai_api_key_env (default OPENAI_API_KEY) names the environment variable holding your API key, and openai_api_key exists as a legacy fallback. Use the environment variable one - the tooltip recommends it explicitly, and it's right: a key typed into a widget is stored in the workflow JSON, which means it travels with every workflow you share. include_timings (default on) controls whether segments carry start/end times, and max_segment_chars splits long segments at a character count if you're building subtitle cards.
Outputs
transcript is the plain text. segments_json is the array of {start, end, text} objects in seconds - that's what you use for anything timed. segment_count is an integer, handy for a quick sanity check that the run wasn't empty. transcribe_report is the human-readable summary and, per the tooltip, the place failures are written instead of being raised. So read it: a node that completed green with an empty transcript is telling you something in the report, not throwing.
That's the opposite choice from some other nodes in this pack, which raise. Both behaviours exist here and it's worth knowing which node does which.
Install
ComfyUI Manager → search Radiance → install → restart → refresh the browser. Manually:
cd ComfyUI/custom_nodes
git clone https://github.com/fxtd-studios/radiance.git
cd radiance
python -m pip install -r requirements.txt
Windows portable users should run that pip command with ComfyUI's bundled python_embeded\python.exe. Then, if you want local inference, pip install openai-whisper into the same environment. The pack does declare transformers, but Whisper isn't a transformers pipeline here - it's the whisper package or the API.
Gotchas
- Auto not doing local inference. No
openai-whispermeans the local branch never runs. Checktranscribe_reportfor which backend actually ran. - Failures are silent-ish. They land in the report string, not in an exception. An empty transcript plus a report line is the normal failure shape.
- API keys in widgets. Use
openai_api_key_env. Anything you type intoopenai_api_keyis serialised into the workflow file you share. - First-run downloads. Local Whisper fetches model weights on first use; if you've set
RADIANCE_ALLOW_DOWNLOADS=0, that's a thing to be aware of in general, though Whisper's own download path is separate from the pack's model fetching. - Model size vs patience.
large-v3on a 30-minute file on CPU is a lunch break. Start atbaseand only go up if the transcript is genuinely too rough.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| audio_filepath | STRING | /path/to/audio.wav | Absolute path to audio/video file |
| backend | COMBO | Auto | Auto: local openai-whisper if installed, else the OpenAI API if a key resolves, else the whisper CLI on PATH. Failures are written to transcribe_report, not raised. |
| model_size | COMBO | base | Whisper model for local_whisper and whisper_cli (larger is slower and more accurate). openai_api always uses whisper-1. |
| language | COMBO | auto | Spoken language code passed to Whisper. auto lets Whisper detect it. |
| openai_api_key_envopt | STRING | OPENAI_API_KEY | Environment variable name containing the OpenAI API key. |
| openai_api_keyopt | STRING | Legacy fallback only. Prefer openai_api_key_env so workflows do not store secrets. | |
| include_timingsopt | BOOLEAN | true | Include per-segment timestamps in segments_json |
| max_segment_charsopt | INT | 0 | If > 0, split long segments at this character count |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| transcript | STRING | — |
| segments_json | STRING | — |
| segment_count | INT | — |
| transcribe_report | STRING | — |