Nodes/radiance/Audio Transcribe
ComfyUI Node

Audio Transcribe

Subtitles out of ComfyUI, no separate app required

By FXTD-Studios·Created 8 months ago·Updated about 18 hours ago· 246
Audio Transcribe
    • transcript
    • segments_json
    • segment_count
    • transcribe_report
    ◄audio_filepath/path/to/audio.wav►
    ◄backendAuto►
    ◄model_sizebase►
    ◄languageauto►
    ◄openai_api_key_envOPENAI_API_KEY►
    ◄openai_api_key►
    ◄include_timingstrue►
    ◄max_segment_chars0►

    What it is

    Speech-to-text via Whisper, from inside a ComfyUI graph. Give it an audio or video file path and it returns a transcript, a JSON array of timestamped segments, a segment count and a report string.

    The reason to have this in a node graph rather than reaching for a subtitle app: the transcript is data. You can wire it into a prompt, feed the segments into a caption burn-in, time cuts to the speech, or check a generated voiceover against a script. That's a different job from "make me an SRT", and it's the job this node does.

    The three backends and how Auto picks

    backend defaults to Auto, which resolves in this order:

    1. local_whisper if the openai-whisper package is importable. Runs the model locally, most accurate, offline.
    2. openai_api if a key resolves - the OpenAI Whisper API, which always uses whisper-1 regardless of model_size.
    3. whisper_cli if the whisper command is on your PATH, run as a subprocess.

    The critical practical detail: openai-whisper is not in Radiance's requirements.txt. The pack detects it, it doesn't install it. So a stock Radiance install will never take the local path - Auto will fall through to the API if you've set a key, and otherwise to a CLI you probably don't have. If you want local transcription, install it yourself into ComfyUI's environment:

    python -m pip install openai-whisper
    

    Whisper's larger models download on first use and the large sizes want real VRAM. base is the default and it's a reasonable starting point; tiny is fine for testing the wiring, and large-v3 is the one you use when accuracy actually matters.

    Inputs

    audio_filepath is an absolute path to audio or video - it will pull the audio track itself. backend and language (16 codes plus auto, which lets Whisper detect the language) are the other required fields. model_size picks from tiny, base, small, medium, large, large-v2, large-v3.

    On the optional side, openai_api_key_env (default OPENAI_API_KEY) names the environment variable holding your API key, and openai_api_key exists as a legacy fallback. Use the environment variable one - the tooltip recommends it explicitly, and it's right: a key typed into a widget is stored in the workflow JSON, which means it travels with every workflow you share. include_timings (default on) controls whether segments carry start/end times, and max_segment_chars splits long segments at a character count if you're building subtitle cards.

    Outputs

    transcript is the plain text. segments_json is the array of {start, end, text} objects in seconds - that's what you use for anything timed. segment_count is an integer, handy for a quick sanity check that the run wasn't empty. transcribe_report is the human-readable summary and, per the tooltip, the place failures are written instead of being raised. So read it: a node that completed green with an empty transcript is telling you something in the report, not throwing.

    That's the opposite choice from some other nodes in this pack, which raise. Both behaviours exist here and it's worth knowing which node does which.

    Install

    ComfyUI Manager → search Radiance → install → restart → refresh the browser. Manually:

    cd ComfyUI/custom_nodes
    git clone https://github.com/fxtd-studios/radiance.git
    cd radiance
    python -m pip install -r requirements.txt
    

    Windows portable users should run that pip command with ComfyUI's bundled python_embeded\python.exe. Then, if you want local inference, pip install openai-whisper into the same environment. The pack does declare transformers, but Whisper isn't a transformers pipeline here - it's the whisper package or the API.

    Gotchas

    • Auto not doing local inference. No openai-whisper means the local branch never runs. Check transcribe_report for which backend actually ran.
    • Failures are silent-ish. They land in the report string, not in an exception. An empty transcript plus a report line is the normal failure shape.
    • API keys in widgets. Use openai_api_key_env. Anything you type into openai_api_key is serialised into the workflow file you share.
    • First-run downloads. Local Whisper fetches model weights on first use; if you've set RADIANCE_ALLOW_DOWNLOADS=0, that's a thing to be aware of in general, though Whisper's own download path is separate from the pack's model fetching.
    • Model size vs patience. large-v3 on a 30-minute file on CPU is a lunch break. Start at base and only go up if the transcript is genuinely too rough.
    CategoryFXTD STUDIOS/Radiance/Pipeline

    Inputs (8)

    NameTypeDefaultDescription
    audio_filepathSTRING/path/to/audio.wavAbsolute path to audio/video file
    backendCOMBOAutoAuto: local openai-whisper if installed, else the OpenAI API if a key resolves, else the whisper CLI on PATH. Failures are written to transcribe_report, not raised.
    model_sizeCOMBObaseWhisper model for local_whisper and whisper_cli (larger is slower and more accurate). openai_api always uses whisper-1.
    languageCOMBOautoSpoken language code passed to Whisper. auto lets Whisper detect it.
    openai_api_key_envoptSTRINGOPENAI_API_KEYEnvironment variable name containing the OpenAI API key.
    openai_api_keyoptSTRINGLegacy fallback only. Prefer openai_api_key_env so workflows do not store secrets.
    include_timingsoptBOOLEANtrueInclude per-segment timestamps in segments_json
    max_segment_charsoptINT0If > 0, split long segments at this character count

    Outputs (4)

    NameTypeDescription
    transcriptSTRING—
    segments_jsonSTRING—
    segment_countINT—
    transcribe_reportSTRING—