Nodes/ComfyUI-NanoBanana2/NanoBanana - Audio Transcribe
ComfyUI Node

NanoBanana - Audio Transcribe

Turn any ComfyUI audio into a timestamped transcript

By IxMxAMAR·Created 7 months ago·Updated 9 days ago· 4
NanoBanana - Audio Transcribe
  • audio
  • network
  • transcript
◄api_key►
◄modelgemini-3.8-flash►
◄custom_model►
◄promptTranscribe this audio verbatim. Include speaker labels if multiple speakers are present.►
◄include_timestampsfalse►
◄language_codes►
◄custom_vocabulary►
◄diarizationfalse►
◄transcription_modeVERBATIM►

ComfyUI generates a lot of things, and apparently one of them is now subtitles. This node takes any standard ComfyUI AUDIO object - a voice recording, a TTS output, a clip from a video workflow, even the music node's output if you're curious what Lyria "sang" - encodes it to WAV, and sends it to Gemini's flash model to transcribe. What comes back is a text transcript, with optional [HH:MM:SS] timestamps on every line.

Where it earns its place: captioning a TTS-generated narration so you can burn subtitles into a video, transcribing a recording you loaded in for a dubbing workflow, or just getting searchable text out of an audio asset that's already sitting in your graph. It's the text-side companion to the pack's TTS nodes - speak, then transcribe, then stitch captions in.

How it works

The mechanism is a single multimodal call. Your AUDIO (waveform + sample rate) gets encoded to WAV bytes and sent to Gemini as an audio part, alongside a transcription prompt. If you flip include_timestamps on, the prompt becomes "transcribe with timestamps in [HH:MM:SS] format at the start of each line" and Gemini does exactly that. Temperature is pinned to 0.0 - for transcription you want deterministic, not creative. The model defaults to gemini-2.5-flash, which is cheap and handles this well.

Inputs:

  • audio - the ComfyUI AUDIO to transcribe.
  • model - default gemini-2.5-flash.
  • prompt - default is "transcribe verbatim, include speaker labels if multiple speakers" - and yes, it genuinely attempts speaker labels on multi-speaker audio.
  • include_timestamps - off by default; flip it when you need subtitle timing.

Output: transcript, a plain STRING - wire it to a text display, a save node, or into a subtitle-generation step.

Installation

Part of the NanoBanana2 pack:

cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-NanoBanana2
pip install google-genai

ComfyUI Manager: search NanoBanana2. Needs a Gemini API key from aistudio.google.com.

Gotchas

Two limits to know before you throw an hour-long podcast at it. First, this node encodes whatever AUDIO you hand it into a single WAV and sends the whole thing - long audio means a big request, and the pack's own docs suggest using the Files Upload node + a file-based vision node for genuinely long recordings instead. Keep this node for clips measured in seconds to a few minutes. Second, it's a paid API call, and audio tokens add up fast - a long clip is not a free transcription. If you get a refusal or a garbled result, remember transcriptions are generative too: Gemini can hallucinate on noisy audio, so for critical text, keep the original and spot-check.

CategoryNanoBanana2/Audio

Inputs (11)

NameTypeDefaultDescription
api_keySTRING—
modelCOMBOgemini-3.8-flash35 options: gemini-pro-latest, gemini-flash-latest, gemini-flash-lite-latest, gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, +29
audioAUDIO—
custom_modeloptSTRING—
promptoptSTRINGTranscribe this audio verbatim. Include speaker labels if multiple speakers are present.—
include_timestampsoptBOOLEANfalseNot used by gemini-3.5-transcribe.
networkoptNB_NETWORKOptional. Wire a NanoBanana - Network Route node here to route this request through that proxy (e.g. US egress).
language_codesoptSTRINGgemini-3.5-transcribe only. Comma-separated BCP-47 codes (e.g. en-US,es-ES). Empty = auto-detect.
custom_vocabularyoptSTRINGgemini-3.5-transcribe only. Comma-separated terms to bias recognition (up to 1000). Cannot be combined with diarization.
diarizationoptBOOLEANfalsegemini-3.5-transcribe only. Label up to 8 distinct speakers.
transcription_modeoptCOMBOVERBATIMgemini-3.5-transcribe only. SMART removes disfluencies; it cannot be combined with diarization.

Outputs (1)

NameTypeDescription
transcriptSTRING—