Nodes/ComfyUI-Gemini_3x_Pro/πŸŽ™οΈ Gemini Live Audio Chat v2
ComfyUI Node

πŸŽ™οΈ Gemini Live Audio Chat v2

A Real Voice Session, Queued Like Any Other Node

By asirusasr-makerΒ·Created 2 months agoΒ·Updated a day agoΒ· 6
πŸŽ™οΈ Gemini Live Audio Chat v2
  • audio_input
  • audio_output
  • text_response
  • session_info
β—„text_inputHello! How are you?β–Ί
β—„system_promptYou are a helpful voice assistant.β–Ί
β—„modelgemini-3.8-liveβ–Ί
β—„voiceKoreβ–Ί
β—„api_keyβ–Ί
β—„proxyβ–Ί
β—„fallback_enabledtrueβ–Ί

What you'd use it for

You want a spoken reply, not a sound file rendered from a script. Give this node text, or a recording from your mic, and it talks back - Gemini's Live API, Google's low-latency conversational endpoint, over a WebSocket. Then you keep the voice as an AUDIO clip and the words as a transcript.

In practice that lands in two places. One is a talking-head pipeline: the transcript and the audio go into whatever drives the face, and because you get both, you can build subtitles or captions off the same run. The other is prototyping a voice assistant inside the graph, where you can see the audio and the text side by side before committing to a local TTS stack.

Set expectations on length. A local TTS model renders a script and costs you nothing per line; this is a hosted session, one per execution. For a fixed narrator across a 40-clip project, run Chatterbox locally. For "say this one thing and sound like a person," this is a lot less setup.

How it works

Each execution opens a short-lived Live session, sends, listens, and closes. Nothing persists between runs - no memory of the last thing it said, even if you re-queue the same prompt with a follow-up.

Inside the session: output modality is AUDIO with output transcription enabled, and your system_prompt becomes the session's system instruction. If you supplied audio, it's converted to 16 kHz mono PCM16 and pushed as realtime input. If you supplied audio and text, the text goes in as realtime text after the audio, so it reads as context rather than as a second turn. Text alone is sent as a normal user turn. The node then reads back message frames, concatenates the audio chunks it receives, collects the transcription, and stops when the server reports the turn complete. Audio comes out decoded as 24 kHz mono PCM16 into ComfyUI's AUDIO format.

One implementation detail worth knowing if you're debugging: the Live session is async, and ComfyUI already runs nodes inside an event loop, so the node runs the session in a worker thread with a 180-second cap. A session that hangs past three minutes is killed rather than wedging your queue forever.

Voice choice is the pack's eight prebuilt voices - Puck, Charon, Kore, Fenrir, Leda, Orus, Aoede, Callirhoe, default Kore. fallback_enabled tries the other Live model when the first shows a transient or model-unavailable failure.

Inputs and outputs

Required: text_input, system_prompt, model (either gemini-3.8-live or gemini-3.8-live-extended-thinking). Optional: audio_input, voice, api_key, proxy, fallback_enabled.

Three outputs, all useful. audio_output goes to any AUDIO consumer - a save node, or the audio input of a lip-sync workflow. text_response is the model's output transcription as a plain string, so it plugs straight into a text node or a caption writer. session_info is JSON: requested model, actual model, whether fallback fired, the voice used, whether audio input was present, and a status field. It's also the only place a failure tells you what happened, which matters - see below.

Install

Manager, if the registry has it: look for the display name ComfyUI Gemini 3x Pro. Otherwise the documented route is a clone:

cd ComfyUI/custom_nodes
git clone https://github.com/asirusasr-maker/ComfyUI-Gemini_3x_Pro

Dependencies are google-genai>=2.27.0,<3.0, pillow, numpy and sounddevice - the WebSocket client comes in with google-genai, you don't install websockets by hand. Portable Windows users install with the embedded interpreter:

python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-Gemini_3x_Pro\requirements.txt

Key resolution, same as every node in the pack: the node's api_key field first, then GEMINI_API_KEY in the pack's config.json, then the GEMINI_API_KEY environment variable. Then restart ComfyUI.

When it doesn't work

A failed session is silent. The node returns a one-frame zero waveform plus the error text in text_response, and it does not raise. Your workflow completes and hands a downstream node a quarter-second of silence. Read session_info when a clip comes out empty - that's the tell between "the API refused" and "your graph is fine."

Blocked WebSockets look like a broken node. The Live API needs outbound WSS through whatever network or proxy sits in front of ComfyUI. Corporate networks, some VPNs and most restrictive firewalls kill it while the REST-based nodes in this pack keep working fine, which makes it look like a bug in this node specifically. It isn't. Note also that proxy here isn't an HTTP proxy - it sets the client's base URL, so it points at an alternate Gemini endpoint, not at an egress proxy.

Don't type the key into the node field on a shared workflow. It's a plain string widget and it gets serialised into the JSON you post. Use config.json or the environment variable.

It is not a conversation. Re-queue and you start fresh. If you want a thread that remembers the last exchange, that's the multimodal node's chat_mode, and even there the memory lives in memory and dies with ComfyUI.

CategoryGemini 3.x

Inputs (8)

NameTypeDefaultDescription
text_inputSTRINGHello! How are you?β€”
system_promptSTRINGYou are a helpful voice assistant.β€”
modelCOMBOgemini-3.8-live2 options: gemini-3.8-live, gemini-3.8-live-extended-thinking
audio_inputoptAUDIOβ€”
voiceoptCOMBOKore8 options: Puck, Charon, Kore, Fenrir, Leda, Orus, +2
api_keyoptSTRINGβ€”
proxyoptSTRINGβ€”
fallback_enabledoptBOOLEANtrueβ€”

Outputs (3)

NameTypeDescription
audio_outputAUDIOβ€”
text_responseSTRINGβ€”
session_infoSTRINGβ€”