π€ Audio Recorder Gemini v2
The Recorder Node That Blocks Your Queue
- audio
What it's for
This one is a microphone. You record into the node, it hands out an AUDIO clip, and you wire that into whatever consumes audio - most obviously the pack's Live Audio Chat node, or a local speech-to-text setup, or a lip-sync pipeline where you'd rather talk than type.
It's the least glamorous node in the pack and the one with the most surprising behaviour, because recording is a blocking, hardware-touching operation sitting inside a system built around queued, mostly-deterministic work.
How it works
It opens a mono float32 input stream through sounddevice at your chosen sample_rate, reading in 100 ms blocks. Then it watches the levels. Nothing happens until a block peaks above silence_threshold - that's the "heard voice" gate - and after that, once it sees silence_duration seconds of quiet, it stops and trims the trailing silence, keeping about 200 ms of tail. max_duration is the hard stop for everything else.
Two consequences fall straight out of that design. First, if you never speak above the threshold it will not stop early: it runs until max_duration and gives you whatever it captured. Second, if your room noise sits above the threshold, the node thinks you're talking the whole time and you get a full-length recording of your fan. The default threshold of 0.01 is quiet-room territory - raise it if you're recording next to a gaming PC.
device is a dropdown built at load time from your system's input devices, listed as index: name, and Default means "let the OS decide". trigger is the counter behind the record button: the node caches the last recording against the trigger value, so if you re-queue without clicking the button again you get the same audio back rather than a fresh take. That's a feature a lot of people mistake for a bug.
Inputs and outputs
Required: device, sample_rate (default 44100), silence_threshold, silence_duration (default 2 s), max_duration (default 10 s), trigger. One output: audio.
The start control is a native π΄ Start Record button drawn on the node canvas. It's a UI-only control: clicking it bumps the hidden trigger value and queues the current workflow for you. In earlier builds the recorder could die with prompt_no_outputs when it was the only output in the graph - the class is now declared as an output node precisely so the queue accepts it, and the button has duplicate-click protection while a run is already queued.
Install
Manager, searching the display name ComfyUI Gemini 3x Pro - or the manual clone:
cd ComfyUI/custom_nodes
git clone https://github.com/asirusasr-maker/ComfyUI-Gemini_3x_Pro
This node is the reason sounddevice is in requirements.txt, and sounddevice alone gets you a device list of exactly Default. Install the pack's dependencies with the interpreter that runs ComfyUI:
python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\ComfyUI-Gemini_3x_Pro\requirements.txt
Add more than one device to the dropdown and you'll want a restart - the list is built when the node class loads, not when you click. There's no API key involved in recording; you only need one for the node downstream.
On Linux, sounddevice needs PortAudio under it (libportaudio2 on Debian/Ubuntu). If speech-to-text or a waveform loader downstream wants a specific format, note that torchaudio is optional and not installed by this pack - it's used for audio handling elsewhere, and it's absent from requirements.txt so a portable CUDA environment doesn't get overwritten.
Where people get burned
Recording happens on the machine running ComfyUI. If you're on a remote box, a cloud instance, or a container, the node records that machine's microphone - which is to say, nothing. There is no browser-side capture here.
It blocks the queue while you talk. The stream read is synchronous, so for the full recording window ComfyUI's execution thread is busy and nothing else in the queue moves. With a 120-second max_duration and a bad threshold, that's two minutes of a stalled queue. Keep max_duration tight.
Headless Linux has no devices. A container without /dev/snd passthrough will only ever offer Default, and the record call fails into a silent one-frame clip rather than an exception. The failure is printed to the ComfyUI console as [Gemini Audio Recorder] Error: ... - that log line is your only clue, because the node still returns an AUDIO object.
A silent take isn't always an error. Empty audio can be a threshold problem, a muted mic, or the wrong entry in the device dropdown. Check the console line reporting the captured length in seconds before you go hunting in the API key settings - the recorder never uses one.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| device | COMBO | Default | 1 options: Default |
| sample_rate | INT | 441008000β96000 | β |
| silence_threshold | FLOAT | 0.0100.001β0.1 | β |
| silence_duration | FLOAT | 2.00.5β5 | β |
| max_duration | FLOAT | 10.01β120 | β |
| trigger | INT | 0 | β |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | β |