TG Send Audio ◀️
Voice note, MP3, or a WAV nobody will play inline
- audio
- trigger
- message
- message_id
- trigger
An AUDIO object in ComfyUI - the thing every TTS and music node emits - isn't a file. It's a tensor plus a sample rate. Telegram wants bytes in a container. This node is the conversion layer, and it gives you three delivery styles because Telegram genuinely treats them as three different kinds of message.
The three ways out
send_as is the whole decision:
- Voice →
sendVoicewith Ogg/Opus. This is the one that renders as a proper voice note with a waveform you can scrub, the way a message recorded in the app looks. - Audio →
sendAudiowith MP3. A music player bubble with metadata-ish presentation. - File →
sendDocumentwith WAV. Lossless, honest, and it arrives as a downloadable file with no inline player.
Mechanically it does this: your tensor is written to WAV bytes first, then, unless you picked File, re-encoded through PyAV to MP3 or Ogg. That's the av>=14.2.0 dependency in requirements.txt - it's already a core ComfyUI dependency for video work, its wheels bundle the codec libraries, and there is no ffmpeg executable or PATH change involved. If PyAV is somehow missing you get an explicit "install this pack's requirements.txt with ComfyUI's Python" error rather than a stack trace.
The WAV step is hand-written rather than delegated to torchaudio - the pack's own comment says it's to dodge torchaudio/torchcodec ABI breakage on recent Torch builds. Small detail, but it's why this node doesn't randomly explode after a Torch upgrade, which is more than you can say for a lot of audio nodes.
Inputs worth naming
audio is a socket (link required) and it takes one item: a batched AUDIO tensor raises Send Audio accepts one AUDIO item (batch size 1). If you've got a folder of takes, deal them out one run at a time - which under Run (Instant) is exactly what the Listener is already doing for you.
chat_id comes from the Listener's chat_id output. file_name defaults to audio and the extension is added for you (.ogg, .mp3 or .wav). caption is 0–1024 characters, parse_mode applies to it, and show_caption_above_media flips it above the player.
Two booleans worth checking before you blame the node: disable_notification defaults to True, so a voice reply arrives silently; and protect_content blocks forwarding and saving, which matters if you're sending someone's personal TTS. message_thread_id takes the Listener's thread ID (-1 = none, non-positive values get stripped from the request).
trigger is the untyped ordering input - pass it through to sequence this send before whatever comes next.
Outputs
message (Telegram response DICT), message_id (INT, feed it to Edit Message Audio when you want to replace the audio in place), trigger. On a total network failure after five retries, message and message_id turn into silent blockers so nothing edits an imaginary message, and the run keeps going.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/CoolBreeze164/ComfyUI-Autonomous-Telegram-Bot
python -m pip install -r ComfyUI-Autonomous-Telegram-Bot/requirements.txt
Windows portable: python_embeded\python.exe -m pip install ... from the portable root. Deps: httpx, numpy, pillow, av>=14.2.0. Manager search: ComfyUI Autonomous Telegram Bot. Restart ComfyUI afterwards.
Common issues
- Voice note won't play / sounds wrong - resampling happens on the way to Opus, and a mono input is what Telegram expects for voice. Feed it a clean single-channel result from your TTS node.
- It's silent on the other end - check
disable_notification, then check you didn't wire an empty AUDIO branch. If the Listener couldn't get audio out of the incoming message, itsmessage_audiooutput is a blocker and this node's branch is skipped entirely, which looks identical to "it did nothing". - Batch error - one AUDIO item per run, as above.
- A 60-second clip takes a while - every audio send goes up as a file upload; the pack allows a 120-second read per attempt and up to five attempts. Music-length files over a slow connection will feel sticky.
- Rate limits - Telegram asks for under one message per second per chat; the pack honours the
retry_afterit's given, which can mean a visible pause.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| bot_token | STRING | Telegram bot token from BotFather | |
| chat_id | INT | Unique identifier for the target chat | |
| audio | AUDIO | the audio to send | |
| caption | STRING | Media caption, 0-1024 characters after entities parsing | |
| parse_mode | COMBO | Mode for parsing entities in the photo caption. See https://core.telegram.org/bots/api#formatting-options for more details. | |
| show_caption_above_media | BOOLEAN | false | Pass True, if the caption must be shown above the message media |
| disable_notification | BOOLEAN | true | Sends the message silently. Users will receive a notification with no sound. |
| protect_content | BOOLEAN | false | Protects the contents of the sent message from forwarding and saving |
| send_as | COMBO | How to send the audio | |
| file_name | STRING | audio | the file name for this media |
| message_thread_id | INT | -1-1–9223372036854776000 | Unique identifier for the target message thread of the forum topic (-1 = None) |
| triggeropt | * | Optional trigger to enforce execution order |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| message | DICT | — |
| message_id | INT | — |
| trigger | * | — |