ComfyUI Node

TG Send Audio ◀️

Voice note, MP3, or a WAV nobody will play inline

By CoolBreeze164·Created 13 days ago·Updated 2 days ago· 3
TG Send Audio ◀️
  • audio
  • trigger
  • message
  • message_id
  • trigger
◄bot_token►
◄chat_id—►
◄caption►
◄parse_mode▾►
◄show_caption_above_mediafalse►
◄disable_notificationtrue►
◄protect_contentfalse►
◄send_as▾►
◄file_nameaudio►
◄message_thread_id-1►

An AUDIO object in ComfyUI - the thing every TTS and music node emits - isn't a file. It's a tensor plus a sample rate. Telegram wants bytes in a container. This node is the conversion layer, and it gives you three delivery styles because Telegram genuinely treats them as three different kinds of message.

The three ways out

send_as is the whole decision:

  • Voice → sendVoice with Ogg/Opus. This is the one that renders as a proper voice note with a waveform you can scrub, the way a message recorded in the app looks.
  • Audio → sendAudio with MP3. A music player bubble with metadata-ish presentation.
  • File → sendDocument with WAV. Lossless, honest, and it arrives as a downloadable file with no inline player.

Mechanically it does this: your tensor is written to WAV bytes first, then, unless you picked File, re-encoded through PyAV to MP3 or Ogg. That's the av>=14.2.0 dependency in requirements.txt - it's already a core ComfyUI dependency for video work, its wheels bundle the codec libraries, and there is no ffmpeg executable or PATH change involved. If PyAV is somehow missing you get an explicit "install this pack's requirements.txt with ComfyUI's Python" error rather than a stack trace.

The WAV step is hand-written rather than delegated to torchaudio - the pack's own comment says it's to dodge torchaudio/torchcodec ABI breakage on recent Torch builds. Small detail, but it's why this node doesn't randomly explode after a Torch upgrade, which is more than you can say for a lot of audio nodes.

Inputs worth naming

audio is a socket (link required) and it takes one item: a batched AUDIO tensor raises Send Audio accepts one AUDIO item (batch size 1). If you've got a folder of takes, deal them out one run at a time - which under Run (Instant) is exactly what the Listener is already doing for you.

chat_id comes from the Listener's chat_id output. file_name defaults to audio and the extension is added for you (.ogg, .mp3 or .wav). caption is 0–1024 characters, parse_mode applies to it, and show_caption_above_media flips it above the player.

Two booleans worth checking before you blame the node: disable_notification defaults to True, so a voice reply arrives silently; and protect_content blocks forwarding and saving, which matters if you're sending someone's personal TTS. message_thread_id takes the Listener's thread ID (-1 = none, non-positive values get stripped from the request).

trigger is the untyped ordering input - pass it through to sequence this send before whatever comes next.

Outputs

message (Telegram response DICT), message_id (INT, feed it to Edit Message Audio when you want to replace the audio in place), trigger. On a total network failure after five retries, message and message_id turn into silent blockers so nothing edits an imaginary message, and the run keeps going.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/CoolBreeze164/ComfyUI-Autonomous-Telegram-Bot
python -m pip install -r ComfyUI-Autonomous-Telegram-Bot/requirements.txt

Windows portable: python_embeded\python.exe -m pip install ... from the portable root. Deps: httpx, numpy, pillow, av>=14.2.0. Manager search: ComfyUI Autonomous Telegram Bot. Restart ComfyUI afterwards.

Common issues

  • Voice note won't play / sounds wrong - resampling happens on the way to Opus, and a mono input is what Telegram expects for voice. Feed it a clean single-channel result from your TTS node.
  • It's silent on the other end - check disable_notification, then check you didn't wire an empty AUDIO branch. If the Listener couldn't get audio out of the incoming message, its message_audio output is a blocker and this node's branch is skipped entirely, which looks identical to "it did nothing".
  • Batch error - one AUDIO item per run, as above.
  • A 60-second clip takes a while - every audio send goes up as a file upload; the pack allows a 120-second read per attempt and up to five attempts. Music-length files over a slow connection will feel sticky.
  • Rate limits - Telegram asks for under one message per second per chat; the pack honours the retry_after it's given, which can mean a visible pause.
CategoryAutonomous Telegram Bot ◀️

Inputs (12)

NameTypeDefaultDescription
bot_tokenSTRINGTelegram bot token from BotFather
chat_idINTUnique identifier for the target chat
audioAUDIOthe audio to send
captionSTRINGMedia caption, 0-1024 characters after entities parsing
parse_modeCOMBOMode for parsing entities in the photo caption. See https://core.telegram.org/bots/api#formatting-options for more details.
show_caption_above_mediaBOOLEANfalsePass True, if the caption must be shown above the message media
disable_notificationBOOLEANtrueSends the message silently. Users will receive a notification with no sound.
protect_contentBOOLEANfalseProtects the contents of the sent message from forwarding and saving
send_asCOMBOHow to send the audio
file_nameSTRINGaudiothe file name for this media
message_thread_idINT-1-1–9223372036854776000Unique identifier for the target message thread of the forum topic (-1 = None)
triggeropt*Optional trigger to enforce execution order

Outputs (3)

NameTypeDescription
messageDICT—
message_idINT—
trigger*—