ComfyUI Node

VRGDG Ensure Video Audio

What VRGDG Ensure Video Audio Is For

By vrgamegirl19·Created about a year ago·Updated a day ago· 742
VRGDG Ensure Video Audio
  • audio
  • video_info
  • audio
  • has_source_audio
  • status
◄fallback_sample_rate32000►
◄fallback_channels2►

The boring problem it solves

Half the video files on your drive have no audio stream. Screen captures, exports from a silent render pipeline, anything that got muxed without a track. That's fine until you plug it into a graph where something downstream treats AUDIO as a required input - an audio-driven video model, a lip-sync pass, a transcription node, a muxer. Then the whole queue dies on a file that is otherwise perfectly good.

VRGDG Ensure Video Audio is the adapter for that situation. Give it an audio stream and the video's info; if the audio is readable it hands it straight back, and if it isn't, it manufactures silence the exact length of the video so the rest of your graph can keep going.

How it works

The mechanism is about as simple as it gets, and reading the source is worth thirty seconds because the failure is subtle:

waveform = audio["waveform"]                       # [batch, channels, samples]
sample_rate = int(audio["sample_rate"])

If that succeeds, the waveform is a 3D tensor, has at least one sample, and the sample rate is positive, the node returns it unchanged - byte-for-byte the same audio, flagged True. No resampling, no remuxing, no normalisation.

If any of that raises, it falls back to a duration lookup on video_info: source_duration first, then loaded_duration, then source_frame_count / source_fps. If none of those give it a number, it raises The video loader did not report a usable video duration. Otherwise it builds a zero-filled float32 tensor on CPU - [1, channels, round(duration × sample_rate)] - and returns that with the flag set to False.

One detail that matters: VideoHelperSuite defers audio extraction until somebody actually touches the AUDIO mapping. That's why the read is wrapped in a try/except rather than checked up front, and why a file that "has audio" can still land on the silent path if the stream is unreadable to ffmpeg.

The inputs and outputs that matter

You set two things, and only if you care. fallback_sample_rate (32000 / 44100 / 48000, default 32000) and fallback_channels (1–8, default 2) only describe the invented silence; internally the node clamps to 8000–192000 Hz and 1–8 channels. Match them to what your downstream model expects - 48k stereo is the safe pick if you're feeding a video model, 32k is fine for speech tooling.

audio takes the AUDIO stream, video_info takes a VHS_VIDEOINFO - that's the dict modern VHS loaders output alongside frames, and it's where the duration comes from.

On the output side, audio is the stream you wire onward. has_source_audio is the one people forget exists and then wish they'd used: wire it into a switch or a boolean-condition node to skip the audio-dependent branch entirely when there was nothing to hear, instead of burning GPU time on generated silence. status is a one-line human string - on the fallback path it includes VHS's own exception text (truncated) so you can read what ffmpeg actually complained about instead of guessing.

Install

ComfyUI Manager → Install Custom Nodes → search vrgamedev, or:

cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl.git
python -m pip install -r comfyui-vrgamedevgirl/requirements.txt

Restart ComfyUI and hard-refresh the browser - this pack ships JavaScript UI panels, and a soft refresh will happily keep serving the old ones.

The real prerequisite here is not this pack: it's ComfyUI-VideoHelperSuite, because VHS_VIDEOINFO comes from VHS's loaders. No VHS, no video_info socket to fill.

When it goes wrong

  • "The video loader did not report a usable video duration." You fed it a video_info that has neither durations nor frame count plus fps. Old or third-party loaders are the usual suspects; use a current VHS load video node.
  • It reports "no readable source audio" on a file that clearly has audio. The stream exists but ffmpeg can't decode it. Run the file through a remux first - ffmpeg -i in.mkv -c:v copy -c:a aac out.mp4 fixes most of these.
  • Transcription returns nothing after the fallback fired. Correct behaviour, not a bug. Silence transcribes to silence.
  • The node re-executes on every queue even when nothing changed. That's intentional: IS_CHANGED returns NaN, so it never caches. Cheap here, worth knowing before you blame your cache settings.

This node is brand new - it landed in the pack mid-September 2026 - so there's no crowd of people who've already hit its edges for you. The source is short enough to read in one sitting, which is more than you can say for most of what ships in video packs.

CategoryVRGDG/Video/Audio

Inputs (4)

NameTypeDefaultDescription
audioAUDIO—
video_infoVHS_VIDEOINFO—
fallback_sample_rateCOMBO320003 options: 32000, 44100, 48000
fallback_channelsINT21–8—

Outputs (3)

NameTypeDescription
audioAUDIO—
has_source_audioBOOLEAN—
statusSTRING—