Audio Beats
Stop guessing where the drop is
- audio
- beats
- audio
- tempo
- bar_seconds
- beat_count
- plot
- report
Every "I cut my video to the music" workflow has the same weak step: you watch the clip back, count beats in your head, and drag a number into a field. Audio Beats does that part for you. Feed it a song and it hands back the beat times, the tempo, the first beat of each bar, and how loud the track is at every single frame of your video. Then you can cut on a downbeat, hold a shot for exactly one bar, or drive a motion amount from the loudness curve instead of vibes.
It is one of the newer, video-shaped corners of WAS Node Suite - a pack most people know for image ops like film grain and blend modes, and which now runs to around 495 nodes with a whole animation category bolted on.
How it works
Under the hood it is a real analysis, not an envelope follower. Onsets are found as log-mel spectral flux, beats are tracked across the whole track by dynamic programming against a tempo prior, and the downbeats - the "one" of each bar - are picked out by bass energy below roughly 200 Hz. Tempo search runs from 50 to 220 BPM, so a 90 BPM hip-hop track and a 174 BPM drum-and-bass track both resolve.
The clever bit: the analysis is re-sampled onto your video's frame rate. That is why the node asks for fps at all. Loudness and onset strength come back as one value per video frame, so a curve output lines up with frame N of your clip exactly, with no offset maths.
The inputs that matter
audio is the song, straight from Load Audio. fps should be the frame rate you will render at - 24 for MiniMax H3, 16 for Wan, 30 for LTX, and so on. Get this wrong and every per-frame number is subtly off.
beats_per_bar defaults to 4, which is right for most music and wrong for a waltz (3). start and length let you analyse a window instead of the whole track - start 31.5, length 48.0 if your video starts half a minute in and runs 48 seconds. length 0 means "the rest of the song".
What comes out
tempo is BPM as a float, and bar_seconds is how long one bar lasts - that second one is the value you actually want, because it turns "maybe four seconds" into a segment length that lines up. beat_count is just the number found.
plot is an IMAGE of the strip you see on the node: loudness, beats and bar lines over time. Wire it to Preview Image when you want to check the tracker caught the right pulse before you trust it. report is a STRING with the tempo, the counts and the first beat and bar times.
audio is a passthrough of the stretch that was analysed - wire it to your video output node so picture and sound cover the same span. And beats is a WAS_BEATS bundle holding times and per-frame curves together. Be aware that nothing else in the pack consumes that type yet, so it is future-proofing rather than a wire you can use today. Use the individual outputs.
Install
In ComfyUI Manager, search for WAS Node Suite v3 and install, then restart. Manually:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
You need ComfyUI 0.14.0 or newer and Python 3.10+. That is the whole install - the pack does not touch pip, and its requirements.txt is literally a one-line comment saying so. The first start is a second or two slower while it writes config.yaml and compiles to bytecode. Audio Beats needs no model files.
Where it gets fiddly
Sparse or beatless material. Ambient, drones, rubato classical and a capella vocal will produce a tempo, because the tracker has to commit to something. Check plot first; if the bar lines drift through the track, do not cut to them.
Half/double confusion. Trackers love to land on 60 or 240 when you hear 120. If bar_seconds feels twice or half as long as the music does, that is what happened - adjust beats_per_bar or just do the arithmetic on your segment lengths.
Mismatched fps. Analyse at 24, render at 16, and the loudness curve is a third too fast. The per-frame outputs are meaningless unless fps matches the clip.
Audio with silence up front. The first beat is measured from second zero of the audio, not of your video. Set start to the audio offset and both line up.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| audio | AUDIO | The song, from Load Audio. | |
| fps | FLOAT | 24.0001–240 | Video frames per second, as `24` for MiniMax H3 or `16` for Wan. |
| beats_per_bar | INT | 41–16 | Beats in one bar, as `4` for most music or `3` for a waltz. |
| start | FLOAT | 0.000–86400 | Seconds into the song the video starts at, as `0.0` or `31.5`. |
| length | FLOAT | 0.000–86400 | Seconds of song the video covers, as `48.0`, or `0.0` for the rest. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| beats | WAS_BEATS | The beats, times and per-frame loudness together. |
| audio | AUDIO | The stretch analysed, from start for length seconds. |
| tempo | FLOAT | Beats per minute. |
| bar_seconds | FLOAT | Seconds in one bar, for lining segment lengths up with bars. |
| beat_count | INT | Beats found. |
| plot | IMAGE | The strip drawn on the node. |
| report | STRING | Tempo, counts and the first beat and bar times. |