Nodes/ComfyUI-SongScribe/Song Analyzer (SongScribe)
ComfyUI Node

Song Analyzer (SongScribe)

Point it at a track and get a prompt that doesn't lie

By TheLocalLab·Created 2 months ago·Updated 20 days ago· 23
Song Analyzer (SongScribe)
  • audio
  • caption
  • lyrics
  • duration
  • duration_int
  • duration_str
  • analysis
◄audio_file▾►
◄describeclap►
◄genre_sourceclap►
◄clap_modelmusic_and_speech►
◄transcribe_lyricsoff►
◄whisper_modelbase►
◄stylebalanced►
◄use_cachetrue►
◄seed0►

Start with what it isn't

Song Analyzer generates no audio. It is the front half of the job - you have a reference track (or you just like how one sounds) and you need a prompt that describes it well enough for a music model to give you something new in the same spirit. That's the whole node.

This matters more than it sounds, because the local music stack has a specific weak spot. As the audio layer of ComfyUI goes, the honest summary is that models like ACE-Step are genuinely good at instrumentals and genuinely bad at vocals and lyrics - the community line people repeat is "the lyrics are absolute garbage," which is also the complaint aimed at Suno. A caption-and-lyrics prep node doesn't fix the model, but it removes the excuse: you hand it structure and words instead of hoping.

Where people reach for it: you have a track you want to riff on, you want its BPM, key and arrangement kept, and its vocal line replaced. MiniMax Music 3's caption format is three labelled sections - Global Metadata, Vocal Details, Arrangement - and this node writes all three from measurements.

How it works, and why it's honest

Three stages, and the split between them is the point. Measured facts come from DSP: librosa does beat tracking for BPM, chroma for key, loudness and crest factor, spectral descriptors, and section boundaries via agglomerative clustering on MFCCs and chroma. Those numbers aren't guessed, so the composer states them exactly.

Descriptors come from CLAP - but not how you'd expect. The model never writes prose and never emits a number; it scores a fixed, hand-authored vocabulary (songscribe/vocab/*.yaml) against 10-second windows, so the caption can only contain phrases that were already in that vocabulary. The failure mode is "picked a less apt word," never "invented a fact." Genre can instead come from MAEST, a supervised tagger trained on 400 Discogs styles, and the author's testing is blunt about the gap: a trap track CLAP called "bossa nova" came back Trap / Cloud Rap / Hardcore Hip-Hop from MAEST.

Two axes are retired, and the reason is worth knowing: mood and vocal timbre returned nearly the same answer for every song tested - "confident and swaggering" on five of six tracks, including a funeral ballad. They're off in the vocab files (enabled: false) rather than left in to produce pretty nonsense. Vocal axes are also skipped outright when vocal-presence detection says the track is instrumental.

The inputs you'll actually touch

audio_file is a dropdown of everything in ComfyUI's input folder that looks like audio - wav, mp3, flac, m4a, ogg, opus, aiff, wma and more - with an upload button, plus a (use AUDIO input) sentinel for when you'd rather drive it from a socket. The optional audio input takes priority over the dropdown when connected.

describe = clap for the full descriptor pass, off for measured facts only. Set genre_source to maest if genre matters and you'll accept trust_remote_code - pinned to one audited commit and only loaded when selected, but it does execute Python from a model repo, so the default stays on CLAP. Use off when a Style Preset already knows the genre.

style is the one that changes your output most. verbatim keeps exact section timings and is the closest thing to cloning the source; balanced keeps tempo and key and drops second-level timings; loose gives vibe only - no tempo, no key, no structure. use_cache stays on (re-runs are roughly 100× faster). seed re-rolls the connective phrasing without re-analysing anything, and transcribe_lyrics stays off unless the file has no embedded lyrics - this is sung ASR, and it will need hand-fixing.

Outputs

caption, lyrics and duration wire straight into MiniMax's caption, lyrics and max_duration. duration_int and duration_str are there for filenames and notes. analysis carries the full payload as a SONGSCRIBE_ANALYSIS type - be aware nothing in the pack consumes it yet, so it's a Show Any / preview node for now.

Lyrics come from embedded tags first, then a sibling .lrc or .txt sitting next to the audio file, and Whisper only as a last resort. That order is deliberate: a tag is an exact transcription someone already made; Whisper on a full mix is an estimate. Note that transcribe_lyrics: always skips both and forces the estimate.

Install

ComfyUI Manager → search SongScribe. Or manually:

cd ComfyUI/custom_nodes
git clone https://github.com/TheLocalLab/ComfyUI-SongScribe
python_embeded/python.exe -m pip install librosa mutagen pyyaml
# optional, only for transcription:
python_embeded/python.exe -m pip install faster-whisper

Restart ComfyUI; the nodes land under the SongScribe category. Analysis runs on CPU, so no GPU requirement at all.

Where it bites

First run feels broken. It isn't - the CLAP checkpoint downloads on first use (the node tooltip says ~600 MB, the checkpoint dropdown says 1.5 GB for music_and_speech and 1.9 GB for general), and MAEST adds ~330 MB the first time you select it. Budget a few GB and a quiet queue.

Your file has to be in ComfyUI's input directory for the dropdown to see it. Use the upload button or the audio socket if it isn't.

The cache writes a .songscribe.json sidecar next to the audio file (or ComfyUI's temp dir if that isn't writable) - if a re-run won't reflect a file you just replaced, look there. Cache keys include describe and clap_model but deliberately not style, so flipping between balanced and loose recomposes instantly from stored measurements.

Speed, for planning: about 5 seconds for a 40-second track and 14 for a five-minute one on CPU, near-instant when cached.

Finally, trust the pack's own honesty section: BPM, key, duration, structure, vocal presence and MAEST genre are solid, CLAP genre is roughly a coin flip. If a reading looks wrong, it probably is.

CategorySongScribe

Inputs (10)

NameTypeDefaultDescription
audio_fileCOMBO1 options: (use AUDIO input)
describeCOMBOclapScore genre, mood, instruments and vocal character with CLAP (CPU, no GPU needed). The first run downloads a ~600 MB model. 'off' emits measured facts only.
genre_sourceCOMBOclapWhere genre comes from. 'maest' is a supervised tagger trained on 400 Discogs styles and is markedly more accurate than CLAP's zero-shot guessing, but it executes custom model code (trust_remote_code) from a pinned, audited revision - opt in knowingly. 'off' omits genre so a preset can supply it.
clap_modelCOMBOmusic_and_speechWhich CLAP checkpoint scores the descriptors. Downloaded on first use: 'music_and_speech' ~1.5 GB (music-specialised, better on genre), 'general' ~1.9 GB (broader audio, weaker on genre). Ignored when describe is off.
transcribe_lyricsCOMBOoffTranscribe sung lyrics with Whisper when the file carries none. CPU, slow, and sung ASR is markedly worse than speech - expect to hand-fix the result. Embedded tags and .lrc files are always preferred over this.
whisper_modelCOMBObaseLarger is more accurate and much slower. Only used when transcribe_lyrics is enabled.
styleCOMBObalancedHow literally the caption reproduces the track. 'verbatim' keeps exact section timings (closest clone); 'balanced' keeps tempo and key; 'loose' gives genre, mood and texture only.
use_cacheBOOLEANtrueReuse a previous analysis of the same file instead of recomputing it on every queue.
seedINT00–18446744073709550000Varies caption phrasing without re-analysing the audio.
audiooptAUDIOAnalyse audio from an upstream node. Takes priority over the file dropdown when connected.

Outputs (6)

NameTypeDescription
captionSTRINGThree-section caption (Global Metadata / Vocal Details / Arrangement).
lyricsSTRINGLyrics from embedded tags or a sidecar .lrc/.txt, if present.
durationFLOATDuration in seconds - wire straight into max_duration.
duration_intINTDuration rounded to whole seconds.
duration_strSTRINGDuration formatted as m:ss.
analysisSONGSCRIBE_ANALYSISFull analysis payload for downstream SongScribe nodes.