Caption Splitter (SongScribe)
Change one line of a music caption, keep the rest
- global_metadata
- vocal_details
- arrangement
A one-input, three-output text knife
Caption Splitter takes one caption string and cuts it into its three MiniMax sections: global_metadata, vocal_details, arrangement. One input, no options, no model, no dependency beyond the pack's own Python. It's the smallest node in SongScribe and it solves the most annoying thing about prompt work: you liked the caption, you want to change exactly one clause of it, and retyping the other 200 words is how you introduce a typo you won't notice.
The description that sells it comes straight from the pack README: split a caption, edit a section, rebuild it. Keeping an analysed arrangement while replacing the vocal description entirely is the demo, and it's a real one. The Song Analyzer's voice description is its weakest output by design - mood and vocal timbre scoring are switched off in this pack because CLAP scored 2/5 on even a binary male/female question - so the voice line is precisely the bit you'll want to overrule.
What the splitting actually does
It's a regex, not a parser, and it's a slightly more careful regex than you'd write on the first attempt. It looks for the three canonical header names at the start of a line, case-insensitively, and it tolerates the markdown bold that shows up when captions get pasted out of a Discord message or a GitHub README - **Arrangement:** and **Arrangement**: both match. Anything that appears before the first header is preserved and prepended to global_metadata rather than silently dropped. Repeated headers append to the same output instead of overwriting.
The one fallback worth knowing: if there are no recognisable headers at all, the entire caption is returned in global_metadata and the other two outputs are empty. Nothing is thrown away, which is the right call - but it's a fallback, not a split, and it's the reason this node and a YuE2-format prompt don't mix.
Wiring it
The outputs are three plain strings, and what they wire into depends on what you're doing.
- Into Caption Composer: the classic round trip. Analyzer
caption→ Splitter → an edit → Composer → MiniMax'scaptioninput. - Into a Show Text / preview node: genuinely worth doing once per project, because it's how you discover that your caption's Arrangement section is carrying five instrument names you didn't want.
- Into Lyrics Structure and then a music model: the sections you care about can be inspected in the same graph as the render, so a bad caption costs you a queue instead of a five-minute song render you throw away.
For the edit step, either type into Composer's three multiline boxes directly, or use core string/primitive nodes so the replacement text comes from somewhere reusable. A text box holding your standard vocal line - "female lead, breathy, double-tracked harmonies in the chorus" - connected into Composer's vocal_details and left in the template saves you the edit on every run.
Install
Same pack, same install. ComfyUI Manager → search SongScribe, or:
cd ComfyUI/custom_nodes
git clone https://github.com/TheLocalLab/ComfyUI-SongScribe
python_embeded/python.exe -m pip install librosa mutagen pyyaml
Restart ComfyUI. The Splitter itself needs nothing downloaded and runs instantly.
Traps
The flat YuE2 format has no sections to find. Run a format: yue2 Style Preset prompt through the Splitter and everything lands in global_metadata with the other two outputs empty. That's not a bug you can configure around - the format genuinely has no headers. If you need sections, use the Style Preset's three dedicated outputs (global_metadata, vocal_details, arrangement), which stay populated even when the prompt string is in YuE2 format.
Second trap, and it's the one to watch: a caption whose headers were hand-deleted. The moment someone rewrites a caption into flowing prose, the Splitter degrades to a pass-through and all three sections collapse into one output. Composer will happily rebuild that as a single-section caption with a Global Metadata: label stuck on the front, which is wrong in a way that doesn't error.
Third: pasted captions arrive with whatever formatting they had. The header matcher handles bold and stray whitespace, but if a caption uses different labels - Style:, Voice:, Structure: - you get the no-headers fallback. Match the pack's three labels exactly and it behaves.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| caption | STRING | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| global_metadata | STRING | Genre, tempo, key, mood, production. |
| vocal_details | STRING | Voice description, or the instrumental note. |
| arrangement | STRING | Instrumentation, groove and section map. |