Nodes/ComfyUI-SongScribe/Caption Splitter (SongScribe)
ComfyUI Node

Caption Splitter (SongScribe)

Change one line of a music caption, keep the rest

By TheLocalLab·Created 2 months ago·Updated 20 days ago· 23
Caption Splitter (SongScribe)
    • global_metadata
    • vocal_details
    • arrangement
    ◄caption►

    A one-input, three-output text knife

    Caption Splitter takes one caption string and cuts it into its three MiniMax sections: global_metadata, vocal_details, arrangement. One input, no options, no model, no dependency beyond the pack's own Python. It's the smallest node in SongScribe and it solves the most annoying thing about prompt work: you liked the caption, you want to change exactly one clause of it, and retyping the other 200 words is how you introduce a typo you won't notice.

    The description that sells it comes straight from the pack README: split a caption, edit a section, rebuild it. Keeping an analysed arrangement while replacing the vocal description entirely is the demo, and it's a real one. The Song Analyzer's voice description is its weakest output by design - mood and vocal timbre scoring are switched off in this pack because CLAP scored 2/5 on even a binary male/female question - so the voice line is precisely the bit you'll want to overrule.

    What the splitting actually does

    It's a regex, not a parser, and it's a slightly more careful regex than you'd write on the first attempt. It looks for the three canonical header names at the start of a line, case-insensitively, and it tolerates the markdown bold that shows up when captions get pasted out of a Discord message or a GitHub README - **Arrangement:** and **Arrangement**: both match. Anything that appears before the first header is preserved and prepended to global_metadata rather than silently dropped. Repeated headers append to the same output instead of overwriting.

    The one fallback worth knowing: if there are no recognisable headers at all, the entire caption is returned in global_metadata and the other two outputs are empty. Nothing is thrown away, which is the right call - but it's a fallback, not a split, and it's the reason this node and a YuE2-format prompt don't mix.

    Wiring it

    The outputs are three plain strings, and what they wire into depends on what you're doing.

    • Into Caption Composer: the classic round trip. Analyzer caption → Splitter → an edit → Composer → MiniMax's caption input.
    • Into a Show Text / preview node: genuinely worth doing once per project, because it's how you discover that your caption's Arrangement section is carrying five instrument names you didn't want.
    • Into Lyrics Structure and then a music model: the sections you care about can be inspected in the same graph as the render, so a bad caption costs you a queue instead of a five-minute song render you throw away.

    For the edit step, either type into Composer's three multiline boxes directly, or use core string/primitive nodes so the replacement text comes from somewhere reusable. A text box holding your standard vocal line - "female lead, breathy, double-tracked harmonies in the chorus" - connected into Composer's vocal_details and left in the template saves you the edit on every run.

    Install

    Same pack, same install. ComfyUI Manager → search SongScribe, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/TheLocalLab/ComfyUI-SongScribe
    python_embeded/python.exe -m pip install librosa mutagen pyyaml
    

    Restart ComfyUI. The Splitter itself needs nothing downloaded and runs instantly.

    Traps

    The flat YuE2 format has no sections to find. Run a format: yue2 Style Preset prompt through the Splitter and everything lands in global_metadata with the other two outputs empty. That's not a bug you can configure around - the format genuinely has no headers. If you need sections, use the Style Preset's three dedicated outputs (global_metadata, vocal_details, arrangement), which stay populated even when the prompt string is in YuE2 format.

    Second trap, and it's the one to watch: a caption whose headers were hand-deleted. The moment someone rewrites a caption into flowing prose, the Splitter degrades to a pass-through and all three sections collapse into one output. Composer will happily rebuild that as a single-section caption with a Global Metadata: label stuck on the front, which is wrong in a way that doesn't error.

    Third: pasted captions arrive with whatever formatting they had. The header matcher handles bold and stray whitespace, but if a caption uses different labels - Style:, Voice:, Structure: - you get the no-headers fallback. Match the pack's three labels exactly and it behaves.

    CategorySongScribe

    Inputs (1)

    NameTypeDefaultDescription
    captionSTRING—

    Outputs (3)

    NameTypeDescription
    global_metadataSTRINGGenre, tempo, key, mood, production.
    vocal_detailsSTRINGVoice description, or the instrumental note.
    arrangementSTRINGInstrumentation, groove and section map.