Nodes/ComfyUI-SongScribe/Caption Composer (SongScribe)
ComfyUI Node

Caption Composer (SongScribe)

Rebuild a music caption without retyping the whole thing

By TheLocalLab·Created 2 months ago·Updated 20 days ago· 23
Caption Composer (SongScribe)
    • caption
    ◄global_metadata►
    ◄vocal_details►
    ◄arrangement►
    ◄headerstrue►

    The boring half of a two-node idea

    Caption Composer takes three strings and glues them into one caption. That's it. It's the reassembly step that pairs with Caption Splitter, and on its own it looks like the least interesting node in the pack.

    It isn't useless, though - it's the node that lets a prompt be a graph instead of a wall of text you edit by hand. MiniMax Music 3's caption format is three labelled sections (Global Metadata, Vocal Details, Arrangement), and the moment you want one of those to come from somewhere other than the text box you last pasted into, you need a node that puts the pieces back together.

    What it does mechanically

    Three inputs - global_metadata, vocal_details, arrangement - all multiline strings, all defaulting to empty, and all connectable. One optional boolean, headers, on by default: it emits the literal Global Metadata: / Vocal Details: / Arrangement: labels. MiniMax expects those labels, so leave it on unless you're targeting something else.

    Two behaviours from the source are worth knowing because they save you from yourself. Empty sections are dropped entirely rather than emitted as a bare Vocal Details: header with nothing after it - so an accidental blank input degrades to a two-section caption instead of a malformed one. And when headers is on, any header text you left inside a body gets stripped first, so pasting a whole caption into arrangement can't produce Arrangement: Arrangement: .... Bodies are joined with blank lines between sections.

    Output is a single STRING named caption. Wire it to MiniMax's caption input. That's the whole contract.

    Where it earns its place

    The obvious use is the round trip: Song Analyzer's caption → Caption Splitter → edit one section → Caption Composer. That's the workflow the pack's README names, and the reason it exists is specific - keeping a measured arrangement while replacing the vocal description entirely. Vocal timbre scoring is switched off in this pack precisely because CLAP was wrong on a binary male/female question two times out of five, so the caption's voice description is the part you're most likely to want to write yourself anyway. Here's the part that surprises people: Style Preset also outputs global_metadata, vocal_details and arrangement separately - in both formats, always, even when you've asked for YuE2's flat string. So Composer is also the way you take a preset's arrangement and a measured caption's vocal description and merge them.

    The third use is the plain one: three text boxes, a caption you typed by hand, and a labelled section format you don't have to remember. That's legitimate. If you keep one Composer node in a template with your favourite vocal description plugged into section two, you've built a reusable "house style" block for MiniMax without writing any code.

    Drop it into your graph

    Install is the pack's shared install. ComfyUI Manager → search SongScribe, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/TheLocalLab/ComfyUI-SongScribe
    python_embeded/python.exe -m pip install librosa mutagen pyyaml
    

    Then restart. Caption Composer has no dependencies of its own - it's pure string handling, no model, no audio, no download. Even the librosa install is only needed because the same pack contains the Song Analyzer; if you only use the caption nodes and the Style Preset, nothing heavy runs.

    Traps

    Composer is a dumb joiner, and it will happily build a caption that reads like nonsense if you feed it mismatched sections - say, a Global Metadata line claiming 76 BPM reggae with an Arrangement section describing a distorted nine-minute prog epic. Music models don't reconcile contradictions; they average them. If you're mixing sources, keep genre, tempo and instrumentation consistent across all three parts.

    Second: headers off is a trap for MiniMax users. The three labels are part of the format the model was trained to read. Turning them off because a caption "looks cleaner" without them just throws away the structure you spent the last two nodes building.

    Third, and it's the one that catches people coming from the Splitter: text with no recognisable headers lands wholly in global_metadata when split. Feed that straight back through the Composer and everything comes out under section one - the round trip "works" and quietly reshuffles your caption into the wrong section. Eyeball the Splitter's outputs before you trust the loop.

    CategorySongScribe

    Inputs (4)

    NameTypeDefaultDescription
    global_metadataSTRING—
    vocal_detailsSTRING—
    arrangementSTRING—
    headersoptBOOLEANtrueEmit the 'Global Metadata:' / 'Vocal Details:' / 'Arrangement:' labels. MiniMax expects them; turn off only for other models.

    Outputs (1)

    NameTypeDescription
    captionSTRING—