Style Preset (SongScribe)
73 genres, zero audio required
- prompt
- global_metadata
- vocal_details
- arrangement
The node most people actually want
Song Analyzer needs a reference track. Style Preset needs nothing - you pick a genre from a dropdown and it hands you a prompt that's already written in the format your music model reads. If you're sitting in front of ACE-Step or YuE2 with an empty style box and no idea what a good prompt for "roots reggae" even looks like, this is the shortcut.
The dropdown is 73 presets across 14 categories - acoustic, cinematic, country, electronic, gospel, hip-hop, jazz, metal, pop, rock, soul, swing, vocal and world, running from Hip-Hop / Trap and Jazz / Bossa Nova to World / Reggae and Vocal / Barbershop Quartet. Each one is a hand-authored YAML file that knows its own genre tags, mood, instrumentation, production, vocal presence and delivery, plus BPM, key and feel. That's the real product here - someone did the prompt engineering for 73 styles so you don't have to guess that a reggae prompt wants "one-drop drums with the kick on the three" and "Hammond organ bubble."
There's a second reason it matters, straight from the author's own limits section: audio-inferred mood and vocal timbre were switched off because they weren't reliable. A preset asserting "laid-back and dreamy" is a human who knows what reggae sounds like, and it beats a model inferring it badly. Presets and analysis are complementary - measured facts from the Analyzer, taste from the preset.
Set format first
Everything else on this node behaves differently depending on one choice, so make it before you touch anything else.
minimax gives you MiniMax Music 3's three-section caption: Global Metadata, Vocal Details, Arrangement, each labelled. It's what the style widget controls - verbatim, balanced, loose - meaning how much detail each section carries.
yue2 gives one flat comma-separated descriptor string, which is what YuE2's style input expects. That's a genuinely different shape: no headers, vocal type stated outright, language named. YuE2's own prompts vary in length, so detail dials yours: tags is genre, tempo, mood and vocal only; full adds instrumentation and production (start here); rich adds scene and texture. vocal picks the voice type explicitly - nine choices from female vocals through choir and rap vocals to instrumental (no vocals) - with auto falling back to whatever the preset declares, so a preset marked instrumental stays instrumental. language matters if your lyrics aren't English; YuE2 supports English, Mandarin, Japanese, Korean, Spanish and Russian, and naming the language steers pronunciation.
Settings for the format you're not using are inert, not broken. Choosing detail: rich in minimax mode does nothing. That's fine, just don't debug it.
The flavour dials and blending
era, texture and mood_shift layer on top of the preset - 1960s through modern, lo-fi / tape / hi-fi / intimate / cavernous, and darker, brighter, sadder, dreamier, harder. They add to what the preset already says; they never replace it. mood_shift: harder on a pop ballad gives you a harder pop ballad, not a metal track.
blend_with plus blend mixes a second preset in at a weight from 0 (first only) to 1 (second only), which is how you get the two things a genre list can't give you: dub-techno, or honky-tonk dream pop. blend: 0.5 on a reggae preset and an ambient one is the sort of thing worth burning a few minutes on.
extra is appended verbatim and is the most underrated field on the node. A reference artist, a negative instruction like "avoid heavy distortion", a specific structure - anything the preset can't know. seed only re-rolls the phrasing, same as on the Analyzer.
Outputs and the YuE2 wiring trap
prompt goes to MiniMax's caption input or YuE2's style input. The three section outputs - global_metadata, vocal_details, arrangement - stay populated in both formats, which is what makes Style Preset a useful partner for Caption Composer: take a preset's arrangement, take a measured caption's vocal line, rebuild. In YuE2 format the pack still computes the three-section form because it costs nothing and keeps the caption nodes usable either way.
For YuE2 there's a step the README flags and it's the number one "why can't I connect this" question in any ComfyUI pack: YuE2's style is a widget, not a socket. Right-click the YuE2 node, choose Convert style to input, and now prompt has somewhere to plug in. That's the standard ComfyUI mechanic for turning a typed field into a wire, and it's not specific to this pack.
Install and adding your own presets
ComfyUI Manager → search SongScribe, or:
cd ComfyUI/custom_nodes
git clone https://github.com/TheLocalLab/ComfyUI-SongScribe
python_embeded/python.exe -m pip install librosa mutagen pyyaml
Only PyYAML matters for this node - no audio, no models, no downloads. Restart ComfyUI and the nodes are under SongScribe.
To add your own genre, drop a .yaml into songscribe/presets/ and restart. Copy world-reggae.yaml as a template: it's a dozen plain keys (genre, mood, scene, production, instruments, vocal presence/timbre/delivery, BPM, key, harmony, feel, drums, low end) and the fields read exactly like the finished prompt. If a preset sounds wrong to you, say so - the README notes preset corrections from users are what fixed several of them.
Inputs (13)
| Name | Type | Default | Description |
|---|---|---|---|
| preset | COMBO | Acoustic / Bluegrass | 73 options: Acoustic / Bluegrass, Acoustic / Country, Acoustic / Indie Folk, Acoustic / Solo Piano, Cinematic / Classical Chamber, Cinematic / Epic Trailer, +67 |
| format | COMBO | minimax | minimax: three labelled sections (Global Metadata / Vocal Details / Arrangement). yue2: one flat comma-separated descriptor string, which is what YuE2's style input expects. |
| detail | COMBO | full | YuE2 only. tags: genre, tempo, mood and vocal. full: adds instrumentation and production. rich: adds scene and extra texture. YuE2's own prompts range across all three. |
| vocal | COMBO | auto | YuE2 states vocal type explicitly in most prompts. 'auto' uses whatever the preset declares. |
| language | COMBO | auto | Sung language. YuE2 supports English, Mandarin, Japanese, Korean, Spanish and Russian; naming it steers pronunciation. |
| style | COMBO | balanced | MiniMax only: how much measured detail the caption carries. |
| era | COMBO | none | Layer a production era on top. |
| texture | COMBO | none | Layer a recording texture on top. |
| mood_shift | COMBO | none | Push the mood in a direction. |
| seed | INT | 00–18446744073709550000 | — |
| blend_withopt | COMBO | none | Optional second preset to mix in. |
| blendopt | FLOAT | 0.500–1 | How much of the second preset to admit. 0 = first only, 1 = second only. |
| extraopt | STRING | Appended verbatim. Use for anything the preset cannot know - a reference artist, a negative instruction like 'avoid heavy distortion', or a specific structure. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| prompt | STRING | The style prompt, in the selected format. Wire to YuE2's style input or MiniMax's caption input. |
| global_metadata | STRING | MiniMax section 1 (empty detail in yue2 format is still filled). |
| vocal_details | STRING | MiniMax section 2. |
| arrangement | STRING | MiniMax section 3. |