🎵Music3音乐提示词生成@炮老师的小课堂
Writing MiniMax Music 3 prompts with a full band's worth of knobs
- 🎼 音乐描述
- 📝 歌词
- 🌐 全局音乐信息
- 🎙️ 人声细节
- 🧱 编曲结构
- 📑 音乐结构分析
- 📄 语言模型完整响应
- ℹ️ 处理信息
Music generation models are like image models circa 2023: the prompt matters enormously and most people write it badly. DapaoMusic3CaptionPromptNode is the dapaoAI pack's answer - it calls an LLM to turn your one-line music idea into a structured MiniMax Music 3 description plus a full set of lyrics, with a frankly absurd number of musical knobs so the output actually matches what you meant.
The knobs (yes, there are a lot)
The required inputs read like a producer's checklist: 🎼 主风格 (20 options: pop, rock, electronic, folk, cinematic…), 🧬 融合风格, 🎯 使用场景, ⏳ 目标曲长, 💫 情绪弧 (the emotional arc - build, release, resolve), ⏱️ 速度 (BPM-ish), 🎹 调性倾向 (key), 🎚️ 拍号/律动 (time signature), 🥁 核心律动 (groove), 🎙️ 人声配置, 📈 人声音域, 🗣️ 人声音色, 🎤 演唱方式, 👥 和声/伴唱, 🎻 核心乐器编制 (17 options), 🧱 歌曲结构, 🎛️ 制作质感, 🌌 空间与混响, 🔥 编曲密度, and 🚫 排除项. Almost everything defaults to 自动识别 (auto), so you can just type a description and let it fill the blanks - or get surgical when you know exactly what you want.
The output-side controls: 🌐 输出语言 (default English, which the tooltip recommends for Music 3), 📏 输出详略 (default "standard 250–450 words"), 📦 输出格式 (structured text default). Plus the standard temperature (0.35), max tokens (4096), Top_P, cache-only seed, and timeout.
Every one of those enum fields has a matching ✍️ 自定义… string input, so when the preset list doesn't have your specific timbre you can type it.
The lyric handling - the part people ask about
📝 歌词 (and its 🔗 外部歌词 twin) has a clean contract: if you provide lyrics, the node preserves them exactly - it won't rewrite your words, it builds the musical description around them. If you don't, it composes original lyrics in the requested language. That's the right behavior for a tool like this: you keep creative control of the words, the LLM handles the sonic scaffolding.
Outputs
The first two are the money outputs - 🎼 音乐描述 and 📝 歌词 - and the tooltip says they can be wired directly into MiniMax's official Music3 Text Encode node, which is the whole point: this is a front-end for the official pipeline, not a standalone. The rest are structure and diagnostics: 🌐 全局音乐信息, 🎙️ 人声细节, 🧱 编曲结构, 📑 音乐结构分析, 📄 语言模型完整响应, ℹ️ 处理信息.
The honest take
Twenty-three musical parameter dropdowns is overkill for "I want a chill lo-fi beat," and the defaults exist precisely so you can ignore most of them. Where this node earns its keep is the specific brief - a client wants "a warm acoustic pop song, intimate female vocal, chorus that builds then resolves softly" and you need a Music 3 prompt that won't come back as elevator music. That's the LLM-as-translator pattern from llm-in-comfyui.md applied to audio, and it works the same way: the model is a translation layer, not a composer, and it can add flavor you didn't ask for. Scan the description before generating if exactness matters.
Gotchas
- Paid calls, no auto-retry. A truncated response means re-running and paying again - check 最大输出令牌 if outputs seem short for the "high detail" setting.
- Seed is cache-only, never sent to the API.
- Lyrics provided = preserved verbatim; leave 歌词 empty to get fresh writing. Both inputs together (📝 歌词 + 🔗 外部歌词) is redundant - use one.
- Chinese-only labels; defaults are sane.
Install: ComfyUI Manager → "dapaoAPI", or git clone https://github.com/paolaoshi/ComfyUI-dapaoAPI.git into ComfyUI/custom_nodes/, pip install -r requirements.txt, restart. dapaoAI key from api.dapaoai.com (register → wallet → redeem code → apply key in the default group).
Inputs (57)
| Name | Type | Default | Description |
|---|---|---|---|
| 🔑 API密钥 | STRING | 密钥仅用于 https://api.dapaoai.com,不写入配置文件。 | |
| 🤖 LLM模型 | COMBO | gemini-3.7-flash | 11 options: gpt-5.5, gpt-5.6-luna, gpt-5.6-terra, gpt-5.6-sol, claude-fable-5, claude-opus-4-8, +5 |
| 📝 原始音乐需求 | STRING | 温暖的原声流行歌曲,亲密女声,副歌逐步扩张并在结尾温柔收束。 | — |
| 🎼 主风格 | COMBO | 自动识别 | 20 options: 自动识别, 东亚现代流行|C-pop/J-pop与电子、R&B、说唱融合, 东亚抒情与国风|华语/日系抒情、原声或管弦, 现代R&B与Neo-Soul, 灵魂、蓝调与福音, 电影感流行抒情, +14 |
| 🧬 融合风格 | COMBO | 无 | 21 options: 无, 自动识别, 东亚现代流行|C-pop/J-pop与电子、R&B、说唱融合, 东亚抒情与国风|华语/日系抒情、原声或管弦, 现代R&B与Neo-Soul, 灵魂、蓝调与福音, +15 |
| 🎯 使用场景 | COMBO | 自动识别 | 11 options: 自动识别, 流媒体完整歌曲, 影视剧情/片尾曲, 电影或游戏配乐, 品牌广告/产品短片, 短视频背景音乐, +5 |
| ⏳ 目标曲长 | COMBO | 自动识别 | 7 options: 自动识别, 30–60秒, 1–2分钟, 2–3分钟, 3–4分钟, 4–5分钟, +1 |
| 💫 情绪弧 | COMBO | 自动识别 | 11 options: 自动识别, 克制铺陈 → 温暖释放 → 余韵收束, 脆弱低语 → 渐强 → 宣言式高潮, 平静神秘 → 紧张堆叠 → 史诗爆发, 忧伤回望 → 希望抬升 → 治愈落地, 明亮轻快 → 律动推进 → 庆祝式收尾, +5 |
| ⏱️ 速度 | COMBO | 自动识别 | 8 options: 自动识别, 慢速|60–78 BPM, 中慢|79–96 BPM, 中速|97–115 BPM, 中快|116–132 BPM, 快速|133–155 BPM, +2 |
| 🎹 调性倾向 | COMBO | 自动识别/不指定 | 10 options: 自动识别/不指定, 明亮大调倾向, 忧郁小调倾向, 大小调转换, 五声音阶/国风调式, 布鲁斯调式, +4 |
| 🎚️ 拍号/律动 | COMBO | 自动识别 | 7 options: 自动识别, 4/4, 3/4, 6/8, 12/8, 自由节拍/无固定律动, +1 |
| 🥁 核心律动 | COMBO | 自动识别 | 12 options: 自动识别, 平稳流行律动, 半拍/半速重心, Swing摇摆律动, Shuffle切分律动, 四拍踩底舞曲律动, +6 |
| 🎙️ 人声配置 | COMBO | 自动识别 | 13 options: 自动识别, 纯器乐|禁止人声与歌词, 单人女声, 单人男声, 中性/不限定性别单人声, 男女对唱, +7 |
| 📈 人声音域 | COMBO | 自动识别 | 10 options: 自动识别, 低沉低音区, 温暖中低音区, 自然中音区, 明亮中高音区, 高音区与假声, +4 |
| 🗣️ 人声音色 | COMBO | 自动识别 | 12 options: 自动识别, 清澈明亮, 温暖柔和, 气声亲密, 醇厚磁性, 沙哑颗粒感, +6 |
| 🎤 演唱方式 | COMBO | 自动识别 | 12 options: 自动识别, 叙述式、克制, 细腻气声、贴耳, 抒情渐强、宽阔副歌, 强力真声与高音爆发, R&B转音与即兴Ad-lib, +6 |
| 👥 和声/伴唱 | COMBO | 自动识别 | 10 options: 自动识别, 无伴唱、单主唱, 副歌轻量叠唱, 贴近三度/六度和声, 宽阔多轨和声墙, 领唱与群体呼应, +4 |
| 🎻 核心乐器编制 | COMBO | 自动识别 | 17 options: 自动识别, 钢琴 + 弦乐 + 克制鼓组, 指弹木吉他 + 轻鼓 + 温暖贝斯, Rhodes电钢琴 + 圆润贝斯 + R&B鼓组, 合成器Pad + 琶音 + 电子鼓, 失真电吉他 + 贝斯 + 现场鼓组, +11 |
| 🧱 歌曲结构 | COMBO | 自动识别 | 10 options: 自动识别, Intro → Verse → Pre-Chorus → Chorus → Verse → Chorus → Bridge → Final Chorus → Outro, Intro → Verse → Chorus → Verse → Chorus → Outro, Intro → Verse → Pre-Chorus → Chorus → Post-Chorus → Bridge → Final Chorus → Outro, Intro → Build → Drop → Breakdown → Final Drop → Outro, Intro → Rap Verse → Hook → Rap Verse → Hook → Bridge → Final Hook → Outro, +4 |
| 🎛️ 制作质感 | COMBO | 自动识别 | 12 options: 自动识别, 自然有机、动态呼吸, 现代流行、清晰宽阔, 温暖复古、磁带/黑胶质感, 电影化宽动态与空间层次, 俱乐部级紧实低频与冲击力, +6 |
| 🌌 空间与混响 | COMBO | 自动识别 | 9 options: 自动识别, 亲密近场、主唱居中, 适度宽声场、清晰层次, 大空间厅堂混响, 电影化深景与环绕感, 俱乐部直接、有力且紧实, +3 |
| 🔥 编曲密度 | COMBO | 自动识别 | 8 options: 自动识别, 极简留白, 由疏到密逐步堆叠, 中等密度、层次清晰, 副歌宽阔、主歌克制, 持续高能密集, +2 |
| 🚫 排除项 | COMBO | 无 | 11 options: 无, 禁止人声, 禁止说唱, 禁止电子鼓与808, 禁止失真吉他, 禁止合唱团, +5 |
| 🌐 输出语言 | COMBO | 英文(Music 3推荐) | 3 options: 英文(Music 3推荐), 中文, 双语:英文为主、中文注释 |
| 📏 输出详略 | COMBO | 标准|250–450词 | 4 options: 精简|180–250词, 标准|250–450词, 详细|450–650词, 自定义 |
| 📦 输出格式 | COMBO | 结构化文本 | 旧工作流兼容控件;当前始终输出可直连Music3的音乐描述与原始歌词。 |
| 🌡️ 温度 | FLOAT | 0.350–2 | — |
| 📝 最大输出令牌 | INT | 4096512–16384 | — |
| 🎲 Top_P | FLOAT | 1.000–1 | — |
| 🎲 随机种 | INT | 00–18446744073709550000 | 只控制ComfyUI缓存,不发送给LLM。 |
| ⌛ 请求超时 | INT | 30030–1200 | — |
| 🔗 外部音乐需求opt | STRING | 连接任意STRING节点;存在时与本节点的原始音乐需求合并。 | |
| 📝 歌词opt | STRING | — | |
| 🔗 外部歌词opt | STRING | 可连接歌词文本;存在时与本节点歌词合并并原样输出。没有任何歌词输入时由LLM自动生成。 | |
| ➕ 补充约束opt | STRING | — | |
| ✍️ 自定义主风格opt | STRING | — | |
| ✍️ 自定义融合风格opt | STRING | — | |
| ✍️ 自定义使用场景opt | STRING | — | |
| ✍️ 自定义目标曲长opt | STRING | — | |
| ✍️ 自定义情绪弧opt | STRING | — | |
| ✍️ 自定义速度opt | STRING | — | |
| ✍️ 自定义调性倾向opt | STRING | — | |
| ✍️ 自定义拍号/律动opt | STRING | — | |
| ✍️ 自定义核心律动opt | STRING | — | |
| ✍️ 自定义人声配置opt | STRING | — | |
| ✍️ 自定义人声音域opt | STRING | — | |
| ✍️ 自定义人声音色opt | STRING | — | |
| ✍️ 自定义演唱方式opt | STRING | — | |
| ✍️ 自定义和声/伴唱opt | STRING | — | |
| ✍️ 自定义乐器编制opt | STRING | — | |
| ✍️ 自定义歌曲结构opt | STRING | — | |
| ✍️ 自定义制作质感opt | STRING | — | |
| ✍️ 自定义空间与混响opt | STRING | — | |
| ✍️ 自定义编曲密度opt | STRING | — | |
| ✍️ 自定义排除项opt | STRING | — | |
| ✍️ 自定义输出详略opt | STRING | — | |
| 🚫 出错时跳过opt | BOOLEAN | false | — |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| 🎼 音乐描述 | STRING | — |
| 📝 歌词 | STRING | — |
| 🌐 全局音乐信息 | STRING | — |
| 🎙️ 人声细节 | STRING | — |
| 🧱 编曲结构 | STRING | — |
| 📑 音乐结构分析 | STRING | — |
| 📄 语言模型完整响应 | STRING | — |
| ℹ️ 处理信息 | STRING | — |