Nodes/ComfyUI-dapaoAPI/🎵Music3音乐提示词生成@炮老师的小课堂
ComfyUI Node

🎵Music3音乐提示词生成@炮老师的小课堂

Writing MiniMax Music 3 prompts with a full band's worth of knobs

By paolaoshi·Created 11 months ago·Updated 13 days ago· 232
🎵Music3音乐提示词生成@炮老师的小课堂
    • 🎼 音乐描述
    • 📝 歌词
    • 🌐 全局音乐信息
    • 🎙️ 人声细节
    • 🧱 编曲结构
    • 📑 音乐结构分析
    • 📄 语言模型完整响应
    • ℹ️ 处理信息
    ◄🔑 API密钥►
    ◄🤖 LLM模型gemini-3.7-flash►
    ◄📝 原始音乐需求温暖的原声流行歌曲,亲密女声,副歌逐步扩张并在结尾温柔收束。►
    ◄🎼 主风格自动识别►
    ◄🧬 融合风格无►
    ◄🎯 使用场景自动识别►
    ◄⏳ 目标曲长自动识别►
    ◄💫 情绪弧自动识别►
    ◄⏱️ 速度自动识别►
    ◄🎹 调性倾向自动识别/不指定►
    ◄🎚️ 拍号/律动自动识别►
    ◄🥁 核心律动自动识别►
    ◄🎙️ 人声配置自动识别►
    ◄📈 人声音域自动识别►
    ◄🗣️ 人声音色自动识别►
    ◄🎤 演唱方式自动识别►
    ◄👥 和声/伴唱自动识别►
    ◄🎻 核心乐器编制自动识别►
    ◄🧱 歌曲结构自动识别►
    ◄🎛️ 制作质感自动识别►
    ◄🌌 空间与混响自动识别►
    ◄🔥 编曲密度自动识别►
    ◄🚫 排除项无►
    ◄🌐 输出语言英文(Music 3推荐)►
    ◄📏 输出详略标准|250–450词►
    ◄📦 输出格式结构化文本►
    ◄🌡️ 温度0.35►
    ◄📝 最大输出令牌4096►
    ◄🎲 Top_P1.00►
    ◄🎲 随机种0►
    ◄⌛ 请求超时300►
    ◄🔗 外部音乐需求►
    ◄📝 歌词►
    ◄🔗 外部歌词►
    ◄➕ 补充约束►
    ◄✍️ 自定义主风格►
    ◄✍️ 自定义融合风格►
    ◄✍️ 自定义使用场景►
    ◄✍️ 自定义目标曲长►
    ◄✍️ 自定义情绪弧►
    ◄✍️ 自定义速度►
    ◄✍️ 自定义调性倾向►
    ◄✍️ 自定义拍号/律动►
    ◄✍️ 自定义核心律动►
    ◄✍️ 自定义人声配置►
    ◄✍️ 自定义人声音域►
    ◄✍️ 自定义人声音色►
    ◄✍️ 自定义演唱方式►
    ◄✍️ 自定义和声/伴唱►
    ◄✍️ 自定义乐器编制►
    ◄✍️ 自定义歌曲结构►
    ◄✍️ 自定义制作质感►
    ◄✍️ 自定义空间与混响►
    ◄✍️ 自定义编曲密度►
    ◄✍️ 自定义排除项►
    ◄✍️ 自定义输出详略►
    ◄🚫 出错时跳过false►

    Music generation models are like image models circa 2023: the prompt matters enormously and most people write it badly. DapaoMusic3CaptionPromptNode is the dapaoAI pack's answer - it calls an LLM to turn your one-line music idea into a structured MiniMax Music 3 description plus a full set of lyrics, with a frankly absurd number of musical knobs so the output actually matches what you meant.

    The knobs (yes, there are a lot)

    The required inputs read like a producer's checklist: 🎼 主风格 (20 options: pop, rock, electronic, folk, cinematic…), 🧬 融合风格, 🎯 使用场景, ⏳ 目标曲长, 💫 情绪弧 (the emotional arc - build, release, resolve), ⏱️ 速度 (BPM-ish), 🎹 调性倾向 (key), 🎚️ 拍号/律动 (time signature), 🥁 核心律动 (groove), 🎙️ 人声配置, 📈 人声音域, 🗣️ 人声音色, 🎤 演唱方式, 👥 和声/伴唱, 🎻 核心乐器编制 (17 options), 🧱 歌曲结构, 🎛️ 制作质感, 🌌 空间与混响, 🔥 编曲密度, and 🚫 排除项. Almost everything defaults to 自动识别 (auto), so you can just type a description and let it fill the blanks - or get surgical when you know exactly what you want.

    The output-side controls: 🌐 输出语言 (default English, which the tooltip recommends for Music 3), 📏 输出详略 (default "standard 250–450 words"), 📦 输出格式 (structured text default). Plus the standard temperature (0.35), max tokens (4096), Top_P, cache-only seed, and timeout.

    Every one of those enum fields has a matching ✍️ 自定义… string input, so when the preset list doesn't have your specific timbre you can type it.

    The lyric handling - the part people ask about

    📝 歌词 (and its 🔗 外部歌词 twin) has a clean contract: if you provide lyrics, the node preserves them exactly - it won't rewrite your words, it builds the musical description around them. If you don't, it composes original lyrics in the requested language. That's the right behavior for a tool like this: you keep creative control of the words, the LLM handles the sonic scaffolding.

    Outputs

    The first two are the money outputs - 🎼 音乐描述 and 📝 歌词 - and the tooltip says they can be wired directly into MiniMax's official Music3 Text Encode node, which is the whole point: this is a front-end for the official pipeline, not a standalone. The rest are structure and diagnostics: 🌐 全局音乐信息, 🎙️ 人声细节, 🧱 编曲结构, 📑 音乐结构分析, 📄 语言模型完整响应, ℹ️ 处理信息.

    The honest take

    Twenty-three musical parameter dropdowns is overkill for "I want a chill lo-fi beat," and the defaults exist precisely so you can ignore most of them. Where this node earns its keep is the specific brief - a client wants "a warm acoustic pop song, intimate female vocal, chorus that builds then resolves softly" and you need a Music 3 prompt that won't come back as elevator music. That's the LLM-as-translator pattern from llm-in-comfyui.md applied to audio, and it works the same way: the model is a translation layer, not a composer, and it can add flavor you didn't ask for. Scan the description before generating if exactness matters.

    Gotchas

    • Paid calls, no auto-retry. A truncated response means re-running and paying again - check 最大输出令牌 if outputs seem short for the "high detail" setting.
    • Seed is cache-only, never sent to the API.
    • Lyrics provided = preserved verbatim; leave 歌词 empty to get fresh writing. Both inputs together (📝 歌词 + 🔗 外部歌词) is redundant - use one.
    • Chinese-only labels; defaults are sane.

    Install: ComfyUI Manager → "dapaoAPI", or git clone https://github.com/paolaoshi/ComfyUI-dapaoAPI.git into ComfyUI/custom_nodes/, pip install -r requirements.txt, restart. dapaoAI key from api.dapaoai.com (register → wallet → redeem code → apply key in the default group).

    Category🤖dapaoAPI/🍬大炮API常用工具🍬

    Inputs (57)

    NameTypeDefaultDescription
    🔑 API密钥STRING密钥仅用于 https://api.dapaoai.com,不写入配置文件。
    🤖 LLM模型COMBOgemini-3.7-flash17 options: gpt-5.5, gpt-5.6-luna, gpt-5.6-terra, gpt-5.6-sol, gpt-6-astra, claude-fable-5, +11
    📝 原始音乐需求STRING温暖的原声流行歌曲,亲密女声,副歌逐步扩张并在结尾温柔收束。—
    🎼 主风格COMBO自动识别20 options: 自动识别, 东亚现代流行|C-pop/J-pop与电子、R&B、说唱融合, 东亚抒情与国风|华语/日系抒情、原声或管弦, 现代R&B与Neo-Soul, 灵魂、蓝调与福音, 电影感流行抒情, +14
    🧬 融合风格COMBO无21 options: 无, 自动识别, 东亚现代流行|C-pop/J-pop与电子、R&B、说唱融合, 东亚抒情与国风|华语/日系抒情、原声或管弦, 现代R&B与Neo-Soul, 灵魂、蓝调与福音, +15
    🎯 使用场景COMBO自动识别11 options: 自动识别, 流媒体完整歌曲, 影视剧情/片尾曲, 电影或游戏配乐, 品牌广告/产品短片, 短视频背景音乐, +5
    ⏳ 目标曲长COMBO自动识别7 options: 自动识别, 30–60秒, 1–2分钟, 2–3分钟, 3–4分钟, 4–5分钟, +1
    💫 情绪弧COMBO自动识别11 options: 自动识别, 克制铺陈 → 温暖释放 → 余韵收束, 脆弱低语 → 渐强 → 宣言式高潮, 平静神秘 → 紧张堆叠 → 史诗爆发, 忧伤回望 → 希望抬升 → 治愈落地, 明亮轻快 → 律动推进 → 庆祝式收尾, +5
    ⏱️ 速度COMBO自动识别8 options: 自动识别, 慢速|60–78 BPM, 中慢|79–96 BPM, 中速|97–115 BPM, 中快|116–132 BPM, 快速|133–155 BPM, +2
    🎹 调性倾向COMBO自动识别/不指定10 options: 自动识别/不指定, 明亮大调倾向, 忧郁小调倾向, 大小调转换, 五声音阶/国风调式, 布鲁斯调式, +4
    🎚️ 拍号/律动COMBO自动识别7 options: 自动识别, 4/4, 3/4, 6/8, 12/8, 自由节拍/无固定律动, +1
    🥁 核心律动COMBO自动识别12 options: 自动识别, 平稳流行律动, 半拍/半速重心, Swing摇摆律动, Shuffle切分律动, 四拍踩底舞曲律动, +6
    🎙️ 人声配置COMBO自动识别13 options: 自动识别, 纯器乐|禁止人声与歌词, 单人女声, 单人男声, 中性/不限定性别单人声, 男女对唱, +7
    📈 人声音域COMBO自动识别10 options: 自动识别, 低沉低音区, 温暖中低音区, 自然中音区, 明亮中高音区, 高音区与假声, +4
    🗣️ 人声音色COMBO自动识别12 options: 自动识别, 清澈明亮, 温暖柔和, 气声亲密, 醇厚磁性, 沙哑颗粒感, +6
    🎤 演唱方式COMBO自动识别12 options: 自动识别, 叙述式、克制, 细腻气声、贴耳, 抒情渐强、宽阔副歌, 强力真声与高音爆发, R&B转音与即兴Ad-lib, +6
    👥 和声/伴唱COMBO自动识别10 options: 自动识别, 无伴唱、单主唱, 副歌轻量叠唱, 贴近三度/六度和声, 宽阔多轨和声墙, 领唱与群体呼应, +4
    🎻 核心乐器编制COMBO自动识别17 options: 自动识别, 钢琴 + 弦乐 + 克制鼓组, 指弹木吉他 + 轻鼓 + 温暖贝斯, Rhodes电钢琴 + 圆润贝斯 + R&B鼓组, 合成器Pad + 琶音 + 电子鼓, 失真电吉他 + 贝斯 + 现场鼓组, +11
    🧱 歌曲结构COMBO自动识别10 options: 自动识别, Intro → Verse → Pre-Chorus → Chorus → Verse → Chorus → Bridge → Final Chorus → Outro, Intro → Verse → Chorus → Verse → Chorus → Outro, Intro → Verse → Pre-Chorus → Chorus → Post-Chorus → Bridge → Final Chorus → Outro, Intro → Build → Drop → Breakdown → Final Drop → Outro, Intro → Rap Verse → Hook → Rap Verse → Hook → Bridge → Final Hook → Outro, +4
    🎛️ 制作质感COMBO自动识别12 options: 自动识别, 自然有机、动态呼吸, 现代流行、清晰宽阔, 温暖复古、磁带/黑胶质感, 电影化宽动态与空间层次, 俱乐部级紧实低频与冲击力, +6
    🌌 空间与混响COMBO自动识别9 options: 自动识别, 亲密近场、主唱居中, 适度宽声场、清晰层次, 大空间厅堂混响, 电影化深景与环绕感, 俱乐部直接、有力且紧实, +3
    🔥 编曲密度COMBO自动识别8 options: 自动识别, 极简留白, 由疏到密逐步堆叠, 中等密度、层次清晰, 副歌宽阔、主歌克制, 持续高能密集, +2
    🚫 排除项COMBO无11 options: 无, 禁止人声, 禁止说唱, 禁止电子鼓与808, 禁止失真吉他, 禁止合唱团, +5
    🌐 输出语言COMBO英文(Music 3推荐)3 options: 英文(Music 3推荐), 中文, 双语:英文为主、中文注释
    📏 输出详略COMBO标准|250–450词4 options: 精简|180–250词, 标准|250–450词, 详细|450–650词, 自定义
    📦 输出格式COMBO结构化文本旧工作流兼容控件;当前始终输出可直连Music3的音乐描述与原始歌词。
    🌡️ 温度FLOAT0.350–2—
    📝 最大输出令牌INT4096512–16384—
    🎲 Top_PFLOAT1.000–1—
    🎲 随机种INT00–18446744073709550000只控制ComfyUI缓存,不发送给LLM。
    ⌛ 请求超时INT30030–1200—
    🔗 外部音乐需求optSTRING连接任意STRING节点;存在时与本节点的原始音乐需求合并。
    📝 歌词optSTRING—
    🔗 外部歌词optSTRING可连接歌词文本;存在时与本节点歌词合并并原样输出。没有任何歌词输入时由LLM自动生成。
    ➕ 补充约束optSTRING—
    ✍️ 自定义主风格optSTRING—
    ✍️ 自定义融合风格optSTRING—
    ✍️ 自定义使用场景optSTRING—
    ✍️ 自定义目标曲长optSTRING—
    ✍️ 自定义情绪弧optSTRING—
    ✍️ 自定义速度optSTRING—
    ✍️ 自定义调性倾向optSTRING—
    ✍️ 自定义拍号/律动optSTRING—
    ✍️ 自定义核心律动optSTRING—
    ✍️ 自定义人声配置optSTRING—
    ✍️ 自定义人声音域optSTRING—
    ✍️ 自定义人声音色optSTRING—
    ✍️ 自定义演唱方式optSTRING—
    ✍️ 自定义和声/伴唱optSTRING—
    ✍️ 自定义乐器编制optSTRING—
    ✍️ 自定义歌曲结构optSTRING—
    ✍️ 自定义制作质感optSTRING—
    ✍️ 自定义空间与混响optSTRING—
    ✍️ 自定义编曲密度optSTRING—
    ✍️ 自定义排除项optSTRING—
    ✍️ 自定义输出详略optSTRING—
    🚫 出错时跳过optBOOLEANfalse—

    Outputs (8)

    NameTypeDescription
    🎼 音乐描述STRING—
    📝 歌词STRING—
    🌐 全局音乐信息STRING—
    🎙️ 人声细节STRING—
    🧱 编曲结构STRING—
    📑 音乐结构分析STRING—
    📄 语言模型完整响应STRING—
    ℹ️ 处理信息STRING—