Nodes/ComfyUI-JZL-MiniMax-H3/JZL - 🎵 MiniMax Music3 提示词预设
ComfyUI Node

JZL - 🎵 MiniMax Music3 提示词预设

Turn a stack of music dropdowns into a MiniMax Music3 structured caption

By wjluoxiao·Created 13 days ago·Updated 3 days ago· 57
JZL - 🎵 MiniMax Music3 提示词预设
    • caption
    高级参数false
    音乐风格[官方] Lo-fi 嘻哈 / Lo-fi Hip-Hop
    风格融合不指定 / Unspecified
    速度不指定 / Unspecified
    调式不指定 / Unspecified
    调号不指定 / Unspecified
    拍号不指定 / Unspecified
    情绪氛围不指定 / Unspecified
    情绪演变不指定 / Unspecified
    使用场景不指定 / Unspecified
    段落结构跟随风格 / Follow Style
    主奏乐器跟随风格 / Follow Style
    辅助乐器跟随风格 / Follow Style
    制作质感跟随风格 / Follow Style
    人声配置跟随风格 / Follow Style
    声部音区跟随风格 / Follow Style
    人声音色跟随风格 / Follow Style
    人声唱法跟随风格 / Follow Style
    和声伴唱跟随风格 / Follow Style
    人声效果跟随风格 / Follow Style

    "JZL - 🎵 MiniMax Music3 提示词预设" (Music3 Caption Preset) is a music-theory form that writes a prompt for you. MiniMax Music3's official caption format is a "structured caption": three stacked sections - Global Metadata (genre/BPM/key/mood/emotion arc/scene/production), Vocal Details (arrangement/range/timbre/delivery/harmony/effects), and Arrangement (the per-section timeline). Getting that format right by hand is a slog. This node turns ~20 dropdowns into that three-part caption, in the shape the model was actually trained to read.

    It's aimed at the same "漫剧" (short drama) pipeline as the rest of the pack - you caption a track, feed it to MiniMax's Music3 node, and score your H3 clips with something that matches the mood you set in the video nodes. The style library is dense: 40 presets covering Lo-fi Hip-Hop, Chillhop, Jazz, Bossa Nova, Soul, R&B, Funk, Hip-Hop, Electronic, Ambient, Synthwave, House and more, each with hand-written global/vocal/arrangement blocks in the density of the official reference examples.

    How it works

    Mostly it's template assembly with taste. The 音乐风格 (music style) dropdown carries the per-style three-part base text; the other dropdowns fill in the structured fields - 风格融合 adds a secondary genre as influences (the official genre-router's dual-style fusion), 速度 picks a BPM band, 调式/调号/拍号 handle key/mode/time signature, 情绪氛围 and 情绪演变 set mood and its arc, 使用场景 covers listening context (studying, driving, gaming…), 段落结构 chooses the arrangement shape, 主奏乐器/辅助乐器 name the lead and support instruments, 制作质感 sets the production feel (clean digital, warm analog, tape lo-fi, vinyl crackle…), and the vocal block covers 人声配置 (instrumental/female/male/choir/acappella), 声部音区, 人声音色, 人声唱法, 和声伴唱, and 人声效果. Flipping 高级参数 reveals the full set; most fields default to "不指定/跟随风格" so you can set two or three and let the style carry the rest.

    Output: caption - the STRING you feed to the Music3 caption input.

    The input that matters

    • 音乐风格 - the one dropdown that does most of the work. Pick this first; everything else is refinement.
    • 高级参数 - toggles the advanced fields (style fusion, time signature, emotion arc, structure, etc.).

    One honest note from the tooltip: this node writes the caption; Music3's own cfg_scale (default 1.5) and top_k (default 50) live on the official node's advanced parameters, and this node doesn't touch them.

    How to install it

    Part of the JZL-MiniMax-H3 pack - ComfyUI Manager, search "ComfyUI-JZL-MiniMax-H3", or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/wjluoxiao/ComfyUI-JZL-MiniMax-H3
    

    Restart. No dependencies - it's pure string building, and it doesn't load or run Music3 itself (that's the official MiniMax music nodes, fed by your caption).

    Common issues

    Don't over-specify. The defaults are "unspecified/follow style" for a reason - set a genre, a mood, maybe a BPM band, and let the style's arrangement text do its job; stacking every dropdown fights the model's sense of the genre. Also, this is a Chinese-first pack, so the dropdown labels are bilingual but lead with Chinese - if a label looks familiar-but-not-quite, match by the English half.

    CategoryJZL/MiniMax

    Inputs (20)

    NameTypeDefaultDescription
    高级参数BOOLEANfalse开启后显示风格融合/拍号/情绪演变/段落结构/声部音区等可选高级项
    音乐风格COMBO[官方] Lo-fi 嘻哈 / Lo-fi Hip-Hop官方 Structured Caption 三段式:Global Metadata(genre/BPM/key/拍号/情绪/情绪演变/场景/制作)/ Vocal Details(配置/音区/音色/唱法/和声/效果)/ Arrangement(分段时间线)
    风格融合COMBO不指定 / Unspecified官方 genre-router 支持双风格融合,副风格以 influences 形式并入 Global Metadata
    速度COMBO不指定 / Unspecified6 options: 不指定 / Unspecified, 慢速 60-75 / Slow 60-75 BPM, 中速 80-100 / Mid 80-100 BPM, 中快 100-110 / Mid-Fast 100-110 BPM, 快速 110-135 / Fast 110-135 BPM, 极快 140+ / Very Fast 140+ BPM
    调式COMBO不指定 / Unspecified16 options: 不指定 / Unspecified, 大调 / Major, 小调 / Minor, 和声小调 / Harmonic Minor, 旋律小调 / Melodic Minor, 和声大调 / Harmonic Major, +10
    调号COMBO不指定 / Unspecified13 options: 不指定 / Unspecified, C / C, 降D / D-flat, D / D, 降E / E-flat, E / E, +7
    拍号COMBO不指定 / Unspecified8 options: 不指定 / Unspecified, 4/4, 3/4, 6/8, 12/8, 2/4, +2
    情绪氛围COMBO不指定 / Unspecified18 options: 不指定 / Unspecified, 梦幻 / Dreamy, 放松 / Laid-back, 温暖 / Warm, 忧郁 / Melancholic, 高能 / Energetic, +12
    情绪演变COMBO不指定 / Unspecified9 options: 不指定 / Unspecified, 由静到扬 / Gradual Rise, 由扬到静 / Gradual Fade, 先抑后扬 / Suppressed→Erupt, 先扬后抑 / Peak→Dissolve, 持续平静 / Steady Calm, +3
    使用场景COMBO不指定 / Unspecified16 options: 不指定 / Unspecified, 学习 / Studying, 雨夜 / Raining Outside, 深夜耳机 / Late-night Headphones, 驾驶 / Driving, 健身 / Workout, +10
    段落结构COMBO跟随风格 / Follow Style6 options: 跟随风格 / Follow Style, 标准流行 / Standard Pop, 简洁 / Compact, 无 Intro / No Intro, 含器乐间奏 / With Instrumental Break, 含 Breakdown / With Breakdown
    主奏乐器COMBO跟随风格 / Follow Style官方 Arrangement 支持指定 primary instrument(主奏乐器)
    辅助乐器COMBO跟随风格 / Follow Style官方 Arrangement 支持指定 supporting instrument;选「无」配合主奏乐器=独奏
    制作质感COMBO跟随风格 / Follow Style覆盖 caption 的制作质感描述。官方节点 cfg_scale(默认 1.5,CFG 引导强度)与 top_k(默认 50,采样候选截断)请在其高级参数中设置,本节点不改动
    人声配置COMBO跟随风格 / Follow Style8 options: 跟随风格 / Follow Style, 无人声 / Instrumental (No Vocals), 女声 / Female Vocal, 男声 / Male Vocal, 童声 / Child Vocal, 柔和中性 / Soft Androgynous Vocal, +2
    声部音区COMBO跟随风格 / Follow Style8 options: 跟随风格 / Follow Style, 女高音 / Soprano, 女中音 / Mezzo-Soprano, 女低音 / Alto, 男高音 / Tenor, 男中音 / Baritone, +2
    人声音色COMBO跟随风格 / Follow Style13 options: 跟随风格 / Follow Style, 清亮 / Clear Bright, 温暖 / Warm, 烟嗓 / Smoky Husky, 沧桑 / Weathered Gravelly, 气声 / Breathy, +7
    人声唱法COMBO跟随风格 / Follow Style15 options: 跟随风格 / Follow Style, 假声 / Falsetto, 半说半唱 / Half-Sung Half-Spoken, 说唱 / Rap Flow, 呼麦 / Throat Singing, 民谣 / Folk Storytelling, +9
    和声伴唱COMBO跟随风格 / Follow Style5 options: 跟随风格 / Follow Style, 无和声 / No Harmony, 双声部和声 / Double-Tracked Harmonies, 多人叠唱 / Stacked Harmonies, 伴唱垫底 / Backing Vocals
    人声效果COMBO跟随风格 / Follow Style6 options: 跟随风格 / Follow Style, 磁带延迟 / Tape Delay, 暖混响 / Warm Reverb, 自动调音 / Auto-Tune, 干声无效果 / Dry (No FX), 回声 / Echo

    Outputs (1)

    NameTypeDescription
    captionSTRING