YuE2 训练歌曲风格
Teach it a genre, not a voice
- style_model
- metadata
Two training nodes in one pack means people pick the wrong one. This is not the singer node - that's YuE2 训练 RVC 音色, which copies a voice. This one trains a LoRA on YuE2's autoregressive stage so generated songs come out sounding like a particular kind of song: production, arrangement, the vibe of your reference tracks. The vocalist stays YuE2's own.
Where it fits
Music generation is the thinnest layer in the ComfyUI stack - audio was bolted on late, and the music tools live at the edge in packs with their own dependency universes. The open-weight reference point has been ACE-Step: good instrumentals, vocals people complain about nearly as much as Suno's. YuE2 is the other bet - a 3B full-song model that starts from your lyrics and generates the singing itself - and this node is the local fine-tune knob for it. Expect the same soft spot either way: a style LoRA changes how a track sounds far more reliably than it fixes diction.
The model you train here has an obvious consumer: the optional 歌曲风格模型 input on YuE2 生成歌曲, strength 0 to 2. You can also pick it from that dropdown in the pack's studio, which is where you'll want to audition it.
What the node does - and what it doesn't
Same design as the rest of the pack: it's a launcher, not a trainer. It checks the worker on 127.0.0.1:8189, verifies training resources are installed, then calls the studio to preprocess and train the run you prepared. Underneath: MERT features from your songs, immutable train/validation snapshots pinned by content version, ~30-second semantic windows sampled out of complete songs, and validation against fixed start/middle/end windows with songs kept disjoint between sets - so you can compare checkpoints across a run instead of watching noise.
Training resources are pinned downloads: the Mothersuperior v4 tokenizer head, its NAR companion and a regularizer, about 414 MB, landing in <node-dir>/models/YuE2-training. The trained adapter currently works only in direct-generation mode - the pack says so up front rather than letting you find out the hard way.
Inputs
training_run is a combo from the studio's style-training records. Note the filter in the code: only runs in draft, failed, cancelled or paused state appear. A completed run drops out of the list - you work with the model in the studio, or feed it to the generation node.
action has three choices, and this is where beginners trip:
- 预处理并训练 (default) - the whole thing, in order.
- 仅预处理 - snapshot prep only, no training. It still returns a handle, but with an empty model asset.
- 训练或继续训练 - resume only. If the record was never preprocessed, it errors: 该训练记录还未预处理;请选择"预处理并训练".
So: wire the handle from a 仅预处理 run into the generation node and it has nothing to load. Start with the default action.
confirmed_materials_and_rights is a BOOLEAN, default off, and the node won't run without it. Unticked, you get 请先在工作台核对训练集、验证集与使用权利,再勾选确认.
Outputs are style_model (type YUE2_STYLE_MODEL, wire it into generation) and metadata (STRING - the prep and train job records). It's flagged as an output node, so it executes even with nothing downstream.
Preparing the run in the studio
This node can only execute a run that already exists, and building one is where the actual decisions are. Import at least two different songs - same song, different segments, can't straddle the train/validation split, and the pack enforces that. Give each vocal song lyrics and a style description (typed, from your library, or from the public list), and explicitly flag pure instrumentals. Then build the snapshot with the rights confirmation ticked and preprocess. Plans default to 200 steps for a quick orientation, with 800 and 1600 available. More steps is not more quality - pick a checkpoint on validation loss plus a fixed-prompt, fixed-seed A/B listen.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/Comfyui-YuE2-T8.git
Then run install_runtime.bat once from the node directory (shared Python 3.12.10 runtime, CUDA 12.8 Torch, FFmpeg, the model bundle) and restart ComfyUI. Manager users can search the pack title. Windows + NVIDIA; 24 GB VRAM and 60 GB disk recommended, and training data and checkpoints on top of that. If the ~414 MB of training resources didn't come down with the install, the studio has an 安装训练资源 button that fetches and verifies exactly that set.
Where it goes wrong
- The run dropdown shows nothing but the placeholder. Either no style run exists yet, every existing run has already completed (see the filter above), or the worker isn't reachable. Refresh the page after creating runs in the studio.
YuE2 训练资源、MERT 模型或 CUDA 环境尚未安装完整- the training resources are the pinned 414 MB set above. Install them from the studio rather than hunting for the files yourself; they're version-verified.- The song list is empty even though audio is in the asset library. Training only sees songs added to the current project. Add them from the asset card.
- The creation button stays disabled with a reason next to it. That reason is the checklist: missing style, missing lyrics, no validation song, or the rights box.
- Pause, don't kill. Safe pause finishes the current gradient accumulation boundary and saves model, optimizer and RNG state, so 继续训练 resumes properly.
And the license, since this is a training node: YuE2's first-party code and weights are CC BY-NC 4.0 - non-commercial - and the pinned NAR companion is Mothersuperior's CC BY-NC 4.0. Your training material's terms are on you. If you're building toward something you intend to sell, this pack is the wrong half of the ecosystem to bet on.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| training_run | COMBO | 1 options: 请先在 YuE2 工作台建立歌曲风格训练 | |
| action | COMBO | 预处理并训练 | 3 options: 预处理并训练, 仅预处理, 训练或继续训练 |
| confirmed_materials_and_rights | BOOLEAN | false | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| style_model | YUE2_STYLE_MODEL | — |
| metadata | STRING | — |