Nodes/YuE2 Music T8/YuE2 RVC 翻唱
ComfyUI Node

YuE2 RVC 翻唱

Swap the singer on a finished song, mix included

By T8mars·Created 23 days ago·Updated 2 days ago· 83
YuE2 RVC 翻唱
  • voice
  • song_audio
  • audio
  • result
  • metadata
◄semitone_shift0►
◄index_rate0.75►
◄protect0.33►
◄vocal_gain_db0.0►
◄accompaniment_gain_db0.0►

Most RVC setups make you do the boring half yourself: split the vocals in UVR or Demucs, convert them in the RVC/Applio GUI, drag the stem into a DAW, line it up against the instrumental. This node does all three steps in one block. Feed it a finished mixed song and an RVC voice, get back that song sung in the new voice with the original backing intact - the payoff node of this pack's RVC trio, where the loader or the trainer hands it a voice and this one spends it.

What it actually does

The node is a thin HTTP client, which is worth understanding before anything goes wrong. It writes the incoming audio to a temp WAV under <node-dir>/uploads/, then POSTs a voice_convert job to the pack's isolated worker on 127.0.0.1:8189. The worker loads Demucs (htdemucs, from models/Demucs/955717e8.safetensors), splits vocals from accompaniment, converts the vocal with RVC using RMVPE pitch extraction and the retrieval index built at training time, then mixes it back over the untouched backing as 48 kHz stereo.

None of that happens inside ComfyUI's Python environment. That's the pack's design argument: node packs normally pip-install into one interpreter with no isolation and end up fighting over transformers and torch. This one ships its own runtime and talks to it over localhost, so installing it doesn't break your other nodes - the trade is a second service that has to be up.

The inputs you'll actually touch

voice comes from YuE2加载 RVC 音色 or straight off YuE2 训练 RVC 音色, and song_audio is any AUDIO - a track you generated with the pack's own song node, or an existing file you loaded. One caveat baked into the code: exactly one clip. The node rejects a batch outright (YuE2 一次只接受一条 AUDIO;请先拆分批次), so don't feed it a loader that queued five files.

Then the interesting ones:

  • semitone_shift (−12 to +12). This is the knob you'll use most, and the pack's UI gives it octave buttons because that's how it's used in practice: −12 to put a female singer's line into a male voice's range, +12 the other way. It only transposes the converted vocal; the accompaniment stays where it was. Non-octave values are fine but can sit sour against the original chords - that's not a bug, that's music theory.
  • index_rate (default 0.75) is retrieval blend - how much the voice's stored feature index nudges the timbre toward the trained speaker. People treat it as a quality slider and max it. Don't. The pack's guide says it plainly: higher is not better, and if you hear garbled diction or buzzing, try turning retrieval down before blaming the model.
  • protect (default 0.33) shields unvoiced consonants and breaths from the timbre transfer - turn it up if word endings sound smeared. vocal_gain_db and accompaniment_gain_db (−18 to +12) are the mix, because conversion often changes how loud the vocal sits. "Vocals in a bucket" is usually just those two left at 0.

Outputs

audio is the finished mix - wire it to PreviewAudio or a save node. result is a YUE2_RESULT handle, which is what the pack's export node (YuE2 导出工件) wants if the artifacts need to outlive the retention window. metadata is the full job status JSON: model, params, paths. Handy for provenance, dead weight otherwise.

Installing

The nodes and the worker are separate installs, and you need both. Through Manager (search the pack title) or by hand:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/Comfyui-YuE2-T8.git

Then, once, from the node directory:

install_runtime.bat

That's the real install - it pulls the unified Python 3.12.10 runtime, CUDA 12.8 Torch, FFmpeg, the YuE2 and Vae weights, Demucs, SheetSage2, MERT, Seed-VC, and the RVC base models (HuBERT, RMVPE, and the G/D pretrained networks for v1 and v2). It's Windows + NVIDIA only, and the README asks for ~24 GB VRAM and 60 GB of free disk. Weights land under <node-dir>/models; move them elsewhere with configure_models.bat or the WebUI setting - and don't put anything in ComfyUI's checkpoints folder.

If you'd rather hear a voice before wiring a graph, start_webui.bat opens the studio, which has the same conversion page.

Where it goes wrong

  • "RVC 推理或人声分离组件尚未安装完整" - the worker's health check says RVC inference or separation isn't installed. Run install_runtime.bat, don't try to pip-install your way out.
  • The voice dropdown just says 没有已训练或导入的 RVC 音色 - the voice library is empty, or the service isn't running. It starts automatically when ComfyUI calls it; failures land in <node-dir>/logs/service.stderr.log.
  • "RVC 音色已被移动或删除" - dropdown entries are names plus a 32-char ID, read once when the graph loads. Train or import a voice, then refresh the page.
  • Screechy, hummy output on a song you made from stems - check what you trained on before blaming the conversion. Vocals-only material must be marked as such in the workbench, because a model trained on vocal tracks with bleed learns the bleed.
  • Strained top notes - a range problem, not a parameter. Try −12 and hear whether it holds up; if it doesn't, the fix is training material that covers the range, not a bigger index_rate.

One thing that isn't a bug: the YuE2 weights and first-party inference code in this pack are CC BY-NC 4.0 - non-commercial. The voice you're using has its own provenance, and the README asks you to only use voices you own or have explicit permission for. Worth knowing before a client hears the demo.

CategoryYuE2 音乐/RVC

Inputs (7)

NameTypeDefaultDescription
voiceYUE2_RVC_VOICE—
song_audioAUDIO—
semitone_shiftINT0-12–12—
index_rateFLOAT0.750–1—
protectFLOAT0.330–0.5—
vocal_gain_dbFLOAT0.0-18–12—
accompaniment_gain_dbFLOAT0.0-18–12—

Outputs (3)

NameTypeDescription
audioAUDIO—
resultYUE2_RESULT—
metadataSTRING—