RVC Voice Convert
ComfyUI Doesn't Do Voice Conversion — This Node Borrows RVC's Brains
- audio
- audio
- info
What it actually is
RVC Voice Convert is the "change the voice, keep the words" node. You feed it any audio - a line of TTS narration, an acapella, a recording - it runs it through an RVC (Retrieval-based Voice Conversion) model, and out comes the same performance in a completely different voice. It's the timbre-swap step that sits between "I have a voice" and "I have that character's voice," the kind of thing people chain in front of a lip-sync model or a song-cover pipeline.
Here's the honest setup: this node is a thin adapter. All the real work happens in a separate RVC-WebUI installation - the community's standalone voice-conversion tool that's been trained, fine-tuned and battle-tested for years outside ComfyUI. The node doesn't train anything, doesn't download models, and doesn't even bring its own inference engine. It's the piece of plumbing that lets a ComfyUI graph reach across and use RVC's. If you already live in RVC's world, this makes it a one-node trip; if you don't, the setup tax is on the RVC side, not here.
How it works (the clever part)
Most audio packs try to vendor torch and the model stack into ComfyUI, and that's exactly how you get the dependency Hell the audio community complains about - add one model, break three because of transformers or torch conflicts. This pack dodges it by cheating structurally: the node writes your audio to a temp WAV, then shells out to a subprocess running the RVC portable package's own Python 3.9 environment (rvc_root/runtime/python.exe), which calls RVC's vc_single() and writes the result back. Your ComfyUI install only needs soundfile. All of torch, faiss and rmvpe stay quarantined inside the portable env, where they belong.
The cost of that isolation is one config file. On first run it copies config.json.example into the pack folder, and until you edit rvc_root to point at a real RVC portable install, the node just shows a placeholder telling you to go configure it.
The inputs that matter
Most of the eleven knobs are RVC's standard tuning surface, and the defaults are sane - you can get good results touching almost nothing:
- audio - the AUDIO to convert (Load Audio → here).
- model_name - your
.pthmodel, scanned fromassets/weights/. Training checkpoints (D_*/G_*) are filtered out automatically. Drop new ones in and restart ComfyUI. - index_path -
auto(default) tries RVC's filename-based index auto-match,noneskips the FAISS index entirely, or pick any.indexfile scanned fromlogs/andassets/indices/. - f0_up_key - pitch shift in semitones; +12 is an octave up. This is the "make them sing it higher" knob.
- f0_method -
rmvpe(default) is the quality pick;pmis fastest;harvesthandles low voices better;crepeis the alternative quality option.
index_rate (0.75) blends how much the index's timbre signature gets mixed in - the main dial for how closely it matches the target voice. protect (0.33) guards consonants from getting mangled, lower = more protection. device flips to cpu when CUDA runs out.
The outputs are audio (wire it to Preview Audio, Save Audio, or VHS's audio output) and info - a STRING with RVC's diagnostics, mostly timing and which index got used.
Installing the real thing
cd ComfyUI/custom_nodes
git clone https://github.com/Wang-Huachen/ComfyUI-RVC-WebUI
…or grab it via ComfyUI Manager by searching "ComfyUI-RVC-WebUI," then restart. That's the easy half. Before anything shows up in the dropdown you need the RVC portable package (the example config points at a Windows install like D:\RVC20240604Nvidia50x0), and your own trained or downloaded .pth voice models inside its assets/weights/. There is no model download here - the pack ships zero weights.
Where people get burned
- Model list is empty -
rvc_rootis wrong or there's no.pthinassets/weights/. The config error message tells you exactly what's missing; read it. - Inference fails on a fresh install - the portable env itself is broken.
runtime/python.exe --versionshould answer before you debug anything else. - CUDA out of memory - set
devicetocpu; RVC is not huge, but it shares your card. - Auto index misses - if your model is
xxx_e250_s6000.pthbut the matching.indexdoesn't carry the_e250_s6000suffix,autofinds nothing. Pick the file manually.
Expect quality to track your RVC model, not this node. Feed it a well-trained voice and the audio-generation layer of ComfyUI finally has a real timbre-swap step; feed it a mediocre one and no amount of index_rate knob-twiddling saves it.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| audio | AUDIO | 输入要转换的音频 | |
| model_name | COMBO | -- 请先配置 config.json -- | RVC 声音模型(.pth 文件) |
| index_path | COMBO | auto | FAISS 索引文件(auto=自动匹配, none=不使用索引) |
| f0_up_key | INT | 0-24–24 | 音高偏移(半音),+12=升高一个八度,-12=降低一个八度 |
| f0_method | COMBO | rmvpe | 音高提取算法:rmvpe(推荐/效果最好) / pm(最快) / harvest(低音好) / crepe(效果好) |
| index_rate | FLOAT | 0.750–1 | 索引特征混合比率(0=不使用索引, 1=完全使用索引) |
| filter_radius | INT | 30–7 | 音高中值滤波半径(0=禁用, ≥3=启用) |
| resample_sr | INT | 00–48000 | 输出重采样率(0=使用模型原始采样率) |
| rms_mix_rate | FLOAT | 0.250–1 | RMS 音量包络混合(0=匹配输入音量包络, 1=保留输出音量包络) |
| protect | FLOAT | 0.330–0.5 | 清辅音保护(0=最大保护, 0.5=不保护) |
| device | COMBO | auto | 推理设备(auto=自动检测 CUDA) |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |
| info | STRING | — |