ComfyUI Node

RVC Voice Convert

ComfyUI Doesn't Do Voice Conversion — This Node Borrows RVC's Brains

By Wang-Huachen·Created 4 months ago·Updated 4 months ago· 0
RVC Voice Convert
  • audio
  • audio
  • info
◄model_name-- 请先配置 config.json --►
◄index_pathauto►
◄f0_up_key0►
◄f0_methodrmvpe►
◄index_rate0.75►
◄filter_radius3►
◄resample_sr0►
◄rms_mix_rate0.25►
◄protect0.33►
◄deviceauto►

What it actually is

RVC Voice Convert is the "change the voice, keep the words" node. You feed it any audio - a line of TTS narration, an acapella, a recording - it runs it through an RVC (Retrieval-based Voice Conversion) model, and out comes the same performance in a completely different voice. It's the timbre-swap step that sits between "I have a voice" and "I have that character's voice," the kind of thing people chain in front of a lip-sync model or a song-cover pipeline.

Here's the honest setup: this node is a thin adapter. All the real work happens in a separate RVC-WebUI installation - the community's standalone voice-conversion tool that's been trained, fine-tuned and battle-tested for years outside ComfyUI. The node doesn't train anything, doesn't download models, and doesn't even bring its own inference engine. It's the piece of plumbing that lets a ComfyUI graph reach across and use RVC's. If you already live in RVC's world, this makes it a one-node trip; if you don't, the setup tax is on the RVC side, not here.

How it works (the clever part)

Most audio packs try to vendor torch and the model stack into ComfyUI, and that's exactly how you get the dependency Hell the audio community complains about - add one model, break three because of transformers or torch conflicts. This pack dodges it by cheating structurally: the node writes your audio to a temp WAV, then shells out to a subprocess running the RVC portable package's own Python 3.9 environment (rvc_root/runtime/python.exe), which calls RVC's vc_single() and writes the result back. Your ComfyUI install only needs soundfile. All of torch, faiss and rmvpe stay quarantined inside the portable env, where they belong.

The cost of that isolation is one config file. On first run it copies config.json.example into the pack folder, and until you edit rvc_root to point at a real RVC portable install, the node just shows a placeholder telling you to go configure it.

The inputs that matter

Most of the eleven knobs are RVC's standard tuning surface, and the defaults are sane - you can get good results touching almost nothing:

  • audio - the AUDIO to convert (Load Audio → here).
  • model_name - your .pth model, scanned from assets/weights/. Training checkpoints (D_*/G_*) are filtered out automatically. Drop new ones in and restart ComfyUI.
  • index_path - auto (default) tries RVC's filename-based index auto-match, none skips the FAISS index entirely, or pick any .index file scanned from logs/ and assets/indices/.
  • f0_up_key - pitch shift in semitones; +12 is an octave up. This is the "make them sing it higher" knob.
  • f0_method - rmvpe (default) is the quality pick; pm is fastest; harvest handles low voices better; crepe is the alternative quality option.

index_rate (0.75) blends how much the index's timbre signature gets mixed in - the main dial for how closely it matches the target voice. protect (0.33) guards consonants from getting mangled, lower = more protection. device flips to cpu when CUDA runs out.

The outputs are audio (wire it to Preview Audio, Save Audio, or VHS's audio output) and info - a STRING with RVC's diagnostics, mostly timing and which index got used.

Installing the real thing

cd ComfyUI/custom_nodes
git clone https://github.com/Wang-Huachen/ComfyUI-RVC-WebUI

…or grab it via ComfyUI Manager by searching "ComfyUI-RVC-WebUI," then restart. That's the easy half. Before anything shows up in the dropdown you need the RVC portable package (the example config points at a Windows install like D:\RVC20240604Nvidia50x0), and your own trained or downloaded .pth voice models inside its assets/weights/. There is no model download here - the pack ships zero weights.

Where people get burned

  • Model list is empty - rvc_root is wrong or there's no .pth in assets/weights/. The config error message tells you exactly what's missing; read it.
  • Inference fails on a fresh install - the portable env itself is broken. runtime/python.exe --version should answer before you debug anything else.
  • CUDA out of memory - set device to cpu; RVC is not huge, but it shares your card.
  • Auto index misses - if your model is xxx_e250_s6000.pth but the matching .index doesn't carry the _e250_s6000 suffix, auto finds nothing. Pick the file manually.

Expect quality to track your RVC model, not this node. Feed it a well-trained voice and the audio-generation layer of ComfyUI finally has a real timbre-swap step; feed it a mediocre one and no amount of index_rate knob-twiddling saves it.

Categoryaudio/voice_conversion

Inputs (11)

NameTypeDefaultDescription
audioAUDIO输入要转换的音频
model_nameCOMBO-- 请先配置 config.json --RVC 声音模型(.pth 文件)
index_pathCOMBOautoFAISS 索引文件(auto=自动匹配, none=不使用索引)
f0_up_keyINT0-24–24音高偏移(半音),+12=升高一个八度,-12=降低一个八度
f0_methodCOMBOrmvpe音高提取算法:rmvpe(推荐/效果最好) / pm(最快) / harvest(低音好) / crepe(效果好)
index_rateFLOAT0.750–1索引特征混合比率(0=不使用索引, 1=完全使用索引)
filter_radiusINT30–7音高中值滤波半径(0=禁用, ≥3=启用)
resample_srINT00–48000输出重采样率(0=使用模型原始采样率)
rms_mix_rateFLOAT0.250–1RMS 音量包络混合(0=匹配输入音量包络, 1=保留输出音量包络)
protectFLOAT0.330–0.5清辅音保护(0=最大保护, 0.5=不保护)
deviceCOMBOauto推理设备(auto=自动检测 CUDA)

Outputs (2)

NameTypeDescription
audioAUDIO—
infoSTRING—