Nodes/TTS Audio Suite/๐Ÿ“ฆ RVC Dataset Prep
ComfyUI Node

๐Ÿ“ฆ RVC Dataset Prep

Get your audio ready to train a voice model

By diodiogodยทCreated about a year agoยทUpdated 22 days agoยท 1,098
๐Ÿ“ฆ RVC Dataset Prep
  • TTS_engine
  • opt_audio1
  • training_dataset
  • dataset_info
โ—„model_nameMyVoiceโ–บ
โ—„dataset_sourceโ–บ
โ—„training_sample_rate40kโ–บ
โ—„cpu_workers1โ–บ
โ—„chunk_seconds3.0โ–บ
โ—„overlap_seconds0.3โ–บ
โ—„max_volume0.95โ–บ
โ—„mute_ratio0.00โ–บ
โ—„reuse_existingtrueโ–บ

Training a voice model is 90% dataset prep and 10% pressing go. This node is the 90%. If you want to train an RVC (Real-time Voice Conversion) model - a .pth that turns any input voice into a target voice - you can't just point a trainer at a folder of clips. The audio has to be sliced into training-sized chunks, resampled to a consistent rate, run through a feature extractor (HuBERT), and pitch-analyzed (F0). This node does all of that and hands you a clean training_dataset for the ๐ŸŽ“ Model Training node.

It's the first node in the suite's training chain: RVC Dataset Prep โ†’ Model Training โ†’ Load RVC Character Model. Get the dataset right and the rest is mostly waiting.

How it works

You give it your source audio and a target sample rate, and it does the mechanical work: slicing long recordings into short chunks (with a little overlap so nothing gets cut mid-word), normalizing volume, extracting HuBERT features and F0 pitch curves, and caching the result so you don't reprocess the same data every time you tweak a training setting. That caching is the quiet win - dataset prep is slow, and reusing it across training runs saves real time.

The inputs and outputs that matter

  • TTS_engine - required; wire in an RVC engine node. This tells the training chain which pipeline you're on.
  • model_name - a name for what you're building (default "MyVoice"). It's how your dataset and eventual model are labeled.
  • dataset_source - a path, a zip, or a folder of audio. Alternatively, feed audio directly into opt_audio1 (an AUDIO input) if it's already in your graph.
  • training_sample_rate - 32k, 40k, or 48k. 40k is the standard default and a safe pick; 48k is higher fidelity but heavier, 32k is lighter. Match whatever you plan to train at.

The optional set is your quality/speed tuning: chunk_seconds and overlap_seconds control the slicing, max_volume and mute_ratio handle normalization and silence, cpu_workers parallelizes the (CPU-bound) processing, and reuse_existing (on by default) reuses that prep cache.

Two outputs: training_dataset (TRAINING_DATASET) โ†’ straight into the Model Training node, and dataset_info (text) telling you how many chunks it produced and whether it reused the cache.

Installing it

Part of the pack. ComfyUI Manager โ†’ search "TTS Audio Suite" โ†’ install โ†’ restart. Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/diodiogod/TTS-Audio-Suite.git
cd TTS-Audio-Suite
python install.py

The HuBERT and RMVPE feature/pitch models this node needs are auto-managed - they download on first prep. On Linux you'll also want the system audio libs the pack asks for (portaudio19-dev, libsamplerate0-dev) so resampling works cleanly.

Common issues & troubleshooting

Garbage dataset from messy audio. RVC training is only as good as its input. Use clean, single-speaker audio with no background music or noise - the slicing and feature extraction can't fix a bad source, they just chop it up. A few minutes of clean speech beats an hour of noisy speech.

It reprocessed everything again. If prep is re-running when you only changed a training parameter, check reuse_existing is on - that's the flag that lets you reuse the cache across runs instead of paying the slicing/feature cost every time.

Wrong sample rate mismatch. Pick training_sample_rate to match your intended training config; a dataset prepped at 40k and a trainer expecting 48k is asking for trouble. Decide the rate once, up front.

Prep is slow. It's CPU-bound feature extraction. Raise cpu_workers to use more cores - that's the dial that actually moves prep time.

CategoryTTS Audio Suite/๐ŸŽ“ Training

Inputs (11)

NameTypeDefaultDescription
TTS_engineTTS_ENGINEConnect the RVC engine here because dataset prep also needs HuBERT and pitch-extraction settings, not just the later training node.
model_nameSTRINGMyVoiceBase name for the trained RVC voice. Changing only this should not force dataset re-prep when the actual data/settings stay the same.
dataset_sourceSTRINGOptional folder, zip, or single audio file path. Relative paths resolve from ComfyUI input/ and input/datasets/. Clean single-speaker speech matters more than having one huge raw recording. Leave empty if you only use opt_audio inputs.
training_sample_rateCOMBO40kTarget training sample rate. 40k is the normal speech default and usually the right place to start.
cpu_workersoptINT11โ€“32CPU workers for slicing and feature extraction. 1-4 is usually enough; more is not automatically better, especially on Windows.
chunk_secondsoptFLOAT3.01โ€“15Target clip length after slicing. RVC usually likes short clips; 3.0s is a good default, and an effective range around 2-8s is normal.
overlap_secondsoptFLOAT0.30โ€“2Small overlap helps avoid harsh cuts between slices. Too much overlap just duplicates data and slows prep.
max_volumeoptFLOAT0.950.1โ€“1Peak normalization ceiling for extracted clips. 0.95 is a safe default; avoid pushing everything to 1.0 unless you like clipping risk.
mute_ratiooptFLOAT0.000โ€“0.5Extra silence augmentation ratio. Leave this at 0 for normal training; raise it only if the model is over-voicing silence or breathing badly.
reuse_existingoptBOOLEANtrueReuse the prepared dataset cache when settings match. Usually keep this on; rebuilding features every run is just wasted time.
opt_audio1optAUDIOOptional training audio clip. Connect one or more AUDIO sources here for quick in-graph tests, or combine them with a dataset path.

Outputs (2)

NameTypeDescription
training_datasetTRAINING_DATASETโ€”
dataset_infoSTRINGโ€”