Nodes/ComfyUI Top TTS/Top TTS 2.5 - Load Model
ComfyUI Node

Top TTS 2.5 - Load Model

The 15 minutes that unlock IndexTTS 2.5 on your own GPU

By whmc76·Created 2 months ago·Updated 2 months ago· 1
Top TTS 2.5 - Load Model
    • model
    • model_info
    ◄model_directoryIndexTTS-2.5►
    ◄deviceauto►
    ◄precisionbf16►
    ◄load_emotion_text_modelfalse►
    ◄use_cuda_kernelfalse►
    ◄use_torch_compilefalse►

    Every workflow in this pack starts here. Top TTS 2.5 - Load Model is the gatekeeper: it loads the official IndexTTS 2.5 weights from disk, and if it refuses to run, the other three nodes are pointless. The name is the whole pitch - this is fully local, offline TTS. No API, no key, no audio leaving your machine. That's the entire reason to tolerate the setup, and this node is the setup.

    How it works

    The loader does two jobs, and both are unusually careful. First it resolves your model folder inside ComfyUI/models/ and checks a hardcoded manifest of files - the main weights (config.yaml, codec.pth, gpt.pth, s2mel.pth, wav2vec2bert_stats.pt, the tiktoken vocab, and friends) plus the helper models (w2v-bert-2.0, BigVGAN, the campplus speaker embedder) and, if you flip on the emotion-text model, a QwenEmotion 0.6B checkpoint. Missing anything, it fails immediately with a list of exactly what's missing rather than half-loading or trying to fetch files mid-inference.

    Second, it spawns a separate subprocess worker that owns the model. This is the smart part and it's the pack's answer to ComfyUI's dependency hell. IndexTTS 2.5 pins transformers==4.52.1, which will happily fight whatever your other nodes want. Instead of downgrading your whole environment, install.py drops that exact transformers into a plugin-private folder and the worker runs with it on its own sys.path, with HF_HUB_OFFLINE=1 set so it never phones home. Your main ComfyUI environment stays untouched.

    The inputs that matter

    • device - auto (default), cuda, or cpu. Leave it on auto. If you pick cuda and PyTorch can't see a GPU, you get a clear error.
    • precision - bf16 (default) or fp32. The upstream community found FP32 noticeably richer on IndexTTS 2. BF16 is the sane default for VRAM, but if quality bugs you, fp32 is a click away.
    • load_emotion_text_model - this one costs you. Turning it on loads a ~0.6B QwenEmotion model that the Synthesize node's emotion_text input requires. Leave it off unless you're actually typing emotion text; it's extra VRAM you don't need for vectors or emotion audio.
    • use_cuda_kernel / use_torch_compile - both default to off, deliberately. The README is blunt: they're disabled to keep first installs stable. Enable them only once you know your CUDA build handles it.

    model_directory is a dropdown of the subfolders in ComfyUI/models, defaulting to IndexTTS-2.5 - that's where the download script puts things, so don't move it unless you know what you're doing.

    Outputs

    You get two: model (the opaque handle that wires into Synthesize and Unload Model) and model_info, a plain string like IndexTTS 2.5 | cuda:0 | bf16 | /path/to/models/IndexTTS-2.5. Stick model_info in a display node if you want to see what you're running.

    Install

    cd ComfyUI/custom_nodes
    git clone https://github.com/whmc76/ComfyUI-Top-TTS.git
    python -m pip install -r ComfyUI-Top-TTS/requirements.txt
    python ComfyUI-Top-TTS/install.py   # must run with ComfyUI's Python, then restart
    

    Then read UPSTREAM_MODEL_LICENSE.txt and download the weights:

    cd ComfyUI-Top-TTS
    python download_models.py --source huggingface --accept-license
    

    Where people get burned

    The classic failure is a FileNotFoundError listing missing files - that means you skipped the download (or the --accept-license flag, which the script refuses to run without). Run the download and restart ComfyUI. If the worker dies mid-session, the real error is in ComfyUI/temp/comfyui_top_tts_worker.log, not the ComfyUI console. And remember the weights are under the upstream bilibili Model Use License, not MIT - fine for tinkering, but read it before you build a product on it.

    Categoryaudio/Top TTS

    Inputs (6)

    NameTypeDefaultDescription
    model_directoryCOMBOIndexTTS-2.527 options: IndexTTS-2.5, audio_encoders, background_removal, checkpoints, clip, clip_vision, +21
    deviceCOMBOauto3 options: auto, cuda, cpu
    precisionCOMBObf162 options: bf16, fp32
    load_emotion_text_modelBOOLEANfalse—
    use_cuda_kernelBOOLEANfalse—
    use_torch_compileBOOLEANfalse—

    Outputs (2)

    NameTypeDescription
    modelTOP_TTS_2_5_MODEL—
    model_infoSTRING—