Nodes/was-node-suite-comfyui/Kandinsky 6 Audio VAE Loader
ComfyUI Node Runs on cloud

Kandinsky 6 Audio VAE Loader

The file that turns the sound half into actual sound

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
Kandinsky 6 Audio VAE Loader
    • VAE
    ◄vae_name▾►

    A Kandinsky 6 latent has two halves, and the picture half gets all the attention. This node loads the decoder for the other half. Without it your clip is silent, or rather it's a waveform nobody has turned into audio yet - the sound exists in the latent, in latent form, and this is what reads it out.

    How it works

    One file, two jobs. kandinsky6_audio_vae.safetensors holds the audio autoencoder and the vocoder that turns the autoencoder's output into an actual waveform. The node wires both up as a ComfyUI VAE, so the core nodes you already know can use it: VAE Decode Audio to get sound out of a sampled latent, VAE Encode Audio to push an existing waveform back in. It's 44.1 kHz, mono.

    The one input is vae_name, a dropdown of the files in models/vae. One audio VAE serves every Kandinsky 6 checkpoint - Lite or Pro, distilled or not - so you load it once and forget it.

    The input

    Just the file. The dropdown lists what's in models/vae; if the node offers you the string put the Kandinsky 6 audio_vae file in models/vae instead, nothing there matches and you need to download it.

    Get it from a kandinskylab/Kandinsky-6.0 repository - the file is audio_vae/diffusion_pytorch_model.safetensors, and you save it as ComfyUI/models/vae/kandinsky6_audio_vae.safetensors. That's the layout the pack's workflows use, and the layout it publishes at Hugging Face WAS/was-node-suite-weights under models/kandinsky6/, if you'd rather not dig through the original repos.

    One output: VAE.

    Installing it

    ComfyUI Manager, search WAS Node Suite v3, install, restart - or the manual route:

    cd ComfyUI/custom_nodes
    git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
    

    ComfyUI 0.14.0+ and Python 3.10+. Nothing gets pip-installed and no weights are fetched for you; the pack's features.network setting is off out of the box and the nodes name the file they want rather than downloading it.

    Where it goes wrong

    "This file is not the Kandinsky 6 audio VAE." The dropdown is set to something else in models/vae - very easy to do, because that folder already holds your HunyuanVideo VAE and half the video VAEs you've ever downloaded. Pick the kandinsky6_audio_vae file.

    Sound that's too quiet or oddly scaled. If you're running a pretrain Kandinsky 6 transformer rather than a released one, the audio scale the pack decodes at is wrong for it: released checkpoints want 0.5302, the pretrained ones 0.417. That's an edge case, but it looks exactly like "the model generates bad audio," so check which transformer you actually loaded before blaming the prompt.

    Commercial use. Worth reading the licence, because this is the one file in the whole Kandinsky 6 set that isn't clean. The transformer weights are MIT. The audio VAE bundles MMAudio's autoencoder under CC BY-NC 4.0 - non-commercial only, with BigVGAN's vocoder alongside it under MIT. If you're shipping something paid, the sound half of this model is the piece your lawyer will ask about.

    It doesn't make sounds by itself. This is a decoder, not a foley model. It has nothing to do with the MMAudio nodes people bolt onto a silent Wan clip; here the audio was already generated jointly with the picture, and this just reads it.

    CategoryWAS Suite/Loaders

    Inputs (1)

    NameTypeDefaultDescription
    vae_nameCOMBOThe audio_vae file from a Kandinsky-6.0 repository, saved under models/vae, as kandinsky6_audio_vae.safetensors. One file serves every Kandinsky 6 checkpoint.

    Outputs (1)

    NameTypeDescription
    VAEVAEThe audio VAE, for VAE Decode Audio and VAE Encode Audio. 44.1 kHz, mono.