Nodes/YuE2-ComfyUI/YuE2 Options
ComfyUI Node

YuE2 Options

The four knobs worth touching, and the two to leave alone

By pytraveler·Created 22 days ago·Updated 5 days ago· 69
YuE2 Options
    • options
    ◄cotfull►
    ◄max_seconds0►
    ◄keep_model_loadedfalse►
    ◄cfg_scale0.00►
    ◄vaestandard►
    ◄deviceauto►
    ◄attention_backendsdpa►
    ◄downloadauto►
    ◄quantizationbf16►
    ◄ode_steps32►
    ◄abc_temperature0.70►
    ◄abc_top_p0.90►
    ◄abc_top_k30►
    ◄temperature1.00►
    ◄top_p0.95►
    ◄top_k100►
    ◄repetition_penalty1.200►
    ◄offloadauto►
    ◄transpose0►
    ◄vocals_onlyfalse►
    ◄low_vramfalse►

    YuE2 Generate Song asks for three things - style, lyrics, seed - and hides everything else on purpose. YuE2 Options is where "everything else" lives, and the important thing about it is what happens when you don't use it: an unconnected options socket is not a special case. The node starts from the released defaults either way. So this isn't a config you have to get right before the first song; it's a panel you open once you have a reason.

    Worth knowing where this sits first. The community's local music default is ACE-Step - fast, instrumental-leaning, weak exactly where YuE2 is aimed: vocals and lyrics. YuE2 is the other bet, a 3B model that writes a readable score and then sings it, and it's new enough that there's no accumulated folklore about which settings are good. The defaults are the model's own. Touch less than you think you need to.

    The inputs that matter

    Three are required, and they're the ones people actually change:

    • cot - how much is planned before anything is sung. full writes a melody-and-chords score first (the default, and where the benchmark numbers come from), melody plans the tune and lets the accompaniment follow the style, and off skips the score entirely and goes straight from lyrics to audio - faster, and no readable plan to edit. Stick with full unless you're chasing speed.
    • max_seconds - the length ceiling, 0 to 360 seconds, 0 meaning "work it out from the lyrics". Here's the part that bites: changing this changes the song, not just its length. The ceiling sizes the static KV cache and the captured CUDA graph, which reorders the attention reduction. Same seed, same lyrics, 40 and 90 second ceilings: identical for 79 tokens, divergent at the 80th. Settle this number before you go seed hunting, or you'll be chasing a moving target.
    • keep_model_loaded - keeps the 6.8 GB on the card after the run. Saves around five seconds per run while you iterate, holds the VRAM until ComfyUI restarts. Leave it off when video or image nodes come next.

    Then the optional fields you'll plausibly touch: attention_backend (sdpa by default and reproducible; cudnn is about 17% faster and produced four different songs from one seed over four runs - use it for exploring only), cfg_scale (0 means the released value, 1.0 normally; anything other than 1.0 runs a second unconditional branch, so roughly double the time and double the KV cache - raise it only when the result is ignoring your style), quantization (bf16 vs int8 - see below), download, vae (standard for listening, legacy to reproduce published numbers), device, ode_steps (32 is the released value; fewer is faster and thinner), and two banks of sampling numbers: abc_temperature / abc_top_p / abc_top_k for the score stage, and temperature / top_p / top_k / repetition_penalty for the song stage.

    The output is one handle, options (YUE2_OPTIONS), and it plugs into the matching socket on YuE2 Generate Song, YuE2 Write Song, YuE2 Plan Batch, YuE2 Render Plan or YuE2 Decode Latents. One options node can feed all of them.

    Two things people expect and don't get

    INT8 does not save VRAM. quantization: int8 downloads Comfy-Org's 3.69 GB build instead of 7.26 GB - real on a slow connection, nothing on your card. This pack's layers are ordinary torch linears, so the weights are restored to BF16 as the file loads and the GPU holds the same 6.8 GB either way. It also isn't quite the same model: the round trip costs about a percent of each weight, so the same seed gives a different song. Download size only.

    A second card is worth it, and a mixed one isn't. device is auto or cpu on a single-GPU install, with more entries when ComfyUI sees more cards; cpu works and takes about an hour a song. A second card stops YuE2 competing with ComfyUI's own VRAM, which is also when keep_model_loaded starts paying - but two YuE2 nodes pinned to different devices reload all 6.8 GB every run, because the loaded model is cached per device.

    Install

    Manager, search YuE2-ComfyUI (YuE2 Music), or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/pytraveler/YuE2-ComfyUI
    

    Restart after. tiktoken is the only new requirement; if it's missing the node prints the correct pip line for the interpreter ComfyUI runs on. Leave download on auto and the pack fetches the 7.26 GB Comfy-Org checkpoint into models/checkpoints/ on first run - the same file ComfyUI's own YuE2 nodes read, so one download serves both. off downloads nothing and tells you which files are missing, with a link and a folder for each.

    One last gotcha: if you wire an options node into YuE2 Render Plan to change cot on a plan that already exists, it won't. cot is part of the prompt the score was written under, so it always comes from the plan, and the node warns you that it's ignoring your change.

    CategoryYuE2

    Inputs (21)

    NameTypeDefaultDescription
    cotCOMBOfullHow much of the composition is planned before any audio is generated. 'full' writes a melody-and-chord score first and is the default for new songs. 'melody' plans the melody only and lets the accompaniment follow the style. 'off' goes straight from the lyrics to audio, which is faster and gives up the readable plan.
    max_secondsFLOAT00–360Ceiling on the length of the song. 0 works it out from the lyrics: about a minute for a verse and a chorus, longer as the words do. The model usually ends the song by itself, well short of the ceiling. The ceiling is for when it does not: on very few lines it can sing on long after the words have run out. Reaching it cuts the song off mid-phrase, and the node says so when that happens. Changing this number changes the song itself, not only its length -- it sizes the attention cache, and the same seed under a different ceiling is a different take. Leave it alone while you are hunting for a seed.
    keep_model_loadedBOOLEANfalseKeep the model loaded after the run. On saves about five seconds per run while you iterate on lyrics or seeds, and holds the memory until Unload Models is pressed, ComfyUI restarts or another YuE2 run needs a different model: on the card, or partly in system RAM when 'offload' has moved a half off. Off frees it immediately, which is what you want when video or image nodes run next in the same graph.
    cfg_scaleoptFLOAT0.000–20Classifier-free guidance. 0 means 'use the value the model was released with' -- 1.0 normally, 1.01 when 'cot' is 'off'. Anything other than 1.0 runs a second, unconditional branch: the song takes about twice as long and needs about twice the KV cache. Raise it only when the result ignores the style.
    vaeoptCOMBOstandardWhich audio decoder turns the latents into sound. 'standard' is the one released for listening. 'legacy' is the decoder the published benchmark numbers were measured with; use it only to reproduce those. They are different weights of the same size, and the same latents decoded by each will not sound identical.
    deviceoptCOMBOautoWhich device generates the song. 'auto' follows ComfyUI. Pick a second card and two things change: the model no longer competes with ComfyUI's own for VRAM, and 'keep_model_loaded' becomes worth turning on, because nothing has to be evicted to make room. Two YuE2 nodes with different devices in one graph will reload 6.8 GB on every run, because the loaded model is cached per device. 'cpu' works and takes roughly an hour per song. No CUDA device is visible to ComfyUI, so only 'cpu' will do anything.
    attention_backendoptCOMBOsdpaHow the model attends while it writes the song, token by token. 'sdpa' is the default and sings a seed as before. 'fast' is quicker and needs nothing extra. 'flash' is the quickest and needs the flash-attn package. Both need an RTX 30 card or newer. Each one repeats a seed, but each sings it its own way.
    downloadoptCOMBOautoWhere to get the weights when they are not on this machine yet. 'auto' fetches Comfy-Org's single checkpoint into ComfyUI/models/checkpoints. That is the same file ComfyUI's own YuE2 nodes read, so one download serves both, and the model manager may well have put it there already. 'original' fetches the three files m-a-p released, into ComfyUI/models/YuE2. The legacy decoder is published only by m-a-p. With 'vae' at 'legacy' and the model already on this machine -- m-a-p's files, or Comfy-Org's BF16 checkpoint, which holds the same weights bit for bit -- that one 0.49 GB file is all that is fetched. Comfy-Org's INT8 checkpoint counts only with 'quantization' at 'int8'. On a machine with nothing yet, 'auto' takes m-a-p's files for 'legacy', the smaller download. 'off' downloads nothing and says instead which files are missing, the direct link to each, and the exact folder to put it in.
    quantizationoptCOMBObf16Which build of the checkpoint to download. 'bf16' is the model as released. 'int8' is Comfy-Org's quantized build: 3.69 GB to fetch instead of 7.26 GB. It saves the download and not the VRAM: an INT8 file is restored to BF16 as it loads and the card holds the same 6.8 GB either way. What keeps weights packed on the card is 'low_vram', which does its own packing from whichever file you have. This one is also not quite the same model -- the round trip costs about a percent of each weight -- so the same seed gives a different song from the two files. Whatever is already on disk is used before anything is downloaded.
    ode_stepsoptINT328–64Solver steps for the acoustic stage. 32 is what the model was released with. Fewer is faster and thinner; more costs time and changes the result rather than clearly improving it.
    abc_temperatureoptFLOAT0.700–5Sampling for the score stage. The defaults are the released values.
    abc_top_poptFLOAT0.900.01–1Sampling for the score stage. The defaults are the released values.
    abc_top_koptINT301–1000Sampling for the score stage. The defaults are the released values.
    temperatureoptFLOAT1.000–5Sampling for the song stage. The defaults are the released values.
    top_poptFLOAT0.950.01–1Sampling for the song stage. The defaults are the released values.
    top_koptINT1001–1000Sampling for the song stage. The defaults are the released values.
    repetition_penaltyoptFLOAT1.2000.1–5Sampling for the song stage. The defaults are the released values.
    offloadoptCOMBOautoWhether the whole model stays on the card, or only the half the running stage uses. YuE2 has one set of weights that writes the score and the performance, and another that turns them into audio; no stage needs both. 'on' keeps only the half the stage needs, and neither during the decode, and it keeps only the rows of the vocabulary the running phase can use: measured, a 40-second song peaked at 4.4 GiB instead of 9.8, and a four-minute one at 4.5 GiB instead of 10.0, a second or two faster. 'off' keeps everything on the card. 'auto' moves a half off only when a stage would not fit beside it. The score, the notes and the words are the same in every mode: the weights are the same wherever they are kept. The audio file matches to the last byte too, as long as the decode has the same room to work in -- on a card so full that cuDNN has to pick a cheaper convolution, the last stage renders a hair differently, measured at 92 dB below the song. The weights waiting their turn sit in system RAM, up to 6.7 GiB.
    transposeoptINT0-12–12Moves the song to another key, in semitones: 2 is a whole tone up, -3 a minor third down, 0 sings the score as the model wrote it. YuE2 takes no key from the style line -- asked for 'A minor' there, it kept its own key every time -- but it follows its score closely, so this moves the score: every note, chord and key by the same step, just before it is sung. Measured, the song lands exactly that far away and the voice moves with it, the octave included: 12 puts the singer a full octave higher. It is a new take of the same tune rather than the old recording pitched up, because the moved score is sung from its first note. It needs a score, so not with 'cot' off, and a score it cannot read note by note is refused rather than guessed at. In 'YuE2 Generate Song' the 'score_abc' output is the moved score; in 'YuE2 Render Plan' the plan's score, or the one pasted in, is the one moved.
    vocals_onlyoptBOOLEANfalseOutputs only the voice. The song is made as always, then Mel-Band RoFormer separates the vocals from the band, and 'audio' carries the voice alone, the same length and rate. The song is the one this seed gives with the switch off, so an ordinary song keeps silence where its intro and instrumental breaks were. For an a cappella song, write 'a cappella' in the style: the model then keeps the voice going, while 'no instruments' in the style was measured to change nothing. Without separating, even an a cappella style leaves a soft pad under the voice in most songs. The first time, this downloads the separator (0.85 GB, MIT) into models/YuE2. For a recording made elsewhere, use 'YuE2 Vocals Only'.
    low_vramoptBOOLEANfalseFor a card of about 4 GB. The models' layers stay on the card as INT8, half their memory, and the work goes in smaller pieces -- the song and the models that hear its words alike. It is slower, and the same seed sings a different take. Leave it off unless the card needs it.

    Outputs (1)

    NameTypeDescription
    optionsYUE2_OPTIONS—