ace_step_15_xl
Pick base or SFT, let the cloud sweat
- continue_from
- api_config
- moderation_status
- epochs
- workflow_id
- raw_json
Same cloud-training deal as the plain ACE-Step node, but for the bigger XL variant - and with one extra decision on top: model_variant, which asks whether you're training on the base model or the sft (supervised fine-tune) one. CivitaiTrainingAiToolkitAceStep15Xl runs Ostris's AI Toolkit trainer on Civitai's fleet to make an ACE-Step XL LoRA from a zip of your audio, billing in Buzz instead of GPU-hours. If you've already decided you want the XL model and its richer capacity, this is the node; if you're not sure, the plain AceStep15 node is the simpler entry point.
It sits in Civitai/Training/ace_step_15_xl, part of Civitai's official ComfyUI pack. ACE-Step, for context, is the open-weights answer to Suno - strong on instrumentals, famously weaker on vocals - and its community treats LoRA training on it the way image folks treat checkpoint LoRAs: capture a genre or an artist's production style and reuse it. XL is the beefier end of that lineup.
How it works
The node submits a training workflow (engine: ai-toolkit, ecosystem: ace_step_15_xl) to Civitai's Orchestration API. You point it at a hosted zip of training audio, choose which checkpoint variant to start from, and the cloud trains the LoRA, delivering one downloadable model per epoch. The model_variant field is the meaningful difference from the non-XL node: base starts from the raw pretrained model, sft from the supervised-fine-tuned release. SFT is usually the better starting point for style capture since it's already been nudged toward quality.
The core input is training_data_json, a JSON object (not a path):
{"type": "zip", "sourceUrl": "urn:air:ace-step-xl:dataset:civitai:98765@1", "count": 20}
sourceUrl is an AIR URN to the zip of training audio; count is the number of items, which is how the API prices the run.
The inputs that matter
- model_variant (required) -
baseorsft. When in doubt,sft. - training_data_json (required) - the zip + count object.
- epochs / steps - epochs = number of saved checkpoints (each a downloadable model), steps = total training length and the primary pricing control. Set one, the other derives.
- trigger_word - the token that calls up your LoRA in prompts later.
- network_dim / network_alpha, lr, batch_size, lr_scheduler, optimizer_type - the standard trainer dials, all optional.
The required-looking storage_buzz_per_epoch, default_steps, uses_step_pricing, and max_batch_size are generated plumbing; leave the defaults alone.
Outputs: moderation_status, epochs, plus the standard workflow_id and raw_json.
Installing it
This is one of ~160 nodes in Civitai Comfy Nodes, Civitai's official pack for their Orchestration API:
- ComfyUI Manager: Manager → Custom Nodes Manager → search Civitai Comfy Nodes → Install, then restart.
- CLI:
comfy node registry-install civitai-comfy-nodes - Source:
cd ComfyUI/custom_nodes && git clone https://github.com/civitai/civitai-comfy-nodes.git && pip install -r civitai-comfy-nodes/requirements.txt(justrequests).
You need a Civitai account with Buzz and credentials - a Civitai Auth node, CIVITAI_API_TOKEN (reliable for headless), or a stored key from the Civitai sidebar.
Where people get burned
- Malformed training_data_json. It must be valid JSON with
type,sourceUrl, andcount; anything else fails with a local parse error. - Variant confusion.
basevssftchanges what you build on and the results differ. Test one clip on each before running a full dataset. - Cloud pricing adds up. Steps are metered and each epoch carries a storage surcharge - check the cost report on the node after a run, because epochs are the line item that creeps.
- Moderation gates the model. Outputs pass through Civitai's content review;
moderation_statustells you the verdict. - Early preview. The README warns of unannounced changes, and early community reports include slow jobs and bugs. Long trainings can hit the default 30-minute timeout - raise it via the Auth node or
CIVITAI_COMFY_TIMEOUT.
If you're serious about ACE-Step style LoRAs and want the XL headroom without owning the hardware, this is the node. Pick SFT, feed it a clean zip, and let the cloud burn the electricity instead of you.
Inputs (25)
| Name | Type | Default | Description |
|---|---|---|---|
| model_variant | COMBO | base | 2 options: base, sft |
| training_data_json | STRING | Represents training data in various formats | |
| storage_buzz_per_epoch | FLOAT | 0.000–2147483647 | Per-epoch surcharge (buzz). Each epoch is a delivered checkpoint plus its preview samples, billed on top of the per-step training cost — so raising the epoch count raises the price by this much each. Override per ecosystem where per-epoch samples are expensive to compute (e.g. video). |
| default_steps | INT | 00–2147483647 | Default total step budget when neither Civitai.Orchestration.Grains.Workflows.Steps.Training.AIToolkit.AIToolkitTrainingInput.Steps nor Civitai.Orchestration.Grains.Workflows.Steps.Training.AIToolkit.AIToolkitTrainingInput.Epochs is supplied. Override per ecosystem where the default training length differs (e.g. video needs more steps, quickly-overtrained models need fewer). |
| uses_step_pricing | BOOLEAN | false | True when billing uses the per-step model. This is the default; the only exception is the legacy path where the caller supplied Civitai.Orchestration.Grains.Workflows.Steps.Training.AIToolkit.AIToolkitTrainingInput.Epochs but no Civitai.Orchestration.Grains.Workflows.Steps.Training.AIToolkit.AIToolkitTrainingInput.Steps (existing consumers), which keeps the historical flat per-epoch price. |
| max_batch_size | INT | 00–2147483647 | Ecosystem-specific maximum training batch size — the upper bound the user's Civitai.Orchestration.Grains.Workflows.Steps.Training.AIToolkit.AIToolkitTrainingInput.BatchSize is clamped to. Most ecosystems cap at 1. |
| samples_jsonopt | STRING | Sample generation configuration for training workflows | |
| epochsopt | INT | 11–200 | Number of training epochs — the number of saved checkpoints produced (each epoch yields one downloadable model). When omitted it is derived from Civitai.Orchestration.Grains.Workflows.Steps.Training.AIToolkit.AIToolkitTrainingInput.Steps; when both are supplied, both are honored (epochs = checkpoint count, steps = total). |
| stepsopt | INT | 11–10000 | Total number of training steps. This is the primary control over training length and determines pricing. When supplied, Civitai.Orchestration.Grains.Workflows.Steps.Training.AIToolkit.AIToolkitTrainingInput.Epochs (the number of saved checkpoints) is derived from it; when omitted, steps are derived from epochs. |
| batch_sizeopt | INT | 11–4 | Training batch size. Defaults to 1; raise it (up to the ecosystem's maximum) to train faster at the cost of more GPU memory. A larger batch sees more images per step, so fewer steps are needed for a comparable result. Values above the ecosystem maximum are clamped down. |
| lropt | FLOAT | 0.000–1 | Sets the learning rate for the model. This is the learning rate when performing additional learning on each attention block (and other blocks depending on the setting). |
| text_encoder_lropt | FLOAT | 0.000–1 | Sets the learning rate for the text encoder. Only used when TrainTextEncoder is true. For models with multiple text encoders, this applies to all of them. |
| train_text_encoderopt | BOOLEAN | false | Whether to train the text encoder(s) alongside the model. Enabling this can improve prompt understanding but increases training time and memory usage. |
| lr_scheduleropt | COMBO | You can change the learning rate in the middle of learning. A scheduler is a setting for how to change the learning rate. | |
| optimizer_typeopt | COMBO | The optimizer determines how to update the neural net weights during training. Various methods have been proposed for smart learning, but the most commonly used in LoRA learning is "adamw8bit". | |
| network_dimopt | INT | 11–256 | The larger the Dim setting, the more learning information can be stored, but the possibility of learning unnecessary information other than the learning target increases. A larger Dim also increases LoRA file size. |
| network_alphaopt | INT | 11–256 | The smaller the Network alpha value, the larger the stored LoRA neural net weights. For example, with an Alpha of 16 and a Dim of 32, the strength of the weight used is 16/32 = 0.5, meaning that the learning rate is only half as powerful as the Learning Rate setting. If Alpha and Dim are the same number, the strength used will be 1 and will have no effect on the learning rate. |
| noise_offsetopt | FLOAT | 0.000–1 | Adds noise to training images. 0 adds no noise at all. A value of 1 adds strong noise. |
| flip_augmentationopt | BOOLEAN | false | If this option is turned on, the image will be horizontally flipped randomly. It can learn left and right angles, which is useful when you want to learn symmetrical people and objects. |
| shuffle_tokensopt | BOOLEAN | false | Randomly changes the order of your tags during training. The intent of shuffling is to improve learning. If you are using captions (sentences), this option has no meaning. |
| keep_tokensopt | INT | 00–10 | If your training images have tags, you can randomly shuffle them. However, if you have words that you want to keep at the beginning, you can use this option to specify "Keep the first 0 words at the beginning". This option does nothing if the Shuffle Tokens option is off. |
| trigger_wordopt | STRING | A trigger word that activates the trained LoRA when used in prompts. Only applicable to certain ecosystems (sd1, sdxl, flux1, chroma, zimagebase, zimageturbo, flux2klein). | |
| continue_fromopt | CIVITAI_AIR | Optional previously-trained LoRA to continue training from ("train further"). When set, the first epoch resumes from this model instead of the base model, and the new epochs build on top of it. | |
| samples_overrides_jsonopt | STRING | — | |
| api_configopt | CIVITAI_CONFIG | Optional Civitai Auth connection; defaults to CIVITAI_API_TOKEN or stored OAuth login. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| moderation_status | STRING | — |
| epochs | STRING | — |
| workflow_id | STRING | — |
| raw_json | STRING | — |