VRGDG Z-Image LoRA Train Chunk
Train a Z-Image character or style LoRA in resumable chunks, ComfyUI-native
- model
- model
- latest_state_path
- log_path
- latest_comfy_lora_path
- output_name
- completed_steps
- total_target_steps
Z-Image is one of the friendlier image models to train against - a Qwen3 text encoder, a 16-channel latent, and community recipes that put a character LoRA at around 32 minutes on a 5090. This node wraps that training in ComfyUI, using a musubi-tuner install as the engine, so you never leave the graph. If the LTX audio node is the weird one in VRGameDevGirl's training trio, this is the straightforward one: point it at a folder of images, set a step budget, and collect a LoRA.
It runs on the same resumable "chunk" philosophy as the rest of the pack: each execution does steps_per_run steps (default 250), saves state, and stops. Re-run the graph and it resumes until completed_steps hits total_target_steps (default 3000). workspace_dir holds cache, logs, config and checkpoints - keep it on a disk with room, because the text-encoder cache and checkpoints add up.
How it works
dataset_images_dir points at your training images (or a parent folder the node organizes into an images subfolder). resolution_width/resolution_height (default 1024×1024) set the training bucket written into the musubi dataset config - 1024 is fine for Z-Image; if you follow the community's small-dataset recipe you'll often train lower and let bucketing handle the rest. num_repeats controls how many times each image-caption pair repeats, and create_captions will write missing .txt files from caption_text (with add_trigger_word/trigger_text if you want a trigger token prepended).
Then the modern-training knobs, and this is where the KB's advice earns its keep: you do not train the text encoder on a Qwen3-class model. The node caches text embeddings and unloads memory first (clear_memory_before_text_encoder on by default) - that caching is the single biggest speed lever. learning_rate_preset offers 1e-4 down to 1e-5 (Custom to set your own); 1e-4 is the default and a reasonable start. blocks_to_swap offloads transformer blocks to CPU - higher values cut VRAM but slow training. fp8_base/fp8_scaled are on by default; fp8_llm additionally loads the Qwen3 encoder in fp8 during caching if you're tight on VRAM.
The three model paths - musubi_root, zimage_checkpoint, vae, text_encoder - default to A:/MUSUBI/... Windows paths that you must repoint at your own downloads. Same as the LTX node, this is the first thing that trips people up.
Outputs and the feedback loop
model returns your base model with the latest LoRA optionally applied at strength_model. latest_comfy_lora_path is the converted ComfyUI-format LoRA (the node converts from musubi's format and, with copy_latest_to_comfy_loras on, drops a copy into your ComfyUI loras folder - keep_only_comfy_lora deletes the non-Comfy versions once the .comfy.safetensors exists). latest_state_path, log_path, output_name, completed_steps and total_target_steps complete the picture; wire the step counts into a compare node so the graph knows when to stop re-running.
Install and gotchas
Install the pack via ComfyUI Manager (search "vrgamedev") or cd ComfyUI/custom_nodes && git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl, restart, and install kornia, librosa, imageio from the README. The real setup is separate: a musubi-tuner install, the Z-Image base checkpoint, its VAE, and the Qwen3 text encoder files - all pointed at via the node's path inputs.
Watch the VRAM math before you start. A 5090 breezes through Z-Image character training, but on a 12GB card you'll want fp8_llm on and blocks_to_swap cranked up, and you'll be trading speed for memory. Keep the dataset small and clean - the community's best results are 15–25 images, well-captioned, and 2200 steps; throwing 200 images at it trains slower and usually worse.
Inputs (32)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | Base model to return downstream with the latest trained LoRA optionally applied. | |
| dataset_images_dir | STRING | Folder containing your training images, or a parent folder that will be organized into an images subfolder. | |
| workspace_dir | STRING | Working folder for cache, logs, config files, checkpoints, and training state. | |
| run_name | STRING | ZImageChunkRun | Name prefix used for the log file. |
| output_name | STRING | ZImageChunkRun | Name prefix used for saved LoRA files and state folders. |
| resolution_width | INT | 102464–8192 | Training bucket width written to the musubi dataset config. |
| resolution_height | INT | 102464–8192 | Training bucket height written to the musubi dataset config. |
| steps_per_run | INT | 2501–100000 | How many steps to train per run, and also when to save the LoRA/state at the end of that run. |
| total_target_steps | INT | 30001–1000000 | Training stops once the latest saved step reaches this total. |
| network_dim | INT | 321–2048 | LoRA rank. |
| network_alpha | INT | 321–2048 | LoRA alpha scaling value. |
| blocks_to_swap | INT | 40–28 | Higher values reduce VRAM usage but usually slow training. |
| clear_memory_before_text_encoder | BOOLEAN | true | Tries to unload ComfyUI models and clear VRAM/RAM before text encoder caching. |
| learning_rate_preset | COMBO | 1e-4 | Quick preset for the training learning rate. Choose Custom to use the float input below. |
| learning_rate | FLOAT | 0.00011e-8–1 | Custom learning rate used only when the preset is set to Custom. |
| num_repeats | INT | 11–1000 | How many times each image-caption pair is repeated in the dataset. |
| cache_strategy | COMBO | auto | Auto builds cache only when needed, Force always rebuilds it, Skip goes straight to training. |
| copy_latest_to_comfy_loras | BOOLEAN | true | Copies the latest Comfy-compatible LoRA into the ComfyUI loras folder after training. |
| keep_only_comfy_lora | BOOLEAN | false | If enabled, deletes the standard .safetensors LoRA files after a matching .comfy.safetensors file exists. |
| strength_model | FLOAT | 1.00-100–100 | Strength used if the node applies the latest LoRA back onto the output model. |
| create_captions | BOOLEAN | false | If enabled, missing caption txt files are created automatically using the caption text input. |
| caption_text | STRING | Base caption text used when create_captions is enabled and an image has no caption file. | |
| add_trigger_word | BOOLEAN | false | If enabled, the trigger text is prepended to each caption. |
| trigger_text | STRING | Trigger word or phrase to prepend to captions when add_trigger_word is enabled. | |
| musubi_root | STRING | A:/MUSUBI/musubi-tuner-ltx2 | Root folder of your musubi-tuner-ltx2 install. |
| zimage_checkpoint | STRING | A:/MUSUBI/models/zimage/zimage-base.safetensors | Path to the base Z-Image DiT checkpoint used for caching and training. |
| vae | STRING | A:/MUSUBI/models/zimage/vae.safetensors | Path to the Z-Image VAE checkpoint. |
| text_encoder | STRING | A:/MUSUBI/models/qwen3 | Path to the Qwen3 text encoder checkpoint or directory. |
| fp8_base | BOOLEAN | true | Enable fp8 base model weights during Z-Image training. |
| fp8_scaled | BOOLEAN | true | Enable scaled fp8 weights during Z-Image training. Requires fp8_base. |
| fp8_llm | BOOLEAN | false | Loads the text encoder in fp8 mode during caching to reduce VRAM usage. |
| use_32bit_attention | BOOLEAN | false | Use 32-bit precision for attention computations in the Z-Image model. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| latest_state_path | STRING | — |
| log_path | STRING | — |
| latest_comfy_lora_path | STRING | — |
| output_name | STRING | — |
| completed_steps | INT | — |
| total_target_steps | INT | — |