DIGIT Dataset Manager
Scan, create, and validate LoRA training datasets without leaving ComfyUI
- dataset_path
- report
- image_count
The DIGIT Dataset Manager is the first step of the pack's LoRA training pipeline, and it's the one that stops you from training on a mess. It does the dataset housekeeping in four actions - scan, create, validate, and stats - so your training folder is a known quantity before a single caption is written or a single step is trained.
Here's the workflow-shaped reason this node exists: training data quality beats every training knob, and the most common dataset failure is just sloppy plumbing - a stray 128px image mixed into a folder of 1024px ones, a missing caption, a source folder you forgot to copy. This node is the plumbing check.
How it works
The four actions:
scan(default) - look atdataset_pathand report what's actually there: image count, resolution info, caption coverage. The "am I about to train on junk?" check.create- build a new dataset by copying images fromsource_pathinto a new dataset folder, optionally filtering out anything belowmin_resolution(default 512). That filter is the point - a LoRA trained on tiny images learns to make tiny images.validate- check an existing dataset for problems: missing captions, wrong formats, undersized images.stats- summary numbers for the dataset.
Inputs worth knowing: dataset_name (default my_dataset) names a create target, caption_ext (.txt) says what a caption file looks like, copy_images (default true) controls whether create copies the actual files or just references them, and min_resolution is the quality gate. Outputs are dataset_path (where the dataset ended up - wire this into the Captioner), report (the human-readable findings), and image_count.
Installing it
Standard pack install:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Or ComfyUI Manager → search comfyui-digit → install → restart. This node is fully local - no cloud credentials - and it pairs with the rest of the training suite: the Dataset Manager feeds dataset_path into the DIGIT Captioner, which feeds into Dataset Prep or straight to the trainer.
How it fits the real workflow
The habit that pays off: run scan before you caption and validate after. Scan catches the 200px screengrab hiding in your character set; validate catches the image whose caption got written to the wrong filename. Both are five-second runs that catch the failure modes that show up later as "the model doesn't look like my character." And when you create a dataset, remember the resolution filter is a floor, not a suggestion - character LoRAs want clean high-res sources with diverse angles, backgrounds, and lighting, and this node is where you enforce that before it costs you a training run.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| action | COMBO | scan | 4 options: scan, create, validate, stats |
| dataset_path | STRING | — | |
| source_pathopt | STRING | — | |
| dataset_nameopt | STRING | my_dataset | — |
| caption_extopt | STRING | .txt | — |
| min_resolutionopt | INT | 51264–4096 | — |
| copy_imagesopt | BOOLEAN | true | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| dataset_path | STRING | — |
| report | STRING | — |
| image_count | INT | — |