DIGIT Dataset Prep
Resize a folder of images into a training-ready dataset in one pass
- log
- processed_count
Before a folder of images can train a LoRA, it has to be the right size and the right format, and your caption sidecars have to travel with it. That's the whole job of DIGIT Dataset Prep: it takes a source folder, resizes every image to your target resolution, and copies the matching .txt captions into a clean output folder. Boring, necessary, and exactly the kind of step that derails a training session when you skip it.
It slots between the Dataset Manager/Captioner and the trainer in the pack's LoRA pipeline. If you've ever hit a training run that choked on a mixed-resolution folder, you know why this node exists.
How it works
For each image it resizes to resolution (default 1024, applied to both dimensions) using the resize_mode you pick:
fit(default) - maintain aspect ratio, no crop. Long images just end up shorter on one side.fill_crop- resize to fill the target, then center-crop. Guarantees exact target dimensions, but you lose edges.stretch- force both dimensions. Fast, but it distorts anything that isn't already the right ratio.pad- fit into the target with solid-color bars (pad_color_r/g/b, default black).
Then it writes to output_folder in output_format (png or jpg, with quality for JPEG) and, with copy_captions on (default), copies each matching .txt so your captions stay paired with the images. overwrite (off by default) controls whether existing output files get replaced.
Outputs are log (per-file progress) and processed_count - a count that should match what you put in.
Installing it
Standard pack install:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
Or ComfyUI Manager → search comfyui-digit → install → restart. Fully local, no API keys, no cloud.
Choosing the resize mode honestly
fit is the safe default and the right call for most training sets - modern trainers bucket across resolutions anyway, so preserving aspect ratio beats force-fitting everything to a square. Use fill_crop when your trainer demands exact dimensions and you're willing to sacrifice edges (and make sure the crop doesn't cut off the thing you're training). pad is for when the subject must survive intact and you don't mind black bars. stretch is the "I know what I'm doing" option and usually a mistake for anything but uniform source material. Whatever you pick, processed_count is your check: if it's lower than your source count, something didn't copy, and that's worth knowing before the trainer sees a hole in your dataset.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| source_folder | STRING | Path to folder containing source images. | |
| output_folder | STRING | Path to output folder. Created if it doesn't exist. | |
| resolution | INT | 1024256–4096 | Target resolution (used for both width and height). |
| resize_mode | COMBO | fit | How to handle aspect ratio differences. |
| output_format | COMBO | png | 2 options: png, jpg |
| quality | INT | 951–100 | JPEG quality (ignored for PNG). |
| overwrite | BOOLEAN | false | Overwrite existing files in output folder. |
| copy_captionsopt | BOOLEAN | true | Copy existing .txt caption files to the output folder. |
| pad_color_ropt | INT | 00–255 | Pad color red (pad mode only). |
| pad_color_gopt | INT | 00–255 | Pad color green (pad mode only). |
| pad_color_bopt | INT | 00–255 | Pad color blue (pad mode only). |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| log | STRING | — |
| processed_count | INT | — |