Nodes/DIGIT Nodes/DIGIT Dataset Manager
ComfyUI Node

DIGIT Dataset Manager

Scan, create, and validate LoRA training datasets without leaving ComfyUI

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT Dataset Manager
    • dataset_path
    • report
    • image_count
    ◄actionscan►
    ◄dataset_path►
    ◄source_path►
    ◄dataset_namemy_dataset►
    ◄caption_ext.txt►
    ◄min_resolution512►
    ◄copy_imagestrue►

    The DIGIT Dataset Manager is the first step of the pack's LoRA training pipeline, and it's the one that stops you from training on a mess. It does the dataset housekeeping in four actions - scan, create, validate, and stats - so your training folder is a known quantity before a single caption is written or a single step is trained.

    Here's the workflow-shaped reason this node exists: training data quality beats every training knob, and the most common dataset failure is just sloppy plumbing - a stray 128px image mixed into a folder of 1024px ones, a missing caption, a source folder you forgot to copy. This node is the plumbing check.

    How it works

    The four actions:

    • scan (default) - look at dataset_path and report what's actually there: image count, resolution info, caption coverage. The "am I about to train on junk?" check.
    • create - build a new dataset by copying images from source_path into a new dataset folder, optionally filtering out anything below min_resolution (default 512). That filter is the point - a LoRA trained on tiny images learns to make tiny images.
    • validate - check an existing dataset for problems: missing captions, wrong formats, undersized images.
    • stats - summary numbers for the dataset.

    Inputs worth knowing: dataset_name (default my_dataset) names a create target, caption_ext (.txt) says what a caption file looks like, copy_images (default true) controls whether create copies the actual files or just references them, and min_resolution is the quality gate. Outputs are dataset_path (where the dataset ended up - wire this into the Captioner), report (the human-readable findings), and image_count.

    Installing it

    Standard pack install:

    cd ComfyUI/custom_nodes
    git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
    cd comfyui-digit
    pip install -r requirements.txt
    

    Or ComfyUI Manager → search comfyui-digit → install → restart. This node is fully local - no cloud credentials - and it pairs with the rest of the training suite: the Dataset Manager feeds dataset_path into the DIGIT Captioner, which feeds into Dataset Prep or straight to the trainer.

    How it fits the real workflow

    The habit that pays off: run scan before you caption and validate after. Scan catches the 200px screengrab hiding in your character set; validate catches the image whose caption got written to the wrong filename. Both are five-second runs that catch the failure modes that show up later as "the model doesn't look like my character." And when you create a dataset, remember the resolution filter is a floor, not a suggestion - character LoRAs want clean high-res sources with diverse angles, backgrounds, and lighting, and this node is where you enforce that before it costs you a training run.

    CategoryDIGIT

    Inputs (7)

    NameTypeDefaultDescription
    actionCOMBOscan4 options: scan, create, validate, stats
    dataset_pathSTRING—
    source_pathoptSTRING—
    dataset_nameoptSTRINGmy_dataset—
    caption_extoptSTRING.txt—
    min_resolutionoptINT51264–4096—
    copy_imagesoptBOOLEANtrue—

    Outputs (3)

    NameTypeDescription
    dataset_pathSTRING—
    reportSTRING—
    image_countINT—