Nodes/DIGIT Nodes/DIGIT Caption Viewer
ComfyUI Node

DIGIT Caption Viewer

Actually look at the captions before you train on them

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT Caption Viewer
    • image
    • caption
    • filename
    • status
    • total
    ◄dataset_folder►
    ◄index0►
    ◄folder_path—►

    Anyone who's trained a LoRA has a horror story about a dataset that looked fine in the file browser and was secretly a mess - captions misaligned to images, a hundred images captioned as "a woman" when they were supposed to be the same character, a batch where the auto-captioner drifted off into describing nothing. The DIGIT Caption Viewer is the boring, necessary node that catches this: it steps through image + caption pairs in a folder so you can actually read what's going to be trained on.

    The best workflows treat captioning as a two-step loop - generate, then QA - and this is the QA half. Wire it after DIGIT Batch Caption and you can flip through every pair in your dataset before it ever reaches a trainer. Skipping this step is how "the model doesn't look like the character" happens.

    How it works

    Give it a dataset_folder containing image + .txt pairs, and use index (0-based) to walk through them. It loads the image and its caption sidecar, shows the preview right on the node, and exposes both to the rest of your graph. The optional folder_path input connects straight to DIGIT Batch Caption's folder_path output, so after a captioning run you just drag a wire and the right folder is already set.

    What comes out:

    • image - the current image as an IMAGE tensor, so you can preview it elsewhere or even feed it onward.
    • caption - the text sidecar for that image.
    • filename - which file you're looking at.
    • status - a readable status string.
    • total - how many pairs are in the folder, so you know where you are.

    The practical workflow is: connect folder_path from Batch Caption, hit run, and bump index to flip through. You're checking two things per image - is the caption accurate, and is it in the style you want - and flagging anything the captioner got wrong for a spot re-caption.

    Installing it

    It's in the digit-comfyui pack:

    cd ComfyUI/custom_nodes
    git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
    cd comfyui-digit
    pip install -r requirements.txt
    

    Or ComfyUI Manager → search comfyui-digit → install → restart. Like the other file-based DIGIT utilities, this node itself is fully local - no GCP credentials needed, it just reads the folder. (You'd only need the GCP setup if you're pairing it with a Gemini captioning node upstream.)

    The honest take

    This node is not clever and does not want to be. It's the closest thing the pack has to a proofreading pass, and it's worth more than it looks because captioning errors are the single most expensive mistake in training - the model happily learns whatever you fed it, including the wrong captions. One habit that pays off: don't just eyeball the first ten images, actually walk the folder, because captioners drift over long runs and the drift is always in the middle. total tells you how long the walk is; a few minutes of flipping usually catches the batch that would have wrecked your likeness.

    CategoryDIGIT

    Inputs (3)

    NameTypeDefaultDescription
    dataset_folderSTRINGPath to dataset folder containing image + .txt pairs.
    indexINT00–99999Index of the image to view (0-based).
    folder_pathoptSTRINGConnect from Batch Caption's folder_path output to auto-set the dataset folder.

    Outputs (5)

    NameTypeDescription
    imageIMAGE—
    captionSTRING—
    filenameSTRING—
    statusSTRING—
    totalINT—