Nodes/DIGIT Nodes/DIGIT Batch SRT From Video
ComfyUI Node

DIGIT Batch SRT From Video

Transcribe a whole folder of videos to SRT in one run

By thedepartmentofexternalservices·Created 7 months ago·Updated 2 months ago· 0
DIGIT Batch SRT From Video
    • log
    • transcribed_count
    • output_folder
    ◄video_folder►
    ◄file_typesall►
    ◄subtitle_outputsrt_only►
    ◄modelgemini-2.5-flash►
    ◄output_modealongside_video►
    ◄overwritefalse►
    ◄gcp_project_id►
    ◄gcp_region►
    ◄extra_instructions►
    ◄identify_speakerstrue►
    ◄pad_frames0►
    ◄frame_rate23.976►
    ◄snap_to_framesfalse►
    ◄max_chars_per_line42►
    ◄max_lines2►
    ◄remove_hallucinationstrue►
    ◄output_formatsrt►
    ◄languageauto►
    ◄translate_tonone►
    ◄font_nameArial►
    ◄font_size24►
    ◄font_colorwhite►
    ◄outline_colorblack►
    ◄outline_width2►
    ◄shadow_depth1►
    ◄positionbottom_center►
    ◄margin_v30►
    ◄delay_seconds1.0►
    ◄projekts_root▾►
    ◄project▾►

    The single-file version of this is handy. This one is the version you want when you have twenty-seven videos sitting in a folder tree and the idea of running them one at a time makes you want to quit the project. DIGIT Batch SRT From Video recursively scans a folder, transcribes every video to subtitles with Gemini, and either saves an .srt next to each clip or collects everything into one project folder. Set it going before lunch, come back to captions.

    It's built for exactly the boring production job: dozens of files, unattended, resumable. The skip logic is smarter than most - it checks for the specific output you asked for (a sidecar .srt, a burned-in _subtitled video, or both), so re-running after a partial failure only touches the files that actually failed.

    How it works

    For each video: ffmpeg extracts a mono 16kHz WAV, Gemini transcribes it with timestamps, then the output runs through the same post-processing pipeline as the single-file node - hallucination removal, line-length enforcement (Netflix's 42-char standard by default), optional frame padding and snap-to-frame. You can also burn the subtitles straight into the video with full styling control instead of (or alongside) saving files.

    The inputs that matter:

    • video_folder - top-level folder; it scans recursively. Required.
    • file_types - all, or filter to mp4, mov, mxf, mkv, avi, m4v, qt when a folder mixes formats you don't all want.
    • subtitle_output - srt_only, burn_in_only, or both.
    • output_mode - alongside_video (.srt next to each clip, wherever it lives) or projekts_auto_srt (all output collected to one project folder).
    • overwrite - off by default, so existing outputs are skipped.
    • identify_speakers - on by default; labels SPEAKER 1, SPEAKER 2, etc.
    • translate_to - transcribe then translate, preserving all timing.
    • output_format - srt, vtt, ass, txt, or all.

    For burn-in there's the styling block: font_name, font_size, font_color, outline_width, shadow_depth, position, margin_v. And model (gemini-2.5-flash default) plus delay_seconds (1.0) to pace the API calls.

    Outputs: log (a per-file OK/SKIPPED/ERROR listing with relative paths), transcribed_count, and output_folder.

    Installing it

    Same pack as the rest of the DIGIT family:

    cd ComfyUI/custom_nodes
    git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
    cd comfyui-digit
    pip install -r requirements.txt
    

    ComfyUI Manager works too - search comfyui-digit. Two environment requirements matter here specifically: ffmpeg must be on your PATH (it's a system dependency, not a pip one), and you need the pack's standard GCP setup - gcloud auth application-default login, a Vertex AI–enabled project. Set gcp_project_id on the node or let it auto-detect.

    Where it trips people up

    Most issues are the two setup ones: ffmpeg missing (you'll see "ffmpeg audio extraction failed" in the log) and credentials not authenticated. Both are one-line fixes. Worth knowing before you start: transcription is per-call billed to your GCP account, and a folder full of long videos adds up, so check delay_seconds and spot-test one file first. The skip logic means you can always re-run and only the failures get processed - that's the safety net.

    CategoryDIGIT

    Inputs (30)

    NameTypeDefaultDescription
    video_folderSTRINGFolder containing video files to transcribe. Scans recursively.
    file_typesCOMBOallWhich video file types to process. 'all' includes mp4, mov, qt, m4v, mkv, avi, mxf.
    subtitle_outputCOMBOsrt_onlysrt_only: sidecar file(s). burn_in_only: hardcode subs into video. both: file(s) + burned-in video.
    modelCOMBOgemini-2.5-flash9 options: gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.0-flash, gemini-2.0-flash-lite, gemini-3.1-pro-preview, +3
    output_modeCOMBOalongside_videoalongside_video: save next to each video. projekts_auto_srt: save all to project auto_srt folder.
    overwriteBOOLEANfalseOverwrite existing output files. If false, skips videos that already have output.
    gcp_project_idSTRINGGCP project ID.
    gcp_regionSTRINGGCP region.
    extra_instructionsoptSTRING—
    identify_speakersoptBOOLEANtrueTry to identify and label different speakers.
    pad_framesoptINT00–120Extend each subtitle by this many frames on both sides (head and tail).
    frame_rateoptFLOAT23.9761–120Frame rate of the video. Used for pad_frames and snap-to-frame.
    snap_to_framesoptBOOLEANfalseSnap all timestamps to nearest frame boundary. Prevents subtitle flicker on frame-accurate systems.
    max_chars_per_lineoptINT420–80Max characters per subtitle line. 42 = Netflix/broadcast standard. 0 = no enforcement.
    max_linesoptINT21–4Max lines per subtitle entry. Entries exceeding this get split.
    remove_hallucinationsoptBOOLEANtrueDetect and remove repeated/hallucinated subtitle entries.
    output_formatoptCOMBOsrtOutput format(s). 'all' saves SRT + VTT + ASS + TXT.
    languageoptCOMBOautoLanguage of the audio. 'auto' lets Gemini detect. Improves accuracy when specified.
    translate_tooptCOMBOnoneTranslate subtitles to this language after transcription. 'none' = no translation.
    font_nameoptSTRINGArialFont family for burn-in subtitles.
    font_sizeoptINT248–120Font size for burn-in subtitles.
    font_coloroptCOMBOwhiteSubtitle text color.
    outline_coloroptCOMBOblackSubtitle outline/border color.
    outline_widthoptINT20–8Outline thickness around subtitle text.
    shadow_depthoptINT10–8Shadow depth behind subtitle text.
    positionoptCOMBObottom_centerWhere to place subtitles on screen.
    margin_voptINT300–200Vertical margin from screen edge (pixels at 1080p).
    delay_secondsoptFLOAT1.00–30Delay between API calls to avoid rate limiting.
    projekts_rootoptCOMBO1 options: /root/PROJEKTS
    projectoptCOMBO1 options: (no projects found)

    Outputs (3)

    NameTypeDescription
    logSTRING—
    transcribed_countINT—
    output_folderSTRING—