DIGIT Batch SRT From Video
Transcribe a whole folder of videos to SRT in one run
- log
- transcribed_count
- output_folder
The single-file version of this is handy. This one is the version you want when you have twenty-seven videos sitting in a folder tree and the idea of running them one at a time makes you want to quit the project. DIGIT Batch SRT From Video recursively scans a folder, transcribes every video to subtitles with Gemini, and either saves an .srt next to each clip or collects everything into one project folder. Set it going before lunch, come back to captions.
It's built for exactly the boring production job: dozens of files, unattended, resumable. The skip logic is smarter than most - it checks for the specific output you asked for (a sidecar .srt, a burned-in _subtitled video, or both), so re-running after a partial failure only touches the files that actually failed.
How it works
For each video: ffmpeg extracts a mono 16kHz WAV, Gemini transcribes it with timestamps, then the output runs through the same post-processing pipeline as the single-file node - hallucination removal, line-length enforcement (Netflix's 42-char standard by default), optional frame padding and snap-to-frame. You can also burn the subtitles straight into the video with full styling control instead of (or alongside) saving files.
The inputs that matter:
video_folder- top-level folder; it scans recursively. Required.file_types-all, or filter tomp4,mov,mxf,mkv,avi,m4v,qtwhen a folder mixes formats you don't all want.subtitle_output-srt_only,burn_in_only, orboth.output_mode-alongside_video(.srtnext to each clip, wherever it lives) orprojekts_auto_srt(all output collected to one project folder).overwrite- off by default, so existing outputs are skipped.identify_speakers- on by default; labels SPEAKER 1, SPEAKER 2, etc.translate_to- transcribe then translate, preserving all timing.output_format-srt,vtt,ass,txt, orall.
For burn-in there's the styling block: font_name, font_size, font_color, outline_width, shadow_depth, position, margin_v. And model (gemini-2.5-flash default) plus delay_seconds (1.0) to pace the API calls.
Outputs: log (a per-file OK/SKIPPED/ERROR listing with relative paths), transcribed_count, and output_folder.
Installing it
Same pack as the rest of the DIGIT family:
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
ComfyUI Manager works too - search comfyui-digit. Two environment requirements matter here specifically: ffmpeg must be on your PATH (it's a system dependency, not a pip one), and you need the pack's standard GCP setup - gcloud auth application-default login, a Vertex AI–enabled project. Set gcp_project_id on the node or let it auto-detect.
Where it trips people up
Most issues are the two setup ones: ffmpeg missing (you'll see "ffmpeg audio extraction failed" in the log) and credentials not authenticated. Both are one-line fixes. Worth knowing before you start: transcription is per-call billed to your GCP account, and a folder full of long videos adds up, so check delay_seconds and spot-test one file first. The skip logic means you can always re-run and only the failures get processed - that's the safety net.
Inputs (30)
| Name | Type | Default | Description |
|---|---|---|---|
| video_folder | STRING | Folder containing video files to transcribe. Scans recursively. | |
| file_types | COMBO | all | Which video file types to process. 'all' includes mp4, mov, qt, m4v, mkv, avi, mxf. |
| subtitle_output | COMBO | srt_only | srt_only: sidecar file(s). burn_in_only: hardcode subs into video. both: file(s) + burned-in video. |
| model | COMBO | gemini-2.5-flash | 9 options: gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.0-flash, gemini-2.0-flash-lite, gemini-3.1-pro-preview, +3 |
| output_mode | COMBO | alongside_video | alongside_video: save next to each video. projekts_auto_srt: save all to project auto_srt folder. |
| overwrite | BOOLEAN | false | Overwrite existing output files. If false, skips videos that already have output. |
| gcp_project_id | STRING | GCP project ID. | |
| gcp_region | STRING | GCP region. | |
| extra_instructionsopt | STRING | — | |
| identify_speakersopt | BOOLEAN | true | Try to identify and label different speakers. |
| pad_framesopt | INT | 00–120 | Extend each subtitle by this many frames on both sides (head and tail). |
| frame_rateopt | FLOAT | 23.9761–120 | Frame rate of the video. Used for pad_frames and snap-to-frame. |
| snap_to_framesopt | BOOLEAN | false | Snap all timestamps to nearest frame boundary. Prevents subtitle flicker on frame-accurate systems. |
| max_chars_per_lineopt | INT | 420–80 | Max characters per subtitle line. 42 = Netflix/broadcast standard. 0 = no enforcement. |
| max_linesopt | INT | 21–4 | Max lines per subtitle entry. Entries exceeding this get split. |
| remove_hallucinationsopt | BOOLEAN | true | Detect and remove repeated/hallucinated subtitle entries. |
| output_formatopt | COMBO | srt | Output format(s). 'all' saves SRT + VTT + ASS + TXT. |
| languageopt | COMBO | auto | Language of the audio. 'auto' lets Gemini detect. Improves accuracy when specified. |
| translate_toopt | COMBO | none | Translate subtitles to this language after transcription. 'none' = no translation. |
| font_nameopt | STRING | Arial | Font family for burn-in subtitles. |
| font_sizeopt | INT | 248–120 | Font size for burn-in subtitles. |
| font_coloropt | COMBO | white | Subtitle text color. |
| outline_coloropt | COMBO | black | Subtitle outline/border color. |
| outline_widthopt | INT | 20–8 | Outline thickness around subtitle text. |
| shadow_depthopt | INT | 10–8 | Shadow depth behind subtitle text. |
| positionopt | COMBO | bottom_center | Where to place subtitles on screen. |
| margin_vopt | INT | 300–200 | Vertical margin from screen edge (pixels at 1080p). |
| delay_secondsopt | FLOAT | 1.00–30 | Delay between API calls to avoid rate limiting. |
| projekts_rootopt | COMBO | 1 options: /root/PROJEKTS | |
| projectopt | COMBO | 1 options: (no projects found) |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| log | STRING | — |
| transcribed_count | INT | — |
| output_folder | STRING | — |