Nodes/ComfyUI-LTX-AudioCaptioner/LTX-2.3 Audio-Video Captioner (Local)
ComfyUI Node

LTX-2.3 Audio-Video Captioner (Local)

A ComfyUI node in LTX-2.3 Dataset Tools with 16 inputs and 1 output.

By nerdydude364·Created 3 months ago·Updated 2 months ago· 0
LTX-2.3 Audio-Video Captioner (Local)
    • final_caption
    video_path
    trigger_namecharacter
    whisper_model_typecharacter
    overwrite_existingtrue
    detect_singingfalse
    avg_segment_length_threshold1.8
    energy_variance_threshold0.0050
    audio_channels1
    audio_sample_rate16000
    silence_rms_threshold0.0010
    low_rms_threshold0.010
    vocal_window_rms_threshold0.015
    audio_window_ms30
    music_rms_threshold0.100
    music_tonal_threshold0.30
    visual_caption
    CategoryLTX-2.3 Dataset Tools

    Inputs (16)

    NameTypeDefaultDescription
    video_pathSTRING
    trigger_nameSTRINGcharacter
    whisper_model_typeCOMBOcharacter5 options: base, tiny, small, medium, character
    overwrite_existingBOOLEANtrue
    detect_singingBOOLEANfalse
    avg_segment_length_thresholdFLOAT1.80.1–10
    energy_variance_thresholdFLOAT0.00500.0001–0.1
    audio_channelsCOMBO12 options: 1, 2
    audio_sample_rateCOMBO160005 options: 8000, 16000, 22050, 44100, 48000
    silence_rms_thresholdFLOAT0.00100–0.01
    low_rms_thresholdFLOAT0.0100.001–0.1
    vocal_window_rms_thresholdFLOAT0.0150.001–0.5
    audio_window_msINT305–200
    music_rms_thresholdFLOAT0.1000.001–0.5
    music_tonal_thresholdFLOAT0.300.01–1
    visual_captionoptSTRING

    Outputs (1)

    NameTypeDescription
    final_captionSTRING