VRGDG_LoadAudioSplitDynamic
Split audio by hand-tuned scene lengths, not a fixed grid
- meta
- total_duration
- audio_1
- audio_2
- audio_3
- audio_4
- audio_5
- audio_6
- audio_7
- audio_8
- audio_9
- audio_10
- audio_11
- audio_12
- audio_13
- audio_14
- audio_15
- audio_16
- audio_17
- audio_18
- audio_19
- audio_20
- audio_21
- audio_22
- audio_23
- audio_24
- audio_25
- audio_26
- audio_27
- audio_28
- audio_29
- audio_30
- audio_31
- audio_32
- audio_33
- audio_34
- audio_35
- audio_36
- audio_37
- audio_38
- audio_39
- audio_40
- audio_41
- audio_42
- audio_43
- audio_44
- audio_45
- audio_46
- audio_47
- audio_48
- audio_49
- audio_50
Most audio splitters work on a grid: "every scene is N seconds." VRGDG_LoadAudioSplitDynamic is the version that lets you play director instead. You pick how many scenes you want, then set each scene's duration individually - scene 1 is 8 seconds, scene 2 is 12, scene 3 is 5 - because real songs don't have uniform verses, and you shouldn't have to pretend they do.
This is one of the pack's audio-to-scene nodes, and it's built around the idea that each scene clip feeds an audio-driven video model. The workflow family here covers two targets: HUMO (ByteDance's human-centric video model, which drives motion from the audio) and talking-head style pipelines. The using_infinite_talk dropdown exists precisely because the splitter was written with one of those in mind first - its own tooltip says "If using HUMO, change this to false." That tooltip is the author telling you which mode is which.
What it does
The inputs that matter:
- path (
STRING, default./audio.mp3) - where the audio lives. This node loads from a path rather than taking anAUDIOinput, which is the key difference from the sibling splitters. - offset_seconds (
FLOAT, default 0) - skip the first N seconds before splitting. Handy for trimming intros with no vocals. - scene_count (
INT, 1–50, default 1) - how many scene clips you want out. - using_infinite_talk (enum
false/true) - set tofalseif you're on the HUMO workflow, per the tooltip. - duration_1 .. duration_50 (
FLOAT, default 3 each) - the per-scene lengths, in seconds. These are your creative control: setscene_countto however many scenes you have, then dial each duration to match the music.
Outputs:
- meta (
DICT) - the scene metadata (counts, offsets), ready forVRGDG_Json2Stringif you want to hand it to an LLM. - total_duration (
FLOAT) - how long the summed scene clips come to; a sanity check that your durations roughly match the track. - audio_1 .. audio_50 (
AUDIO) - the individual scene clips. Wire each into whatever generates video from a scene.
How to use it like a human
Load the track, set scene_count, and start with the verse-chorus-verse structure you actually hear. Intro 4s, verse 8s, chorus 12s, verse 8s, chorus 12s, outro 6s - that's a six-scene split with six different durations. The defaults are all 3 seconds, which is a fine placeholder, but the whole point of this node is that you replace them with real timings. Check total_duration against the track length after you're done.
Installing it
Part of the comfyui-vrgamedevgirl pack. ComfyUI Manager: search "vrgamedev", install, restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/vrgamegirl19/comfyui-vrgamedevgirl
then install the README's requirements - this node is why librosa is on the list:
pip install -r custom_nodes/comfyui-vrgamedevgirl/requirements.txt
If it's not working
- Sum of durations doesn't match the track - that's on you, not the node; the last scene will just be shorter (or the split overlaps the track end). Verify against
total_duration. - Wrong target workflow - if your clips come out mistimed for the video side, re-check
using_infinite_talk. The tooltip is explicit: HUMO wantsfalse. - Path issues - it's a raw path string, so relative paths resolve against ComfyUI's working directory. If the node can't find the file, give it an absolute path.
Inputs (54)
| Name | Type | Default | Description |
|---|---|---|---|
| path | STRING | ./audio.mp3 | — |
| offset_seconds | FLOAT | 0.00 | — |
| scene_count | INT | 11–50 | — |
| using_infinite_talk | COMBO | false | If using HUMO, change this to false. |
| duration_1opt | FLOAT | 3.00 | — |
| duration_2opt | FLOAT | 3.00 | — |
| duration_3opt | FLOAT | 3.00 | — |
| duration_4opt | FLOAT | 3.00 | — |
| duration_5opt | FLOAT | 3.00 | — |
| duration_6opt | FLOAT | 3.00 | — |
| duration_7opt | FLOAT | 3.00 | — |
| duration_8opt | FLOAT | 3.00 | — |
| duration_9opt | FLOAT | 3.00 | — |
| duration_10opt | FLOAT | 3.00 | — |
| duration_11opt | FLOAT | 3.00 | — |
| duration_12opt | FLOAT | 3.00 | — |
| duration_13opt | FLOAT | 3.00 | — |
| duration_14opt | FLOAT | 3.00 | — |
| duration_15opt | FLOAT | 3.00 | — |
| duration_16opt | FLOAT | 3.00 | — |
| duration_17opt | FLOAT | 3.00 | — |
| duration_18opt | FLOAT | 3.00 | — |
| duration_19opt | FLOAT | 3.00 | — |
| duration_20opt | FLOAT | 3.00 | — |
| duration_21opt | FLOAT | 3.00 | — |
| duration_22opt | FLOAT | 3.00 | — |
| duration_23opt | FLOAT | 3.00 | — |
| duration_24opt | FLOAT | 3.00 | — |
| duration_25opt | FLOAT | 3.00 | — |
| duration_26opt | FLOAT | 3.00 | — |
| duration_27opt | FLOAT | 3.00 | — |
| duration_28opt | FLOAT | 3.00 | — |
| duration_29opt | FLOAT | 3.00 | — |
| duration_30opt | FLOAT | 3.00 | — |
| duration_31opt | FLOAT | 3.00 | — |
| duration_32opt | FLOAT | 3.00 | — |
| duration_33opt | FLOAT | 3.00 | — |
| duration_34opt | FLOAT | 3.00 | — |
| duration_35opt | FLOAT | 3.00 | — |
| duration_36opt | FLOAT | 3.00 | — |
| duration_37opt | FLOAT | 3.00 | — |
| duration_38opt | FLOAT | 3.00 | — |
| duration_39opt | FLOAT | 3.00 | — |
| duration_40opt | FLOAT | 3.00 | — |
| duration_41opt | FLOAT | 3.00 | — |
| duration_42opt | FLOAT | 3.00 | — |
| duration_43opt | FLOAT | 3.00 | — |
| duration_44opt | FLOAT | 3.00 | — |
| duration_45opt | FLOAT | 3.00 | — |
| duration_46opt | FLOAT | 3.00 | — |
| duration_47opt | FLOAT | 3.00 | — |
| duration_48opt | FLOAT | 3.00 | — |
| duration_49opt | FLOAT | 3.00 | — |
| duration_50opt | FLOAT | 3.00 | — |
Outputs (52)
| Name | Type | Description |
|---|---|---|
| meta | DICT | — |
| total_duration | FLOAT | — |
| audio_1 | AUDIO | — |
| audio_2 | AUDIO | — |
| audio_3 | AUDIO | — |
| audio_4 | AUDIO | — |
| audio_5 | AUDIO | — |
| audio_6 | AUDIO | — |
| audio_7 | AUDIO | — |
| audio_8 | AUDIO | — |
| audio_9 | AUDIO | — |
| audio_10 | AUDIO | — |
| audio_11 | AUDIO | — |
| audio_12 | AUDIO | — |
| audio_13 | AUDIO | — |
| audio_14 | AUDIO | — |
| audio_15 | AUDIO | — |
| audio_16 | AUDIO | — |
| audio_17 | AUDIO | — |
| audio_18 | AUDIO | — |
| audio_19 | AUDIO | — |
| audio_20 | AUDIO | — |
| audio_21 | AUDIO | — |
| audio_22 | AUDIO | — |
| audio_23 | AUDIO | — |
| audio_24 | AUDIO | — |
| audio_25 | AUDIO | — |
| audio_26 | AUDIO | — |
| audio_27 | AUDIO | — |
| audio_28 | AUDIO | — |
| audio_29 | AUDIO | — |
| audio_30 | AUDIO | — |
| audio_31 | AUDIO | — |
| audio_32 | AUDIO | — |
| audio_33 | AUDIO | — |
| audio_34 | AUDIO | — |
| audio_35 | AUDIO | — |
| audio_36 | AUDIO | — |
| audio_37 | AUDIO | — |
| audio_38 | AUDIO | — |
| audio_39 | AUDIO | — |
| audio_40 | AUDIO | — |
| audio_41 | AUDIO | — |
| audio_42 | AUDIO | — |
| audio_43 | AUDIO | — |
| audio_44 | AUDIO | — |
| audio_45 | AUDIO | — |
| audio_46 | AUDIO | — |
| audio_47 | AUDIO | — |
| audio_48 | AUDIO | — |
| audio_49 | AUDIO | — |
| audio_50 | AUDIO | — |