Extensions/ComfyUI-vlo
ComfyUI Extension

ComfyUI-vlo

Utility nodes providing quality-of-life features for interaction with vlo, including memory loaders for media handling.

By PxTicks·Created 4 months ago·Updated 8 days ago· 0
PxTicks/ComfyUI-vlo
Nodes26
On cloudLocal install
Categorylatent/audio, utils/logic
Stars0
Updated8 days ago

Nodes (26)

LTX Set Audio Latent Binary Masks

Put silence (or sound) exactly where you want it in LTX's audio latent

latent/audio
vlo Gate None

A mute button for any wire in your graph

utils/logic
vlo Latent Composite Masked

Paste latent patches back without re-rendering the whole frame

latent/composite
vlo Logic Not

The boolean flip you'll wire a hundred times

utils/logic
vlo Mask to Latent Mask

Give video inpainting a pixel-space mask without fighting the VAE's time math

latent/mask
vlo Memory Load Audio

Pump audio into a ComfyUI graph without a file in sight

audio
vlo Memory Load Audio Batch

Feed audio to your graph from memory instead of a swelling input folder

audio
vlo Memory Load Image

Feed ComfyUI an image that was never written to your input folder

image
vlo Memory Load Image Batch

Feed ComfyUI batches of images without junking up your input folder

image
vlo Memory Load Video

Hand a video to ComfyUI from memory, not from your input folder

image/video
vlo Memory Load Video Batch

Batch video into your graph from memory — and decide per-video whether its sound comes along

image/video
MiniMax H3 Add Guides from Spec (experimental)

The latent half of an H3 guide spec — add pre-planned masked clips to your conditioning

model/conditioning/minimax
MiniMax H3 Add Masked Guide (experimental)

Anchor an image guide inside an H3 video — but only trust part of it

model/conditioning/minimax
MiniMax H3 Add Masked Guides from Video (experimental)

Drive an H3 video from a masked video — trust some moments, not all of them

model/conditioning/minimax
MiniMax H3 Apply Semantic Guides (experimental)

Let H3's Qwen encoder see your guide too — with the honesty rules attached

model/conditioning/minimax
MiniMax H3 Build Guide Spec (experimental)

Plan an H3 guide once so the latent half and the Qwen half tell the same story

model/conditioning/minimax
MiniMax H3 Build Guide Spec from Video (experimental)

Plan a whole masked video of H3 guides once, before anything touches the conditioning

model/conditioning/minimax
MiniMax H3 Guide Token Mask Preview

See the mask the way the model sees it — on H3's token grid, not your pixels

model/conditioning/minimax
MiniMax H3 Masked Guide: Pixel Fill (baseline)

The baseline every fancy H3 masked-guide trick has to beat — and it's this boring node

model/conditioning/minimax
MiniMax H3 Patch Masked Guides (experimental)

The node that makes H3 masked guides actually do something (and why it's pinned to a ComfyUI commit)

model/advanced/minimax
vlo MiniMax H3 Reference to Video (Batch)

Throw a whole batch of reference videos and audio at MiniMax H3 in one execution

model/conditioning/minimax
vlo Save Image Websocket (BMP)

Images to the frontend without ever hitting disk or PNG encode

api/image
vlo Save Video Websocket

A 'Save Video' that hands the clip to the frontend from memory

api/video
vlo Set Audio Latent Binary Masks

Mask the sound, not just the picture

latent/audio
vlo Time-to-Move (TTM)

Steal the motion from a reference clip and give it new content

advanced/model
vlo Video Convert FPS

Make a clip's frame rate match what the model expects

image/video
Readme

ComfyUI-vlo

Utility nodes for vlo. The workflow provides quality-of-life nodes for interaction with vlo, such as the memory loaders described below. It is not required for vlo to work with Comfy, but it IS used in the default workflows.

Memory loaders

The vlo Memory Load family of nodes (vlo Memory Load Image, vlo Memory Load Audio, vlo Memory Load Video).

Ordinarily, feeding media into ComfyUI means uploading it into the ComfyUI/input folder. When an external app like vlo is generating many short-lived inputs per run, that folder fills up quickly with throwaway files. The memory loaders avoid this: media is held in an in-memory registry and referenced by id, so nothing is written to ComfyUI/input just to be passed into a graph.

Each loader also keeps a disable_in_memory toggle, which falls back to loading the selected file from the normal input directory — handy for testing workflows by hand without going through vlo.

Batch loaders

The corresponding Batch nodes load an ordered multi-selection:

  • vlo Memory Load Image Batch outputs ordered IMAGE and MASK lists.
  • vlo Memory Load Audio Batch outputs an ordered AUDIO list.
  • vlo Memory Load Video Batch outputs an ordered VIDEO list plus a matching BOOLEAN "use audio" list, one flag per video.

These are ComfyUI list outputs rather than concatenated tensors. Images may therefore have different dimensions, audio clips may have different durations, and each video remains an independent video. Downstream nodes that need the whole collection in one execution must opt into ComfyUI list inputs; ordinary nodes will execute once per list item.

The batch nodes replace ComfyUI's static multi-select with an ordered selector owned by this extension. In the default mode it refreshes directly from vlo's in-memory registry. The disable_in_memory toggle reuses the same selector for files already present in ComfyUI/input. Arrow controls determine the exact list order sent downstream.

Selections are capped at 100 items as a general safety bound. Model-specific nodes should enforce their own lower limits when consuming a collection.

Per-video audio flags

The video loader also carries a per-item switch. include_audio is a comma-separated flag list in selection order (1,0,1); the selector renders it as a checkbox on each row, and vlo writes it from the speaker toggles on its batch slot. Unset items are false, and the flags travel with their video through adds, removals, and reordering. The loader emits them as its second output, so a consumer that takes a BOOLEAN list — such as the MiniMax H3 adapter's use_embedded_video_audio — receives one flag per delivered video.

MiniMax H3 batch adapter

vlo MiniMax H3 Reference to Video (Batch) wraps ComfyUI's native MiniMax H3 reference-conditioning node so the three batch loaders can feed it directly. It consumes each connected list in one execution and preserves its order. The wrapper reads reference limits and socket prefixes from the installed native node schema, and stops with a compatibility error if that contract changes.

The adapter converts each VIDEO to the native node's expected 24 fps image frames. The ref_audios socket is for standalone audio references.

Reference video audio

MiniMax treats a reference video's own soundtrack as a separate <Audio N> reference that has to be enabled: an ordinary reference video does not become an audio reference merely because its file contains sound. Enabling one also consumes an <Audio N> ordinal, which shifts the numbering of every later audio tag, because <Video N> and <Audio N> are numbered independently and the indices do not encode the pairing. The association is carried structurally, not by the tag numbers.

use_embedded_video_audio therefore defaults to off. It accepts either form:

  • a single value, which applies to every reference video;
  • a BOOLEAN list with one entry per reference video, bound positionally.

Both work because the node uses Comfy list inputs, so a widget arrives as a one-item list and a connected list arrives with one entry per video. The shipped vlo workflow uses the second form: the video batch loader's "use audio" output is linked to this input, so inclusion is decided per video rather than once for the whole batch.

An AUDIO list connected to ref_video_audios overrides soundtracks positionally and always wins, whether or not embedded audio is enabled for that video. Videos with neither an override nor enabled embedded audio are passed as video-only references.

The wrapper expands to a real native node in the execution graph rather than calling its Python method directly. ComfyUI therefore applies the native node's normal V3 lifecycle, validation, caching, and resource handling.

MiniMax H3 masked guides (experimental)

MiniMax H3 Add Masked Guide gives an H3 image guide a continuous spatial confidence mask: 1 keeps the guide at full strength, 0 corrupts that part of it to noise. It works by giving each guide token its own condition noise level and a matching condition timestep, generalizing the per-token modulation ComfyUI already uses for masked target rows.

The mask does nothing until the model passes through MiniMax H3 Patch Masked Guides, which installs a forked H3 forward pass. Samples without a masked guide take the stock path untouched, and a fully open mask is bit-identical to a stock MiniMaxH3AddGuide.

This is research code: it carries a copy of ComfyUI's MiniMaxH3Model._forward and is tied to the ComfyUI version it was forked from. See nodes/minimax_masked_guide/README.md for the semantics, the compatibility rules and the experiment protocol.

Installation

Clone (or symlink) this repository into your ComfyUI custom_nodes directory and restart ComfyUI:

cd ComfyUI/custom_nodes
git clone https://github.com/PxTicks/ComfyUI-vlo.git

The web extension under web/ is registered automatically.