Extensions/ComfyUI-llama-multimodal
ComfyUI Extension

ComfyUI-llama-multimodal

Analyze ComfyUI image, audio, and video lists with Ollama, llama.cpp GGUF, or native generative CLIP backends

By craftingmod·Created 2 months ago·Updated about 15 hours ago· 2
craftingmod/ComfyUI-llama-multimodal
Nodes—
On cloudLocal install
Stars2
Updatedabout 15 hours ago
Readme

ComfyUI llama multimodal

Thumbnail

English | 한국어

LLM and multimodal nodes for ComfyUI.

Pass images of different sizes, video, and audio to multimodal LLMs through llama.cpp, with control over model loading and generation settings. Limited support for CLIP and Ollama is also available.

Useful for writing video prompts, describing media, and translating text.

Supported Runtimes

  • llama.cpp: Connect to an HTTP(S) server or use the llama executable on your PATH.
  • CLIP: Basic support for ComfyUI CLIP models with media of different resolutions.
  • Ollama: Basic support for multiple images. Audio and video are not supported.

If llama is not on your PATH, download it from Settings → llama-multimodal → Download llama.cpp.

For local GGUF models, place the model and its matching multimodal projector (mmproj) in ComfyUI/models/LLM.

Image, video, and audio support depends on the selected model.

Examples

Vision example

Download workflow

Generate text from multiple media inputs using a vision LLM.

More examples are available in EXAMPLES.md.

Install

Requires ComfyUI 0.19.3 or later.

  • ComfyUI Manager

Search for llama multimodal and install ComfyUI-llama-multimodal.

  • Comfy CLI
comfy node install ollama-image-list
  • Manual install
cd ComfyUI/custom_nodes
git clone https://github.com/craftingmod/ComfyUI-llama-multimodal.git

Development

After cloning the repository as described under Manual install, open the repository directory and run:

uv venv .venv
./.venv/Scripts/Activate
uv sync
bun install

Then build the frontend:

bun run build

More scripts are available in package.json.