Extensions/ComfyUI-llama-multimodal
ComfyUI Extension

ComfyUI-llama-multimodal

ComfyUI Ollama node for image list

By craftingmod·Created 2 months ago·Updated 5 days ago· 3
craftingmod/ComfyUI-Ollama-ImageList
Nodes35
On cloudLocal install
Categoryllama_cpp/session, llama_cpp/decision
Stars3
Updated5 days ago

Nodes (35)

[llama.cpp] Connect Session

Attach ComfyUI to a llama.cpp Server You Already Have Running

llama_cpp/session
[llama.cpp] Create Question From Input

Build a Question Your llama.cpp Decide Node Can Actually Answer

llama_cpp/decision
[llama.cpp] Create Runtime Session

The Session Node That Starts Its Own llama.cpp Server

llama_cpp/session
[llama.cpp] Create Native Session

[llama.cpp] Create Native Session

llama_cpp/session
[llama.cpp] Decide

Make the Model Pick A, B, or C — and Show You the Odds

llama_cpp/decision
[llama.cpp] Decide (Media Sequential)

Run a Decision Over a Whole Media List

llama_cpp/decision
[llama.cpp] Decide (Prompt Sequential)

Multiple Choice, Batched, With the Probabilities Printed

llama_cpp/decision
[llama.cpp] Generate

[llama.cpp] Generate

llama_cpp/generate
[llama.cpp] Generate (Prompt Sequential)

Six Captions From One Image Without Reloading the Model

llama_cpp/generate
[llama.cpp] Generate (Media Sequential)

Caption Forty Images Without Reloading the Model Forty Times

llama_cpp/generate
[llama.cpp] Hardware Runtime Profile

The VRAM Knobs the Session Node Hides in Advanced

llama_cpp/profile
[llama.cpp] Media Diagnostics

Did the Model Actually See Your Image? Read the Receipt

llama_cpp/utils
[llama.cpp] Model Profile

Pick the Right Chat Handler Without Knowing It Exists

llama_cpp/profile
Muse Glimmer Response Parser

When Your Model Talks to Itself, This Pulls the Answer Out

llama_cpp/utils
[llama.cpp] Native Speculative Profile

Speculative Decoding as a Typed Socket (Mostly Leave It Off)

llama_cpp/profile
[llama.cpp] Prefill Profile

How Many Tokens Is That Image Worth? Prefill Profile Decides

llama_cpp/profile
[llama.cpp] Thinking / Reasoning Profile

Turn Thinking Off Before Your Captions Get Weird

llama_cpp/profile
[llama.cpp] Unload Session

Free the VRAM Before the Sampler Asks For It

llama_cpp/session
CLIP Text Encode (Multimodal)

The CLIP node in the Ollama pack — no server, no key, just your model's own tokenizer

model/conditioning/multimodal
Ollama Connectivity (images)

A fetch button for your Ollama server, so you stop mistyping model names

Ollama/images
Ollama Generate (with images)

The Ollama vision node that finally handles batches

Ollama/images
Jinja Chat Template Preset

The one-node reason Qwen chat templates stop lying to you

Ollama/preset
[llama.cpp] Gemma 4 Runtime Preset

Three runtime presets and the numbers behind them

llama_cpp/legacy
[llama.cpp] Generate (Multimodal)

The GGUF node that runs a vision model inside ComfyUI — no Ollama server required

llama_cpp/legacy
[llama.cpp] Media Diagnostics

Did it even see the media?

llama_cpp/utils
Llama.cpp Model Profile

The preset node that does the thinking about your model for you

llama_cpp/profile
[llama.cpp] N-gram Speculative Config

Model-free n-gram speculative decoding

llama_cpp/compact
[llama.cpp] N-gram Speculative Preset

N-gram speculative speedup for the detailed Generate, with a separate socket to prove it's off by default

llama_cpp/legacy
[llama.cpp] Generate

Model, hardware and reasoning decisions behind typed sockets

llama_cpp/compact
Llama.cpp Thinking / Reasoning Config

Thinking, effort and budget — your reasoning controls, decoupled from the model

llama_cpp/profile
[llama.cpp] Sampling Preset

The preset that keeps generation settings honest

llama_cpp/legacy
[llama.cpp] Media Sequential Generate

The sequential variant for batches

llama_cpp/compact
MiniMax System Prompt Preset

I2V, FL2V, T2V and friends, pre-written

Ollama/preset
Muse Glimmer Response Parser

Muse Glimmer talks in channels — this node separates the thinking from the answer

llama_cpp/utils
Ollama Options (images)

Ollama options without the guesswork — disabled means omitted

Ollama/images
Readme

ComfyUI llama multimodal

Thumbnail

English | 한국어

LLM and multimodal nodes for ComfyUI.

Pass images of different sizes, video, and audio to multimodal LLMs through llama.cpp, with control over model loading and generation settings. Limited support for CLIP and Ollama is also available.

Useful for writing video prompts, describing media, and translating text.

Supported Runtimes

  • llama.cpp: Connect to an HTTP(S) server or use the llama executable on your PATH.
  • CLIP: Basic support for ComfyUI CLIP models with media of different resolutions.
  • Ollama: Basic support for multiple images. Audio and video are not supported.

If llama is not on your PATH, download it from Settings → llama-multimodal → Download llama.cpp.

For local GGUF models, place the model and its matching multimodal projector (mmproj) in ComfyUI/models/LLM.

Image, video, and audio support depends on the selected model.

Examples

Vision example

Download workflow

Generate text from multiple media inputs using a vision LLM.

More examples are available in EXAMPLES.md.

Install

Requires ComfyUI 0.19.3 or later.

  • ComfyUI Manager

Search for llama multimodal and install ComfyUI-llama-multimodal.

  • Comfy CLI
comfy node install ollama-image-list
  • Manual install
cd ComfyUI/custom_nodes
git clone https://github.com/craftingmod/ComfyUI-llama-multimodal.git

Development

After cloning the repository as described under Manual install, open the repository directory and run:

uv venv .venv
./.venv/Scripts/Activate
uv sync
bun install

Then build the frontend:

bun run build

More scripts are available in package.json.