[llama.cpp] Media Diagnostics
Did the Model Actually See Your Image? Read the Receipt
- media_diagnostics
- all_media_evaluated
- vision_available
- audio_available
- video_available
- audio_count
- image_count
- video_count
- json
- formatted_text
Here's the failure this node exists to catch. You wire an image into Generate, pick a vision GGUF, hit queue, and the model writes you a beautiful paragraph of fiction about a picture it never received. No error. No warning. Just a confident caption of the void.
Vision-language models fail silently by default, and a missing mmproj projector is not a crash - it's a text model answering a text question. Media Diagnostics turns that into a boolean you can branch on.
What's inside the receipt
Every Generate and Decide node in this pack returns a media_diagnostics output. It's the ingestion receipt from the fork-specific multimodal path: which modalities the loaded model claims to support (vision, audio, video), how many items you asked it to ingest versus how many were actually evaluated, whether verification passed, which handler was in play, and whether the model was unloaded after the response.
This node expands that object into things a graph can use.
Input: media_diagnostics - one value, straight off a generate or decide node. That's the whole node.
Outputs, in order:
all_media_evaluated- the one that matters. True when everything you connected actually reached the model.vision_available,audio_available,video_available- capability flags reported by the loaded model.audio_count,image_count,video_count- how many items were evaluated, not how many you connected. Swap the two numbers and read that again; that's the diagnostic.json- the raw receipt, if you want to log it.formatted_text- a human-readable report you can preview directly.
The formatted version opens with a status line: MTMD MEDIA INGESTION: PASS, NO MEDIA, or UNAVAILABLE, then capability lines reading available/unavailable, then evaluated counts in n/N form for each modality, then the handler and the unload state. It's the single most useful thing to stare at when a caption looks unhinged.
How to wire it without being precious about it
Wire formatted_text into a preview/show-text node and read it after the first successful run of any new media workflow. That's the five-second habit that stops you burning an hour tuning prompts against a model that never got your image.
Once you trust a graph, branch on all_media_evaluated through a switch or logic node so a bad run fails loudly instead of shipping a hallucinated prompt downstream. The counts are useful in loops - if you fed six frames in and image_count reads one, you've found why your video prompt describes a single still.
Reading the three states
PASS means the media landed. Good.
NO MEDIA is normal when nothing was connected - a text-only rewrite run. Nothing to fix.
UNAVAILABLE is the interesting one. You asked for media and it wasn't ingested. The usual causes, in order of likelihood: no mmproj_path set on Create Native Session; a projector that doesn't match the model file; a modality your model simply doesn't handle (audio and video are model-specific - an image-only VLM will ignore a waveform forever); or a model profile whose handler doesn't cover what you're feeding it. The capability flags usually tell you which: audio_available: unavailable with an audio count of zero is a straight answer, not a bug.
Practical notes
Install the pack and you're done - this node has no dependencies of its own, no model, no llama binary. Manager: search llama multimodal, or:
cd ComfyUI/custom_nodes
git clone https://github.com/craftingmod/ComfyUI-llama-multimodal.git
Restart afterwards, and note it needs ComfyUI 0.19.3+. If you're migrating an older workflow, there's a deprecated alias of this node (the pre-rename OllamaImageList_LlamaCppMediaDiagnostics) kept alive for saved-workflow compatibility - same behaviour, and you should replace it with the current one when you next open that graph.
One honest caveat: the receipt reflects what the fork's multimodal path reported, and a model that ingests an image and then ignores it will still say PASS. It answers "did the media arrive", not "did the model use it well". That's still the question worth answering first. Multi-subject attribution - the layered failure the KB documents across every captioner - happens after ingestion, and no diagnostic node can save you from that one.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| media_diagnostics | OLLAMA_IMAGE_LIST_LLAMA_CPP_MEDIA_DIAGNOSTICS | — |
Outputs (9)
| Name | Type | Description |
|---|---|---|
| all_media_evaluated | BOOLEAN | — |
| vision_available | BOOLEAN | — |
| audio_available | BOOLEAN | — |
| video_available | BOOLEAN | — |
| audio_count | INT | — |
| image_count | INT | — |
| video_count | INT | — |
| json | STRING | — |
| formatted_text | STRING | — |