Nodes/GPM-for-ComfyUI/GPM VLM Internal Diagnostics
ComfyUI Node

GPM VLM Internal Diagnostics

Run this the moment the scanner refuses to work

By Mystic419·Created 6 months ago·Updated 4 days ago· 1
GPM VLM Internal Diagnostics
    • summary_json
    • status_text
    ◄model_name<no gguf models found>►
    ◄mmproj_name(auto)►
    ◄n_ctx4096►
    ◄n_gpu_layers-1►
    ◄n_batch512►
    ◄threads0►

    Every pack that loads a model has a node like this, and this one is better than most. It answers the only question that matters when a local vision model isn't captioning your folder: is this a missing dependency, a bad model pair, or a build of llama-cpp-python that can't see your GPU?

    It doesn't caption anything and it doesn't write sidecars. It inspects, resolves, and prints a JSON report you can read or paste into a bug report.

    Why you'd reach for it

    The internal scanner is not Ollama and not an API - it's a local GGUF running through llama-cpp-python's vision path, in a worker process. That stack fails quietly in three ways: llama-cpp-python isn't installed at all, it's installed but built CPU-only (so n_gpu_layers is a wish and a twenty-minute scan takes all night), or your mmproj came from a different release than your model quant.

    That last one is the sneakiest. A multimodal GGUF pair is a language model plus a projector - the part that actually looks at the image. Mismatched, they may still load and produce fluent, hallucinated captions, which is worse than an error. Use a matching pair from the same release.

    How it works

    The node probes four layers and reports each separately:

    The import. Whether llama_cpp comes in at all, and its version.

    The backend. It calls the capability probes the binding exposes (llama_supports_cuda, llama_supports_metal, llama_supports_vulkan, llama_supports_hipblas, and friends), reads llama_print_system_info() if present, and scans module attributes as a fallback. Out of that it produces a status_text line you should read literally: CUDA-capable llama_cpp build detected, CPU-only llama_cpp build detected, GPU offload requested but backend capability could not be confirmed.

    The file resolution. Where model_name and mmproj_name actually resolve to on disk, and whether those paths exist. It searches ComfyUI/models/llm/, ComfyUI/models/llm/GGUF/ and ComfyUI/models/GGUF/.

    The handler. It infers the multimodal family from the filename (Qwen2.5-VL, Qwen3-VL, Gliese/Qwen3.5-VL, LLaVA, unverified Qwen-VL) and checks whether your llama-cpp-python exposes a chat handler class for that family. It also builds the constructor kwargs it would pass, then filters them against the signature of the installed Llama.__init__ and lists anything it had to drop. That's the honest part of this node: it tells you when your version of the binding silently ignores an argument you set.

    Inputs and outputs

    • model_name / mmproj_name - same dropdowns as the scanners, straight from the model folders above. mmproj_name defaults to (auto); leave it there unless you have several projectors and know which belongs to which model.
    • n_ctx - context size, 256–32768, default 4096. This is the window the scan prompt runs in, not a per-image budget.
    • n_gpu_layers - default -1, which means "offload everything you can". Keep it there unless you're deliberately splitting.
    • n_batch - default 512, the batch size the binding would be constructed with.
    • threads - default 0 (let the runtime decide); set it if you're CPU-bound and want to pin core count.

    Outputs are summary_json (the full report, pretty-printed) and status_text (the one-liner). Wire neither anywhere; use a preview or read the console.

    Install

    ComfyUI Manager → Gallery Prompt Manager, or manually, and this is the node that actually needs the special install step:

    cd ComfyUI/custom_nodes
    git clone https://github.com/Mystic419/GPM-for-ComfyUI GPM
    cd GPM
    python -m pip install -r requirements.txt
    python install.py
    

    install.py is the whole reason it exists: plain pip install llama-cpp-python gets you a CPU build. The installer checks whether llama_cpp already imports and leaves it alone if so, otherwise tries a cuBLAS wheel index, then archived wheels, then a CPU wheel. You can force the path:

    GPM_LLAMA_INSTALL_MODE=cpu python install.py
    GPM_LLAMA_INSTALL_MODE=cuda python install.py
    

    Model side, you need one GGUF + matching mmproj. The validated family is Qwen2.5-VL, with the Qwen2.5-VL Abliterated Caption GGUF used by the pack's own scanner workflow as the reference pair.

    Where people get burned

    • CPU-only build, and everything "works". The most common real failure here. status_text says CPU-only, and the fix is a CUDA wheel, not a bigger n_gpu_layers.
    • Qwen3-VL is not approved for scanning. The diagnostics node will happily detect a Qwen3-VL family and report that a handler exists - but the scanner gate blocks it, because the author hasn't validated its image-input behaviour. This is the one place the pack is genuinely behind the community: in this corpus Qwen3-VL gets far more captioning talk than Qwen2.5-VL, and it's the captioner people reach for now. If you have a Qwen3-VL GGUF you want to use for captions, expect this pack to decline until the author validates it.
    • Correct-looking paths that don't exist. The report resolves paths and checks model_path_exists/mmproj_path_exists separately. A stale dropdown entry from a moved file will show up here and nowhere else.
    • Diagnostics is visibility-only. It runs no installs and repairs nothing. It also prints a dependency block at ComfyUI startup regardless of whether you use it - if you see [GPM startup] lines in your console, that's normal.
    CategoryGPM

    Inputs (6)

    NameTypeDefaultDescription
    model_nameCOMBO<no gguf models found>1 options: <no gguf models found>
    mmproj_nameCOMBO(auto)2 options: (auto), <no mmproj models found>
    n_ctxINT4096256–32768—
    n_gpu_layersINT-1-1–200—
    n_batchINT51232–8192—
    threadsINT00–128—

    Outputs (2)

    NameTypeDescription
    summary_jsonSTRING—
    status_textSTRING—