Nodes/ComfyUI-llama-multimodal/[llama.cpp] Prefill Profile
ComfyUI Node

[llama.cpp] Prefill Profile

How Many Tokens Is That Image Worth? Prefill Profile Decides

By craftingmod·Created 2 months ago·Updated 5 days ago· 3
[llama.cpp] Prefill Profile
    • prefill_profile
    ◄profileCustom►
    ◄n_batch512►
    ◄n_ubatch0►
    ◄image_min_tokens0►
    ◄image_max_tokens0►

    A vision model doesn't read your image. It reads a number of tokens representing your image, decided by the projector, and that number is the whole trade: more tokens means it can see the small text and the fine detail, and it also means more compute, more memory and a slower queue.

    This node is where you set that number, along with the batch sizes llama.cpp uses to chew through the prompt. It's advanced-adjacent plumbing - you'll ignore it for weeks, then hit a case where it's the only fix.

    How it works

    Four integers, bundled into one OLLAMA_IMAGE_LIST_LLAMA_CPP_PREFILL_PROFILE cable you connect to Create Runtime Session (in the pack's own words, batch and image-token limits are shared by the llama.cpp Generate and session nodes; context size stays on each consumer). Named presets override the editable values, so Custom is where your numbers actually apply:

    • n_batch - prompt-processing chunk in tokens. Larger processes the prompt faster and wants more memory.
    • n_ubatch - the physical micro-batch. 0 lets llama.cpp use its own default.
    • image_min_tokens / image_max_tokens - the floor and ceiling on image tokens. 0 means "use the mmproj's or the handler's default".

    The presets are family-shaped: Gemma4 Medium (1024 batch / 768 ubatch, 280–560 image tokens), Gemma4 High (1536/1280, 560–1120), Qwen3.8 Medium (2048/1024, 1024–2048) and Qwen3.8 High (4096/1024, 1024–4096). Look at those image-token ranges - the Qwen profiles start where the Gemma ones stop. That's the concrete reason a Qwen-VL model can read a dense OCR page and a Gemma can't, and why the KB flags Qwen3-VL as the captioner people reach for when description accuracy matters.

    Two rules are enforced before anything runs: n_ubatch can't exceed n_batch, and image_min_tokens can't exceed image_max_tokens. Deliberately, context size is not here - n_ctx stays on each consumer, so one profile can drive several sessions with different context budgets.

    What to actually set

    Default (Custom, 512/0/0/0) is fine for casual use. Reach for this node when:

    A big image is blowing up prefill. Cap image_max_tokens and let the model get a coarser look. Blurry-but-answered beats an out-of-memory kill.

    You're doing OCR, UI screenshots, or anything with text in it. Nudge image_min_tokens up so the model isn't handed a thumbnail of a page. This is the single cheapest accuracy win in the whole pack for document-ish inputs.

    You want throughput on a batch of simple images. Lower the image-token ceiling and raise n_batch; captioning a set of portraits doesn't need 4096 vision tokens per frame.

    You're on a card that's already full. Drop n_batch first. A prefill batch that doesn't fit is the usual reason a run dies right at the start, before any tokens are generated.

    Install

    Manager → search llama multimodal → ComfyUI-llama-multimodal. Or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/craftingmod/ComfyUI-llama-multimodal.git
    

    ComfyUI 0.19.3 or later. No model needed for this node; it's only meaningful next to a session. Sessions need the llama runtime (Settings → llama-multimodal → Download llama.cpp, or your own on PATH) and a GGUF plus matching mmproj in ComfyUI/models/LLM.

    Where people get burned

    Values only apply when the session is built. Edit, then re-queue. A running server doesn't pick up a new batch size.

    Turning image_max_tokens up is not free. People read "more tokens, better vision" and set 8192, then wonder why a vision-only workflow now takes four times as long. The ceiling is a budget, not a quality slider.

    Zero doesn't mean zero. 0 means "let llama.cpp or the handler choose", not "send no image tokens". The author's tooltips say so, and misreading it as "off" has cost people an afternoon.

    The mmproj decides the available range. If your projector only supports a fixed token count, this node can't invent a bigger one - the model-side default wins. If bumping the numbers changes nothing measurable, that's why: check what your mmproj supports before assuming the node is broken.

    Categoryllama_cpp/profile

    Inputs (5)

    NameTypeDefaultDescription
    profileCOMBOCustom5 options: Custom, Gemma4 Medium, Gemma4 High, Qwen3.8 Medium, Qwen3.8 High
    n_batchINT5120–65536—
    n_ubatchINT00–655360 uses llama.cpp's default physical batch size.
    image_min_tokensINT00–65536—
    image_max_tokensINT00–655360 uses the mmproj or handler default.

    Outputs (1)

    NameTypeDescription
    prefill_profileOLLAMA_IMAGE_LIST_LLAMA_CPP_PREFILL_PROFILE—