[llama.cpp] Native Speculative Profile
Speculative Decoding as a Typed Socket (Mostly Leave It Off)
- speculative
Speculative decoding, wrapped in a node so you don't have to learn llama.cpp's flags. A small draft model proposes several tokens ahead, the real model checks them in one pass, and you keep the ones it agrees with. When the draft is well matched, decode gets faster; when it isn't, you've added a second model's worth of memory for a speedup that never shows up.
You plug one output - speculative - into the speculative input on [llama.cpp] Create Native Session. That's the entire interface.
The presets, and what they map to
The preset combo is the part you actually choose. Each preset sets a spec type and an MTP provider behind the scenes:
- Off - the default, and it means it. Whatever else you've set, this node emits a disabled config and the session loads plain.
- External MTP - multi-token prediction with an external assistant GGUF. This one does use
draft_model. - Internal MTP - MTP where the target model carries the prediction layers itself. Per the tooltip on
draft_model, Internal MTP doesn't use a draft GGUF at all. - DFlash and DSpark - the two draft-model paths, both of which use
draft_model. - Custom - exposes the raw pair:
custom_spec_type(none,draft-dflash,draft-dspark,draft-mtp) andcustom_mtp_provider(off,external,internal). Picknonehere and you're back to Off.
draft_model is a combo of the GGUFs in ComfyUI/models/LLM. The knobs are draft_n_max (default 2, up to 64) for how many tokens the draft proposes, draft_p_min for the acceptance floor, draft_n_gpu_layers (auto or all), and draft_backend_sampling. Presets fill in the strategy; these widgets still apply on top, so a preset isn't a sealed box.
Why you probably shouldn't touch it yet
This is the most experimental corner of the pack, and the honest read is: it exists for people running exotic builds on purpose.
Native MTP in particular needs an experimental llama-cpp-python build that exposes speculative ABI v2 and MTP bridging - the earlier DFlash/DSpark-only wheel isn't enough. It's text-only, wants all layers on the GPU, and disables context shifting, grammar constraints, custom logits processors, state-cache reuse and multi-sequence batching while it's active. Your draft model also has to genuinely match, and matching means specific pairings: an assistant GGUF for the Gemma 4 external path, a target with embedded NextN layers for Qwen 3.5 internal MTP.
Get any of that wrong and the usual outcome isn't a crash - it's a slower run with the same output quality, or output that quietly degrades. If you're not swapping GGUFs and measuring tokens per second already, set preset to Off and spend your afternoon on n_ctx and your prompt instead.
There's also a model-free alternative in the pack: a deprecated N-gram speculative config node that drafts from prompt history with no second model at all. It's the cheaper experiment if you just want to see whether drafting helps your workload - the native node here supersedes it.
Installing and setup
Nothing special for this node, but the pack's moving parts are the usual suspects. Manager → search llama multimodal → install ComfyUI-llama-multimodal, or:
cd ComfyUI/custom_nodes
git clone https://github.com/craftingmod/ComfyUI-llama-multimodal.git
Restart. ComfyUI 0.19.3+ required, or none of these nodes register. Put both the model GGUF and any draft GGUF in ComfyUI/models/LLM/ and refresh the node list - a combo reading [no GGUF models found] means the folder's empty as far as ComfyUI is concerned, and a missing draft model is the most common reason a preset produces nothing.
When it actually breaks
If the session fails at load with a message about speculative bindings rather than about your model, your llama-cpp-python build lacks the required API - that's a wheel problem, and it's named in the error with a pointer to the prebuilt releases. If it loads but generation is slower, your draft model is mismatched or draft_n_max is set too aggressively for how often the target accepts the proposals; drop it back toward 2 and retest. And if you're mixing this with vision, remember the text-only constraint: speculative decoding here doesn't buy you anything on the image prefill, which is where your time actually goes on a media-heavy graph.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| preset | COMBO | Off | 6 options: Off, External MTP, Internal MTP, DFlash, DSpark, Custom |
| draft_model | COMBO | [no GGUF models found] | Used by External MTP, DFlash, and DSpark. Internal MTP does not use a draft GGUF. |
| draft_n_max | INT | 21–64 | — |
| custom_spec_type | COMBO | draft-dflash | 4 options: none, draft-dflash, draft-dspark, draft-mtp |
| custom_mtp_provider | COMBO | off | 3 options: off, external, internal |
| draft_p_min | FLOAT | 0.000–1 | — |
| draft_n_gpu_layers | COMBO | all | 2 options: auto, all |
| draft_backend_sampling | BOOLEAN | true | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| speculative | OLLAMA_IMAGE_LIST_LLAMA_CPP_SPECULATIVE_CONFIG | — |