ComfyUI Node
arkennemasis Local LLM (GGUF, llama.cpp)
Ask a local GGUF model (llama.cpp, GPU) and get text back. Runs in its own process so all of its VRAM is freed before the next node. Wire a JSON schema to force a JSON answer.
arkennemasis Local LLM (GGUF, llama.cpp)
- text
- report
◄model▾►
◄system►
◄prompt►
◄temperature0.70►
◄max_tokens4096►
◄seed0►
◄json_schema—►
◄context_tokens16384►
Categoryarkennemasis/LLM
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | Any .gguf under ComfyUI/models/LLM. Instruction-tuned models only; the chat template is read from the file. | |
| system | STRING | — | |
| prompt | STRING | — | |
| temperature | FLOAT | 0.700–2 | — |
| max_tokens | INT | 409664–32768 | — |
| seed | INT | 00–2147483647 | — |
| json_schemaopt | STRING | A JSON schema. When wired, decoding is constrained to it, so the answer is always parseable JSON of that shape. | |
| context_tokensopt | INT | 163842048–131072 | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| text | STRING | — |
| report | STRING | — |