Nodes/comfyui-arkennemasis/arkennemasis Local LLM (GGUF, llama.cpp)
ComfyUI Node

arkennemasis Local LLM (GGUF, llama.cpp)

Ask a local GGUF model (llama.cpp, GPU) and get text back. Runs in its own process so all of its VRAM is freed before the next node. Wire a JSON schema to force a JSON answer.

By Hishamahmer·Created 2 months ago·Updated 4 days ago· 10
arkennemasis Local LLM (GGUF, llama.cpp)
    • text
    • report
    ◄model▾►
    ◄system►
    ◄prompt►
    ◄temperature0.70►
    ◄max_tokens4096►
    ◄seed0►
    ◄json_schema—►
    ◄context_tokens16384►
    Categoryarkennemasis/LLM

    Inputs (8)

    NameTypeDefaultDescription
    modelCOMBOAny .gguf under ComfyUI/models/LLM. Instruction-tuned models only; the chat template is read from the file.
    systemSTRING—
    promptSTRING—
    temperatureFLOAT0.700–2—
    max_tokensINT409664–32768—
    seedINT00–2147483647—
    json_schemaoptSTRINGA JSON schema. When wired, decoding is constrained to it, so the answer is always parseable JSON of that shape.
    context_tokensoptINT163842048–131072—

    Outputs (2)

    NameTypeDescription
    textSTRING—
    reportSTRING—