Nodes/Kinburg-Nodes/Local LLM Chat (GGUF)
ComfyUI Node

Local LLM Chat (GGUF)

A ComfyUI node in Kinburg-Nodes/LLM with 9 inputs and 2 outputs.

By Kinburg·Created 2 months ago·Updated 3 days ago· 1
Local LLM Chat (GGUF)
  • persona_1
  • image
  • persona_2
  • persona_3
  • persona_4
  • persona_5
  • persona_6
  • text
  • help
unload_on_approvetrue
chat_state
CategoryKinburg-Nodes/LLM

Inputs (9)

NameTypeDefaultDescription
persona_1KINBURG_LLM_CONFIGPersona #1: wire a whole 'Local LLM Settings (GGUF)' here — its system prompt, model and sampling become the config for every turn this persona speaks. A chip for it appears in the chat as soon as a second persona is wired. Keep the loader fields (model, n_ctx, n_gpu_layers, flash_attn, kv_cache_type) identical across personas and switching between them costs no model reload. This one is also the node's default config.
imageoptIMAGEOptional image for vision, attached to every turn while it stays connected. Prefer pasting or dropping a picture straight into the chat window instead — that attaches it to ONE turn and doesn't re-run this branch on every message. Needs an mmproj on the active Settings node either way.
persona_2optKINBURG_LLM_CONFIGPersona #2: wire a whole 'Local LLM Settings (GGUF)' here — its system prompt, model and sampling become the config for every turn this persona speaks. A chip for it appears in the chat as soon as a second persona is wired. Keep the loader fields (model, n_ctx, n_gpu_layers, flash_attn, kv_cache_type) identical across personas and switching between them costs no model reload.
persona_3optKINBURG_LLM_CONFIGPersona #3: wire a whole 'Local LLM Settings (GGUF)' here — its system prompt, model and sampling become the config for every turn this persona speaks. A chip for it appears in the chat as soon as a second persona is wired. Keep the loader fields (model, n_ctx, n_gpu_layers, flash_attn, kv_cache_type) identical across personas and switching between them costs no model reload.
persona_4optKINBURG_LLM_CONFIGPersona #4: wire a whole 'Local LLM Settings (GGUF)' here — its system prompt, model and sampling become the config for every turn this persona speaks. A chip for it appears in the chat as soon as a second persona is wired. Keep the loader fields (model, n_ctx, n_gpu_layers, flash_attn, kv_cache_type) identical across personas and switching between them costs no model reload.
persona_5optKINBURG_LLM_CONFIGPersona #5: wire a whole 'Local LLM Settings (GGUF)' here — its system prompt, model and sampling become the config for every turn this persona speaks. A chip for it appears in the chat as soon as a second persona is wired. Keep the loader fields (model, n_ctx, n_gpu_layers, flash_attn, kv_cache_type) identical across personas and switching between them costs no model reload.
persona_6optKINBURG_LLM_CONFIGPersona #6: wire a whole 'Local LLM Settings (GGUF)' here — its system prompt, model and sampling become the config for every turn this persona speaks. A chip for it appears in the chat as soon as a second persona is wired. Keep the loader fields (model, n_ctx, n_gpu_layers, flash_attn, kv_cache_type) identical across personas and switching between them costs no model reload.
unload_on_approveoptBOOLEANtrueFree the LLM from VRAM when you press ✅ Approve, so the image model downstream has room. Turn off to keep chatting at full speed when nothing heavy follows.
chat_stateoptSTRINGThe conversation, managed by the chat window. You don't edit this directly.

Outputs (2)

NameTypeDescription
textSTRING
helpSTRING