Kimi K2.6
Kimi K2.6 — a 1M-token-reasoning model as a ComfyUI node
- messages
- images
- text
Kimi is Moonshot AI's model line, and the K2 series has a reputation for strong open-weights reasoning - the kind of model people use for the thinking-heavy jobs that small local LLMs embarrass themselves on. K2.6 is the current tier, and this node puts it inside ComfyUI as a text node with a serious spec: an up-to-1,000,000-token output ceiling via settings.maxTokens_value, five levels of settings.thinkingLevel, tool calling, JSON schema output, and vision via an IMAGE input.
Why would a ComfyUI user want this? Same reason you'd want any good LLM node: prompt rewriting, captioning, structured JSON generation for downstream nodes, and decision logic in multi-step workflows. K2.6's edge is the combination of a huge context, cheap reasoning levels, and JSON output - which makes it genuinely useful as a workflow brain that hands structured data to the rest of the graph.
How it works
The standard pack text-node path: required messages socket (Runware Messages builder: role + content, stackable), a textInference request, and a STRING text output. Where K2.6 stands out is the control surface - it exposes more thinking and output control than the other text nodes in the pack.
Inputs that matter
messages(required) +settings.systemPrompt- conversation and system instruction. For structured-output workflows, put the format spec in the system prompt and useoutputFormat: JSON.settings.thinkingLevel-none,low,medium,high,xhigh(defaultnone). This is the knob that makes K2.6 interesting: it's a reasoning model, but unlike some peers it lets you dial reasoning from off to extreme. For pipeline tasks,none/lowis fast and cheap; for hard analysis,high/xhigh. Defaulting tononeis a nice "don't pay for thinking you didn't ask for" default.outputFormat-TEXTorJSON. JSON mode is the underrated one: combined with theadvanced_jsonjsonSchemafield, you can get reliably-structured output out of a text node - hugely useful when the result feeds other nodes.settings.maxTokens- off-by-default gate withmaxTokens_value(up to 1,000,000). That ceiling is enormous; leave the gate off for the model default, and only flip it on if you're doing something genuinely long. One careless 1M-token call is a bill, not a result.images- anIMAGEinput, so K2.6 is vision-capable: wire a rendered image in and ask for analysis, captioning, or critique.settings.promptCacheKey- a cache key for reusing prompt processing across requests. If you're running the same big system prompt repeatedly (batch captioning, same-context loop), set a key and the repeated prefix gets cached - a real cost saver for cloud LLM use.toolChoice/toolChoice.type/toolChoice.name- tool calling viaadvanced_json(tools).auto/any/tool/none.settings.temperature(default 1),settings.topK(default 0),settings.topP(default 1),minP, the three penalties - full sampler set. For deterministic pipeline work, drop temperature to ~0.2.numberResults(1–4),seed,includeUsage- variations and spend visibility.
Single output: text (a STRING - if you chose JSON mode, it's a JSON string; parse it in a downstream node).
Installing
cd ComfyUI/custom_nodes
git clone https://github.com/Runware/ComfyUI-Runware
pip install -r ComfyUI-Runware/requirements.txt
Or ComfyUI Manager → search Runware → install → restart. API key via Settings, RUNWARE_API_KEY, or runware auth login.
The honest caveats
The 1M-token output ceiling is real but you will almost never touch it - and the cost scales with what you use, so default to small. The advanced_json fields (jsonSchema, tools, settings.stopSequences) are where half the power lives, and they're easy to miss if you only look at the visible widgets. And remember the output is a plain string: wiring K2.6's answer into a workflow that needs numbers or booleans means you're parsing text (or using JSON mode) yourself. For "big-brain reasoning that hands structured answers to my graph," this is one of the best values in the pack.
Inputs (22)
| Name | Type | Default | Description |
|---|---|---|---|
| messages | RUNWARE_MESSAGES | — | |
| imagesopt | IMAGE | — | |
| seedopt | INT | 00–9223372036854776000 | Random seed for reproducible generation. When not provided, a random seed is generated in the unsigned 32-bit range. |
| numberResultsopt | INT | 11–4 | Number of results to generate. Each result uses a different seed, producing variations of the same parameters. |
| includeUsageopt | BOOLEAN | false | Include token usage statistics in the response. |
| settings.frequencyPenaltyopt | FLOAT | 0.00-2–2 | Penalizes tokens based on their frequency in the output so far. A value of 0.0 disables the penalty. |
| settings.maxTokensopt | BOOLEAN | false | Enable to set settings.maxTokens. Off uses the model's default. |
| settings.maxTokens_valueopt | INT | 11–1000000 | Maximum number of tokens to generate in the response. |
| settings.minPopt | FLOAT | 0.000–1 | Minimum probability threshold. Tokens with probability below this value are excluded from sampling. |
| settings.presencePenaltyopt | FLOAT | 0.00-2–2 | Encourages the model to introduce new topics. A value of 0.0 disables the penalty. |
| settings.promptCacheKeyopt | STRING | Cache key for reusing prompt cache across requests. Requests sharing the same key reuse cached prompt processing. | |
| settings.repetitionPenaltyopt | FLOAT | 1.000–5 | Penalizes tokens that have already appeared in the output. A value of 1.0 disables the penalty. |
| settings.systemPromptopt | STRING | System-level instruction that guides the model's behavior and output style across the entire generation. | |
| settings.temperatureopt | FLOAT | 1.000–2 | Controls randomness in generation. Lower values produce more deterministic outputs, higher values increase variation and creativity. |
| settings.thinkingLevelopt | COMBO | none | Controls the depth of internal reasoning the model performs before generating a response. |
| toolChoiceopt | BOOLEAN | false | Enable to set toolChoice. Off uses the model's default. |
| toolChoice.nameopt | STRING | Name of the specific tool the model must call. Required when type is `tool`. | |
| toolChoice.typeopt | COMBO | (default) | Strategy the model uses to decide when and which tools to call. |
| settings.topKopt | INT | 00–999 | Top-K sampling parameter that limits the number of highest-probability tokens considered at each step. |
| settings.topPopt | FLOAT | 1.000–1 | Nucleus sampling parameter that controls diversity by limiting the probability mass. Lower values make outputs more focused, higher values increase diversity. |
| outputFormatopt | COMBO | TEXT | Output format for the generated text. |
| advanced_jsonopt | STRING | Optional JSON merged into the request. For: settings.stopSequences, jsonSchema, tools |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| text | STRING | — |