LLM Lifecycle: LM Studio
Make LM Studio let go of VRAM when you're done
- lifecycle
LLM Lifecycle: LM Studio is the pack's answer to a specific annoyance: LM Studio keeps whatever model you loaded resident on your GPU until you tell it otherwise. If your card is also trying to run a diffusion model, that's a permanently split VRAM budget - your LLM hoards memory while the sampler sits idle, and your sampler evicts the LLM the moment you actually use it. This node exists to make the LLM give the memory back.
The mechanism is a time-to-live, not an unload command. This lifecycle node produces an LLM_LIFECYCLE object that you connect to the lifecycle input on LLM Provider: OAI Compatible (it's specifically for the LM Studio backend that provider detected). When the lifecycle is present, the adapter manages the model's residency against the TTL - after the idle window elapses with no generation, the model is left to unload. It's the same VRAM-aware instinct the KB's LLM essay describes (llm-in-comfyui.md: two models on one card is the failure mode, so good nodes automate unload/reload instead of holding everything resident) - but for LM Studio the lever is idle-timeout rather than explicit load/unload hooks, because LM Studio is a GUI app you control externally, not a headless server.
The two widgets
ttl- the idle timeout in seconds, default 30, minimum 0. 0 means unload immediately; larger values keep the model warm between generations. If you're chaining several generations back to back, the chain-aware deferral in the generation nodes already keeps things loaded across the chain, so a modest TTL is fine.context_length- an optional override for the model's context window, default 0. Zero means "don't override - use whatever the model/server wants." Max is 1048576. You'd set this to shrink the context if you're fighting for VRAM, since a huge context reservation can be the difference between fitting and not.
Both are simple integers because that's genuinely all this node needs to say. The output is a single lifecycle socket of type LLM_LIFECYCLE - connect it to the provider, that's the whole wiring.
When you'd bother
Honestly, only if you share the GPU. If the LLM runs on the same card as the sampler and you do a lot of toggling between "write a prompt with the LLM" and "actually render," this keeps your diffusion runs from being throttled by a resident model. If you have VRAM to burn or a separate card for the LLM, the node is dead weight - skip it. The default 30s TTL is a reasonable middle ground: long enough that a single generation's output isn't instantly evicted, short enough that a render session gets its VRAM back.
One sharp edge worth knowing: this lifecycle only does anything when the backend at your provider's URL is actually LM Studio. Detection says so - the OAI Compatible provider fingerprints the endpoint - but if you've pointed it at a generic OpenAI-compatible server, a connected LM Studio lifecycle is simply ignored.
Install
Pack standard: ComfyUI Manager search "comfyui-llm-bikeshed", or clone into custom_nodes/ and pip install -r requirements.txt (only pyyaml and requests, both already in ComfyUI core). No model downloads, no extra config needed for this node beyond having LM Studio running.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| ttl | INT | 30 | Time-to-live in seconds — how long LM Studio keeps the model loaded after each request. Timer resets on each request. 0 = unload immediately. |
| context_length | INT | 00–1048576 | Context window for explicit model load via LM Studio REST API. 0 = use the model default. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| lifecycle | LLM_LIFECYCLE | Connect to the lifecycle input on LLM Provider: OAI Compatible when the detected backend is LM Studio. |