Unload drops the models from VRAM. Free runs collection and empties the CUDA cache. Both use ComfyUI's own endpoint. Restart needs ComfyUI-Manager and asks twice before it does anything.
Runs both of the above once the last job finishes, not between queued jobs, so a batch keeps its models warm and only the idle rig gives the memory back.
Vision Model
The VLM that runs the Tools. Any Qwen3-VL or Gemma instruct repack from your text_encoders folder. Leave unset to use whichever the node defaults to.
Holds the encoder in memory between calls. Much faster on a large model, at the cost of the VRAM it occupies.
Minutes to wait before giving up on a reply. Raise it if a big encoder is slow to load.
Lets the model reason before answering on text tasks. Slower, and only worth it on models trained for it.
Gallery
Seconds each image is held when ▸ auto-advance runs.
How far back ⧖ in the gallery reaches. Anything older is left alone.
Outputs
Consolidated outputs
Routes all output to output/<folder>/<model>/, overriding workflow paths.