llama-swap
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
About this project
llama-swap Run multiple generative AI models on your machine and hot-swap between them on demand. llama-swap works with any OpenAI and Anthropic API compatible server and is used by thousands of people to power their local AI workflows. Built in Go for performance and simplicity, llama-swap has zero dependencies and is incredibly easy to set up. Get started in minutes - just one binary and one configuration file. Features: — ✅ Easy to deploy and configure: one binary, one configuration file. no external dependencies — ✅ On-demand model switching for many local AI servers (llama.cpp + forks, vllm, stable-diffusion.cpp, audio.cpp, ComfyUI, etc.) — future proof, upgrade your inference…
Technologies
Project health
GitHub
Reviews
Built by
Maintain mostlygeek/llama-swap? Claiming verifies admin access through your GitHub account and gives you control of this listing.





