
Description
Using several local models when VRAM fits only one means stopping, reconfiguring and restarting servers. llama-swap is a proxy providing reliable model swapping for local OpenAI-compatible servers.
It loads whichever model is requested, supports llama.cpp and vLLM and speaks OpenAI and Anthropic APIs.
Auto swap:On demand.
Backends:llama.cpp and vLLM.
APIs:OpenAI and Anthropic.
Single binary:Go.
It loads whichever model is requested, supports llama.cpp and vLLM and speaks OpenAI and Anthropic APIs.
Features
Auto swap:On demand.
Backends:llama.cpp and vLLM.
APIs:OpenAI and Anthropic.
Single binary:Go.

