llama-swap

llama-swap

Reliable model swapping for local LLM servers

Description

Using several local models when VRAM fits only one means stopping, reconfiguring and restarting servers. llama-swap is a proxy providing reliable model swapping for local OpenAI-compatible servers.

It loads whichever model is requested, supports llama.cpp and vLLM and speaks OpenAI and Anthropic APIs.

Features



Auto swap:On demand.

Backends:llama.cpp and vLLM.

APIs:OpenAI and Anthropic.

Single binary:Go.