vMLX

vMLX

MLX inference server for Apple Silicon

Description

Running local models on a Mac and letting every tool call them requires an OpenAI-compatible server. vMLX is a self-hosted inference server for Apple Silicon for LLMs, VLMs and image generation.

It exposes OpenAI, Anthropic and Ollama compatible HTTP APIs, supports KV cache compression and reuse, prefix caching and the GGUF-like JANGQ format, with no third-party keys.

Features



Compatible APIs:OpenAI, Anthropic and Ollama style.

Multimodal:Text, vision and image generation.

Cache tuning:KV compression, reuse and prefix cache.

MLX-native:Built for Apple Silicon.