
Description
Running local models on a Mac and letting every tool call them requires an OpenAI-compatible server. vMLX is a self-hosted inference server for Apple Silicon for LLMs, VLMs and image generation.
It exposes OpenAI, Anthropic and Ollama compatible HTTP APIs, supports KV cache compression and reuse, prefix caching and the GGUF-like JANGQ format, with no third-party keys.
Compatible APIs:OpenAI, Anthropic and Ollama style.
Multimodal:Text, vision and image generation.
Cache tuning:KV compression, reuse and prefix cache.
MLX-native:Built for Apple Silicon.
It exposes OpenAI, Anthropic and Ollama compatible HTTP APIs, supports KV cache compression and reuse, prefix caching and the GGUF-like JANGQ format, with no third-party keys.
Features
Compatible APIs:OpenAI, Anthropic and Ollama style.
Multimodal:Text, vision and image generation.
Cache tuning:KV compression, reuse and prefix cache.
MLX-native:Built for Apple Silicon.

