
Description
Serving local models to AI coding tools on a Mac queues concurrent requests and recomputes long contexts. oMLX is an LLM inference server for Apple Silicon managed from the menu bar.
Continuous batching and SSD caching boost concurrency and long-context speed behind an OpenAI-compatible API.
Batching:Concurrent.
SSD cache:Long context.
Menu bar:Native.
OpenAI API:Drop-in.
Continuous batching and SSD caching boost concurrency and long-context speed behind an OpenAI-compatible API.
Features
Batching:Concurrent.
SSD cache:Long context.
Menu bar:Native.
OpenAI API:Drop-in.

