oMLX

oMLX

LLM inference server for Apple Silicon

Description

Serving local models to AI coding tools on a Mac queues concurrent requests and recomputes long contexts. oMLX is an LLM inference server for Apple Silicon managed from the menu bar.

Continuous batching and SSD caching boost concurrency and long-context speed behind an OpenAI-compatible API.

Features



Batching:Concurrent.

SSD cache:Long context.

Menu bar:Native.

OpenAI API:Drop-in.