
Description
Running a model with hundreds of billions of parameters locally usually takes a rack of GPUs. Flash-MoE is a pure C and Metal inference engine that runs the 397B-parameter Qwen3.5-397B-A17B on a MacBook Pro with 48GB of RAM at 4.4+ tokens per second, with tool calling.
The 209GB model streams from SSD through a custom Metal pipeline with no Python or frameworks, and the author published a paper covering 90+ experiments.
Huge model:A 397B MoE on a laptop.
SSD streaming:Expert weights read on demand.
Pure C and Metal:No Python or frameworks.
Tool calling:Production-quality output.
The 209GB model streams from SSD through a custom Metal pipeline with no Python or frameworks, and the author published a paper covering 90+ experiments.
Features
Huge model:A 397B MoE on a laptop.
SSD streaming:Expert weights read on demand.
Pure C and Metal:No Python or frameworks.
Tool calling:Production-quality output.
Screenshots
Tags:local-llm
