Flash-MoE

Flash-MoE

Run a 397B model on a laptop

Description

Running a model with hundreds of billions of parameters locally usually takes a rack of GPUs. Flash-MoE is a pure C and Metal inference engine that runs the 397B-parameter Qwen3.5-397B-A17B on a MacBook Pro with 48GB of RAM at 4.4+ tokens per second, with tool calling.

The 209GB model streams from SSD through a custom Metal pipeline with no Python or frameworks, and the author published a paper covering 90+ experiments.

Features



Huge model:A 397B MoE on a laptop.

SSD streaming:Expert weights read on demand.

Pure C and Metal:No Python or frameworks.

Tool calling:Production-quality output.