FreeToken

FreeToken

Run massive MoE models locally on consumer hardware

Description

Open-weight models like DeepSeek, Qwen and GLM run to hundreds of GB, too big or too slow locally. FreeToken brings datacenter-scale model serving to your desktop, optimized for mixture-of-experts models on consumer hardware.

It loads experts on demand, manages KV and expert caches across VRAM and RAM, and offers a desktop console with usage stats and a local API.

Features



Big models:DeepSeek, Qwen, GLM.

MoE:On-demand experts.

Caches:KV and expert.

Console:Status and usage.

Local API:For other apps.