
Description
Open-weight models like DeepSeek, Qwen and GLM run to hundreds of GB, too big or too slow locally. FreeToken brings datacenter-scale model serving to your desktop, optimized for mixture-of-experts models on consumer hardware.
It loads experts on demand, manages KV and expert caches across VRAM and RAM, and offers a desktop console with usage stats and a local API.
Big models:DeepSeek, Qwen, GLM.
MoE:On-demand experts.
Caches:KV and expert.
Console:Status and usage.
Local API:For other apps.
It loads experts on demand, manages KV and expert caches across VRAM and RAM, and offers a desktop console with usage stats and a local API.
Features
Big models:DeepSeek, Qwen, GLM.
MoE:On-demand experts.
Caches:KV and expert.
Console:Status and usage.
Local API:For other apps.
