
Description
Running LLMs on phones and embedded devices with typical frameworks is bulky, slow and power-hungry. MNN is Alibaba's battle-tested, lightweight, blazing-fast inference engine for on-device LLMs and edge AI.
It supports ARM, Vulkan and more, with an Android app for running Qwen and other models offline.
Fast and light:On-device optimized.
Backends:ARM, Vulkan and GPU.
On-device LLMs:Offline on phones.
Proven:Used at Taobao scale.
It supports ARM, Vulkan and more, with an Android app for running Qwen and other models offline.
Features
Fast and light:On-device optimized.
Backends:ARM, Vulkan and GPU.
On-device LLMs:Offline on phones.
Proven:Used at Taobao scale.

