MNN

MNN

Lightweight on-device inference engine

Description

Running LLMs on phones and embedded devices with typical frameworks is bulky, slow and power-hungry. MNN is Alibaba's battle-tested, lightweight, blazing-fast inference engine for on-device LLMs and edge AI.

It supports ARM, Vulkan and more, with an Android app for running Qwen and other models offline.

Features



Fast and light:On-device optimized.

Backends:ARM, Vulkan and GPU.

On-device LLMs:Offline on phones.

Proven:Used at Taobao scale.