
Description
Running LLMs on phones, watches or robots with ordinary frameworks is bulky, slow and battery-draining. Cactus is an on-device AI inference engine for mobiles, wearables, smart homes and robots with quantization, kernels and a runtime.
It supports LLMs, speech recognition and RAG with Android, iOS, Flutter and React Native bindings.
On-device:Offline.
Quantization:Memory and power saved.
Multimodal:Text and speech.
Frameworks:Flutter and RN.
It supports LLMs, speech recognition and RAG with Android, iOS, Flutter and React Native bindings.
Features
On-device:Offline.
Quantization:Memory and power saved.
Multimodal:Text and speech.
Frameworks:Flutter and RN.
