
Description
LLMs usually need GPUs, and running one on a microcontroller sounds impossible. This project runs a 28.9M-parameter language model on an ESP32-S3 microcontroller, entirely on the chip with no server.
It generates 9.88 tokens per second and shows text on a display, pushing the limits of on-device AI.
On-chip:No server needed.
28.9M parameters:On an MCU.
Live display:Text on screen.
It generates 9.88 tokens per second and shows text on a display, pushing the limits of on-device AI.
Features
On-chip:No server needed.
28.9M parameters:On an MCU.
Live display:Text on screen.
