
Description
LLMs feel like black boxes, and papers are too abstract to show how text becomes weights becomes output. GuppyLM is a roughly 9M parameter language model that pretends to be a small fish, trained from scratch in about five minutes on a single GPU in one Colab notebook.
Data generation, tokenizer, architecture, training loop and inference are all laid out. The model only chats in short lowercase sentences about water, food and tank life, but every step is clear.
Full pipeline:Synthetic data, tokenizer, training and inference.
Low barrier:Trains in about 5 minutes on one GPU.
Educational:Small code built to explain how LLMs work.
Data generation, tokenizer, architecture, training loop and inference are all laid out. The model only chats in short lowercase sentences about water, food and tank life, but every step is clear.
Features
Full pipeline:Synthetic data, tokenizer, training and inference.
Low barrier:Trains in about 5 minutes on one GPU.
Educational:Small code built to explain how LLMs work.
Screenshots
Tags:llm
