
Description
After many articles on LLM training, it is still unclear how each step from downloading data to generating text is written. This project hand-writes the whole pipeline from raw text to an aligned reasoning model in plain PyTorch, training million to billion parameter models on a single GPU.
It implements the Transformer from Attention is All You Need and continues through alignment and reasoning-style training, with every algorithm written by hand.
Full pipeline:Data, tokenizer, pretraining and alignment.
Single GPU:Million to billion parameter models.
Hand-written:No high-level training frameworks.
Reasoning stage:Alignment and reasoning training.
It implements the Transformer from Attention is All You Need and continues through alignment and reasoning-style training, with every algorithm written by hand.
Features
Full pipeline:Data, tokenizer, pretraining and alignment.
Single GPU:Million to billion parameter models.
Hand-written:No high-level training frameworks.
Reasoning stage:Alignment and reasoning training.
