Train LLM From Scratch

Train LLM From Scratch

Train your own LLM from scratch on one GPU

Description

After many articles on LLM training, it is still unclear how each step from downloading data to generating text is written. This project hand-writes the whole pipeline from raw text to an aligned reasoning model in plain PyTorch, training million to billion parameter models on a single GPU.

It implements the Transformer from Attention is All You Need and continues through alignment and reasoning-style training, with every algorithm written by hand.

Features



Full pipeline:Data, tokenizer, pretraining and alignment.

Single GPU:Million to billion parameter models.

Hand-written:No high-level training frameworks.

Reasoning stage:Alignment and reasoning training.