Decoder-only Transformer • PyTorch • From Scratch
LLM from Scratch
Build, train, benchmark, and fine-tune a modern language model from first principles.
Architecture
Tokenizer, transformer blocks, training loop, kernels, and distributed training.
SFT & RLFT
Qwen2.5-Math-1.5B fine-tuning story with concrete results on gsm8k.
Benchmarks
Loss curves, learning-rate schedule, and performance context.
Docs
Structured markdown chapters for tokenizer, transformer internals, training, scaling, and alignment.