Transformer
Introduced the Transformer, replacing recurrence with self-attention and enabling highly parallel training. Its architecture became the foundation of BERT, GPT, and nearly every modern language model.
Study this paperAn open curriculum for modern AI
A clear, guided path through the research that built large language models — from attention and scaling to alignment, retrieval, and agents.
THE FIELD, IN SEQUENCE
Attention · Transformers · Pre-training
Scaling laws · Instruction tuning · RLHF
Retrieval · Multimodality · Agents
SELECTED PAPERS
Introduced the Transformer, replacing recurrence with self-attention and enabling highly parallel training. Its architecture became the foundation of BERT, GPT, and nearly every modern language model.
Study this paperHOW IT WORKS
Understand what problem the paper set out to solve.
Work through the architecture, equations, and key decisions.
Check your understanding with focused questions and walkthroughs.
YOUR NEXT PAPER
No noise. No endless feed. Just the research, explained in the right order.
Begin the learning path