Research Library

-- FLOPs

Curated research curriculum

The papers behind
modern intelligence.

Study the field in the order its ideas emerged. Each paper includes guided notes, focused learning modules, and questions that test real understanding.

70Guided papers
180Learning modules
2017 — 2026Research span

Read in sequence.

DATE
012017V100Research paper · 1 learning moduleHardware
022017RLHFDeep reinforcement learning from human preferences · 1 learning moduleAlignment
032017TransformerAttention Is All You Need · 4 learning modulesArchitecture
042017PPOProximal Policy Optimization Algorithms · 1 learning moduleAlignment
052017QATQuantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference · 1 learning moduleQuantization
062018Top-kHierarchical Neural Story Generation · 1 learning moduleInference
072018GPT-1Improving Language Understanding by Generative Pre-Training · 7 learning modulesLanguage ModelingFine-Tuning
082018Tesla T4NVIDIA · 1 learning moduleHardware
092018BERTBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding · 7 learning modulesLanguage ModelingFine-Tuning
102018GPipeGPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · 1 learning moduleSystems
112018Hugging FaceHugging Face · 1 learning moduleTooling
122019GPT-2Language Models are Unsupervised Multitask Learners · 5 learning modulesLanguage Modeling
132019Top-pThe Curious Case of Neural Text Degeneration · 1 learning moduleInference
142019Tensor ParallelMegatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism · 1 learning moduleSystems
152019Data ParallelZeRO: Memory Optimizations Toward Training Trillion Parameter Models · 1 learning moduleSystems
162019T5Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer · 7 learning modulesLanguage Modeling
172019MQAFast Transformer Decoding: One Write-Head is All You Need · 1 learning moduleArchitectureInference
182020Scaling LawScaling Laws for Neural Language Models · 1 learning moduleScalingLanguage Modeling
192020A100Research paper · 1 learning moduleHardware
202020GPT-3Language Models are Few-Shot Learners · 5 learning modulesScalingLanguage Modeling
212020GShardGShard: Scaling Giant Models with Conditional Computation and Automatic Sharding · 1 learning moduleSystemsArchitecture
222020ViTAn Image is Worth 16x16 Words: Transformers for Image Recognition at Scale · 3 learning modulesArchitectureVision
232021Switch TransformerSwitch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · 1 learning moduleArchitectureScaling
242021ZeRO-OffloadZeRO-Offload: Democratizing Billion-Scale Model Training · 1 learning moduleSystems
252021ViLTViLT: Vision-and-Language Transformer Without Convolution or Region Supervision · 3 learning modulesVisionMultimodal
262021DALL·E 1Zero-Shot Text-to-Image Generation · 6 learning modulesMultimodalImage Generation
272021CLIPLearning Transferable Visual Models From Natural Language Supervision · 3 learning modulesVisionMultimodal
282021Megatron-LM: 3D ParallelismEfficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM · 1 learning moduleSystems
292021ZeRO-InfinityZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning · 1 learning moduleSystems
302021LoRALoRA: Low-Rank Adaptation of Large Language Models · 1 learning moduleFine-Tuning
312021CodeXEvaluating Large Language Models Trained on Code · 4 learning modulesLanguage ModelingCode
322021GLaMGLaM: Efficient Scaling of Language Models with Mixture-of-Experts · 1 learning moduleArchitectureScaling
332021Stable DiffusionHigh-Resolution Image Synthesis with Latent Diffusion Models · 2 learning modulesMultimodalImage Generation
342022AlphaCodeCompetition-Level Code Generation with AlphaCode · 5 learning modulesReasoningCode
352022COTChain-of-Thought Prompting Elicits Reasoning in Large Language Models · 1 learning moduleInferenceReasoning
362022InstructGPTTraining language models to follow instructions with human feedback · 5 learning modulesAlignmentFine-Tuning
372022H100Research paper · 1 learning moduleHardware
382022ChinchillaTraining Compute-Optimal Large Language Models · 1 learning moduleScalingLanguage Modeling
392022DALL·E 2Hierarchical Text-Conditional Image Generation with CLIP Latents · 3 learning modulesMultimodalImage Generation
402022FlashAttentionFlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness · 1 learning moduleSystemsInference
412022LLM.int8()LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale · 1 learning moduleQuantizationInference
422022GPTQGPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers · 1 learning moduleQuantizationInference
432022WhisperRobust Speech Recognition via Large-Scale Weak Supervision · 2 learning modulesSpeech
442023LLaMA-1LLaMA: Open and Efficient Foundation Language Models · 5 learning modulesScalingLanguage Modeling
452023LLaVAVisual Instruction Tuning · 2 learning modulesVisionMultimodal
462023DPODirect Preference Optimization: Your Language Model is Secretly a Reward Model · 1 learning moduleAlignmentFine-Tuning
472023QLoRAQLoRA: Efficient Finetuning of Quantized LLMs · 1 learning moduleFine-TuningQuantization
482023AWQAWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration · 1 learning moduleQuantizationInference
492023FlashAttention-2FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning · 1 learning moduleSystemsInference
502023LLaMA-2Llama 2: Open Foundation and Fine-Tuned Chat Models · 7 learning modulesLanguage ModelingAlignment
512023Qwen-VLQwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond · 3 learning modulesVisionMultimodal
522023Qwen-1Qwen Technical Report · 7 learning modulesLanguage Modeling
532023Mistral 7BMistral 7B · 5 learning modulesArchitectureInference
542023H200NVIDIA H200 Tensor Core GPU Datasheet · 1 learning moduleHardware
552023LVMSequential Modeling Enables Scalable Learning for Large Vision Models · 3 learning modulesArchitectureVision
562024KTOKTO: Model Alignment as Prospect Theoretic Optimization · 1 learning moduleAlignmentFine-Tuning
572024Mixtral 8x7BMixtral of Experts · 2 learning modulesArchitectureScaling
582024ORPOORPO: Monolithic Preference Optimization without Reference Model · 1 learning moduleAlignmentFine-Tuning
592024Gemma 1Gemma: Open Models Based on Gemini Research and Technology · 7 learning modulesLanguage Modeling
602024DeepSeek-V2DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model · 2 learning modulesArchitectureInference
612024SimPOSimPO: Simple Preference Optimization with a Reference-Free Reward · 1 learning moduleAlignmentFine-Tuning
622024B200Research paper · 1 learning moduleHardware
632024ChatGLMChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools · 3 learning modulesLanguage Modeling
642024Llama 3The Llama 3 Herd of Models · 6 learning modulesScalingLanguage Modeling
652024Gemma 2Gemma 2: Improving Open Language Models at a Practical Size · 6 learning modulesArchitectureLanguage Modeling
662024DeepSeek-V3DeepSeek-V3 Technical Report · 4 learning modulesSystemsArchitecture
672025DeepSeek-R1DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · 3 learning modulesLanguage ModelingReasoning
682025B300NVIDIA GTC 2025 Keynote & Blackwell Ultra Architecture Brief · 1 learning moduleHardware
692025Gemma 3Gemma 3 Technical Report · 6 learning modulesLanguage ModelingMultimodal
702026Rubin GPUNVIDIA Vera Rubin NVL72 Product Page · 1 learning moduleHardware

One library. Every guided paper.

Unlock the complete reading path, every learning module, personal notes, quizzes, and future paper updates.

See full access
LLM Research Library & Guided Paper Path - LearnLLM.AI