参考
How to Scale Your Model:JAX Scaling Book
Transformer training memory:Transformer memory usage overview
Transformer FLOPs:Transformer FLOPs accounting overview
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
UvA DL Notebooks: Profiling and Scaling Single-GPU Transformer Models
UvA DL Notebooks: Part 2.2: (Fully-Sharded) Data Parallelism
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Accelerating Large Language Model Decoding with Speculative Sampling
Efficient Memory Management for Large Language Model Serving with PagedAttention
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Establishing Best Practices for Building Rigorous Agentic Benchmarks
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming
AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
DataComp-LM: In Search of the Next Generation of Training Sets for Language Models
OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
UniMax: Fairer and More Effective Language Sampling for Large-Scale Multilingual Pretraining
RegMix: Data Mixture as Regression for Language Model Pre-training
Training Language Models to Follow Instructions with Human Feedback
Self-Instruct: Aligning Language Models with Self-Generated Instructions
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Training Language Models to Follow Instructions with Human Feedback
UltraFeedback: Boosting Language Models with Scaled AI Feedback
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
SimPO: Simple Preference Optimization with a Reference-Free Reward
Tülu 3: Pushing Frontiers in Open Language Model Post-Training
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Stanford CS336 Spring 2026, Lecture 16: Reinforcement Learning from Verifiable Rewards
High-Dimensional Continuous Control Using Generalized Advantage Estimation
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Learning Transferable Visual Models From Natural Language Supervision(CLIP)
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution