The Art of Scaling Reinforcement Learning Compute for LLMs
Meta AI and collaborators establish the first systematic scaling laws for reinforcement learning in large language models, based on over 400,000 GPU-hours of experiments.
Authors: Khatri et al. (Meta, UT Austin, UCL, UC Berkeley, Harvard, Periodic Labs)
Pre-training has well-established power-law relationships like Chinchilla scaling, but RL has lacked predictive models due to shifting data distributions and bounded reward functions.
Pre-training vs RL Scaling: Fundamental Differences
Experimental Scope: 400,000 GPU-Hours
The Sigmoid Framework
Performance follows an S-curve with three parameters: A (asymptotic ceiling), B (compute efficiency), and C_mid (midpoint). This framework enables extrapolation from small-scale experiments to predict large-scale performance.
Sigmoid Parameters: Interactive Exploration
Extrapolation Validation: Fit Early, Predict Late
Fitted sigmoid curve
Training checkpoints (fitted)
Validation checkpoints (predicted)
Algorithm Comparison: Asymptotic Performance
Not all RL methods converge to the same ceiling. The choice of algorithm fundamentally determines parameter A, the asymptotic performance limit.
Asymptotic Pass Rates by Method
Performance Heatmap: Methods × Compute Budget
ScaleRL Components: Best-Practice Recipe
ScaleRL integrates known techniques to maximize both asymptotic performance and compute efficiency. Each component was validated with 16,000-GPU-hour ablations.
Architecture: Asynchronous Pipeline-RL
Component Impact: Ablation Results
Extrapolation Results: Validation Across Scales
The sigmoid framework accurately predicts performance at 100k GPU-hours from fits at 50k, demonstrating reliable extrapolation within the tested regime.
8B Dense Model: 100k GPU-Hour Validation
Scout MoE Model: 50k GPU-Hour Validation
Key References
Primary Paper: Khatri et al. (2025) — The Art of Scaling Reinforcement Learning Compute for LLMs arXiv:2510.13786
Scaling Laws: Kaplan et al. (2020) — Scaling Laws for Neural Language Models
Hoffmann et al. (2022) — Training Compute-Optimal Large Language Models (Chinchilla)
RL Methods: Schulman et al. (2017) — Proximal Policy Optimization (PPO)
Rafailov et al. (2023) — Direct Preference Optimization (DPO)
Reasoning: Wei et al. (2022) — Chain-of-Thought Prompting
Zelikman et al. (2022) — STaR: Self-Taught Reasoner
Lightman et al. (2023) — Let's Verify Step by Step
Production Systems: OpenAI (2024) — o1 System Card
DeepSeek-AI (2025) — DeepSeek-R1