The Art of Scaling Reinforcement Learning Compute for LLMs

Meta AI and collaborators establish the first systematic scaling laws for reinforcement learning in large language models, based on over 400,000 GPU-hours of experiments.

Authors: Khatri et al. (Meta, UT Austin, UCL, UC Berkeley, Harvard, Periodic Labs)
Published: October 2025

The Scaling Challenge

Pre-training has well-established power-law relationships like Chinchilla scaling, but RL has lacked predictive models due to shifting data distributions and bounded reward functions.

Pre-training vs RL Scaling: Fundamental Differences

PRE-TRAINING Compute (GPU-hours) Loss ✓ Static dataset ✓ Smooth power law: L ∝ C^(-α) REINFORCEMENT LEARNING Asymptote A Compute (GPU-hours) Performance ✗ Shifting data distribution ✗ Bounded rewards (saturation) → Sigmoid curve: A / (1 + e^(-B(C-C_mid)))

Experimental Scope: 400,000 GPU-Hours

Ablation Phase 16,000 GPU-hours per run GRPO, DAPO, PPO, REINFORCE ~240,000 GPU-hours total ScaleRL 8B 100,000 GPU-hours 7,400 training steps 3.5× longer than ProRL ScaleRL Scout MoE 50,000 GPU-hours 17B × 16 experts 7,100 training steps Task: AIME Math Reasoning Dataset: Qwen 2.5-Math-Instruct (synthetic) Validation: AIME 2024 (out-of-distribution)

The Sigmoid Framework

Performance follows an S-curve with three parameters: A (asymptotic ceiling), B (compute efficiency), and C_mid (midpoint). This framework enables extrapolation from small-scale experiments to predict large-scale performance.

Sigmoid Parameters: Interactive Exploration

Compute (GPU-hours ×1000) Pass Rate 0 25 50 75 100 0.0 0.2 0.4 0.6 0.8 Performance = A / (1 + e^(-B(C - C_mid)))

Extrapolation Validation: Fit Early, Predict Late

FITTING REGION (0-50k GPU-hrs) EXTRAPOLATION (50k-100k GPU-hrs) FIT → PREDICT GPU-hours (×1000) Pass Rate
Fitted sigmoid curve
Training checkpoints (fitted)
Validation checkpoints (predicted)

Algorithm Comparison: Asymptotic Performance

Not all RL methods converge to the same ceiling. The choice of algorithm fundamentally determines parameter A, the asymptotic performance limit.

Asymptotic Pass Rates by Method

ScaleRL: A=0.61 MiniMax: A=0.58 DAPO: A=0.54 GRPO: A=0.52 Magistral: A=0.49 GPU-hours (×1000) Pass Rate 0 20 40 60 80 0.0 0.2 0.4 0.6 0.8

Performance Heatmap: Methods × Compute Budget

5k 10k 20k 40k 60k 80k 100k GPU-hours ScaleRL MiniMax DAPO GRPO Magistral 0.25 0.35 0.44 0.52 0.57 0.60 0.61 0.23 0.33 0.42 0.49 0.54 0.57 0.58 0.21 0.30 0.39 0.46 0.51 0.53 0.54 0.19 0.28 0.36 0.43 0.48 0.51 0.52 0.17 0.26 0.34 0.40 0.45 0.48 0.49 Pass Rate 0.6 0.4 0.2

ScaleRL Components: Best-Practice Recipe

ScaleRL integrates known techniques to maximize both asymptotic performance and compute efficiency. Each component was validated with 16,000-GPU-hour ablations.

Architecture: Asynchronous Pipeline-RL

GENERATION CLUSTER Policy Model (frozen) Generate rollouts Reward Model Score trajectories Data Queue (prompt, response, reward) async transfer TRAINING CLUSTER Sample Batch Zero-variance filtering Forced length interruptions Loss Computation CISPO (truncated importance sampling) Prompt-level averaging, FP32 logits Gradient Update Batch-level advantage normalization No-Positive-Resampling periodic sync Updated Policy Copy to generation cluster

Component Impact: Ablation Results

Extrapolation Results: Validation Across Scales

The sigmoid framework accurately predicts performance at 100k GPU-hours from fits at 50k, demonstrating reliable extrapolation within the tested regime.

8B Dense Model: 100k GPU-Hour Validation

A = 0.61 FIT @ 50k 65k 85k 100k Extrapolation Accuracy Predicted vs Actual @ 100k: 0.608 vs 0.610 (0.3% error) GPU-hours (×1000) Pass Rate 0 25 50 75 100 0.0 0.2 0.4 0.6 0.8

Scout MoE Model: 50k GPU-Hour Validation

A = 0.64 FIT @ 16k 25k 35k 50k Scout MoE Architecture 17B params × 16 experts 7,100 training steps Higher asymptote than 8B dense GPU-hours (×1000) Pass Rate 0 15 25 35 50 0.0 0.2 0.4 0.6 0.8

Key References

Primary Paper: Khatri et al. (2025) — The Art of Scaling Reinforcement Learning Compute for LLMs
arXiv:2510.13786
Scaling Laws: Kaplan et al. (2020) — Scaling Laws for Neural Language Models
Hoffmann et al. (2022) — Training Compute-Optimal Large Language Models (Chinchilla)
RL Methods: Schulman et al. (2017) — Proximal Policy Optimization (PPO)
Rafailov et al. (2023) — Direct Preference Optimization (DPO)
Reasoning: Wei et al. (2022) — Chain-of-Thought Prompting
Zelikman et al. (2022) — STaR: Self-Taught Reasoner
Lightman et al. (2023) — Let's Verify Step by Step
Production Systems: OpenAI (2024) — o1 System Card
DeepSeek-AI (2025) — DeepSeek-R1