Traditional vs. Context-Aware Scaling Laws
The gap between upstream metrics (training loss) and downstream metrics (task accuracy)
Two-Dimensional Scaling Space
Context-aware scaling laws model performance as a function of two independent axes
Functional Forms Compared
Mathematical structure of traditional vs. context-aware scaling laws
Performance Prediction: Context-Blind vs. Context-Aware
Simulated accuracy across context lengths for arithmetic task (inspired by paper Figure 1)
Observed accuracy (ground truth)
Context-blind: flat prediction (Chen 2024 style)
Context-aware: saturating curve (Montgomery 2025)
Task Accuracy Heatmap: Compute vs. Context
Simulated downstream accuracy across training compute budgets and context lengths (Llama-2-7B, Arithmetic task)
Task-Specific Exponents
Fitted power-law exponents vary by task type, reflecting different saturation behaviors
Interactive Saturation Explorer
Toggle between compute scaling and context scaling to see saturation behavior
Compute scaling: performance vs. training FLOPs
Context scaling: performance vs. prompt length
Penalty region: exceeding trained context limit
Penalty Term Visualization
Performance degradation when prompt length exceeds trained context window
Experimental Pipeline
How the paper generated checkpoints and fitted the context-aware scaling law
Missing Citation: Gemini 1.5 Context Scaling
Google DeepMind demonstrated context-length power laws 18 months before this paper