Predicting Task Performance with Context-aware Scaling Laws

Extending neural scaling laws to predict downstream task accuracy while modeling context length as a first-order variable

📄 arXiv:2510.14919 Montgomery et al. • Oct 2025 UC Santa Cruz • Databricks • Google DeepMind • Berkeley

Traditional vs. Context-Aware Scaling Laws

The gap between upstream metrics (training loss) and downstream metrics (task accuracy)

Traditional Scaling Laws (Kaplan 2020, Chinchilla 2022) Upstream Metrics Cross-entropy loss Perplexity ? Downstream Tasks Accuracy on real-world tasks Context-Aware Scaling Laws (Montgomery 2025) Direct prediction Bypasses upstream loss Task Performance f(compute, context) + penalty term for exceeding n_ctx Key Insight Traditional scaling laws predict training loss but ignore (1) downstream task accuracy and (2) context length effects. Context-aware laws model both compute and context as saturating power laws, directly predicting what users care about.

Two-Dimensional Scaling Space

Context-aware scaling laws model performance as a function of two independent axes

Training Compute (C) Context Length (n_pmt) Low compute Low context Medium compute Low context High compute Low context Low compute Medium context Medium compute Medium context High compute Medium context Low compute High context Medium compute High context Penalty zone: n_pmt exceeds trained n_ctx Compute saturation Context saturation
Poor performance
Moderate performance
High performance

Functional Forms Compared

Mathematical structure of traditional vs. context-aware scaling laws

Kaplan et al. 2020 Loss(C) = (C_0 / C)^α Only predicts upstream cross-entropy loss. Context ignored. Chinchilla (Hoffmann et al.) 2022 Loss(N, D) = E + A/N^α + B/D^β Refines compute-optimal N/D ratio. Still upstream loss only. Chen et al. 2024 (Two-stage) Stage 1: Loss(C) = classical scaling law Stage 2: Accuracy = f(Loss) Maps upstream to downstream, but context-blind in both stages. Montgomery et al. 2025 (Context-Aware) Accuracy(C, n_pmt, n_ctx) = a · (1 - (C_0/C)^α_C) · (1 - (n_0/n_pmt)^α_n) · penalty(n_pmt, n_ctx) Direct downstream prediction. Dual saturating power laws. Penalty term models degradation when n_pmt exceeds n_ctx.

Performance Prediction: Context-Blind vs. Context-Aware

Simulated accuracy across context lengths for arithmetic task (inspired by paper Figure 1)

0% 20% 40% 60% 80% Task Accuracy 0 8 16 32 64 128 Prompt Tokens (k)
Observed accuracy (ground truth)
Context-blind: flat prediction (Chen 2024 style)
Context-aware: saturating curve (Montgomery 2025)

Task Accuracy Heatmap: Compute vs. Context

Simulated downstream accuracy across training compute budgets and context lengths (Llama-2-7B, Arithmetic task)

128k 64k 32k 16k 8k 4k Context Window (n_ctx) 400 800 1.6k 3.2k 6.4k 12.8k 25.6k Training Steps (Continued Training) Accuracy 0% 50% 100%

Task-Specific Exponents

Fitted power-law exponents vary by task type, reflecting different saturation behaviors

Power-law Exponent (α) 0.00 0.10 0.20 0.30 0.40 Arithmetic α=0.25 Commonsense α=0.20 Translation α=0.32 Higher α = faster saturation with added context

Interactive Saturation Explorer

Toggle between compute scaling and context scaling to see saturation behavior

0 25 50 75 100 Accuracy (%) Training Compute (FLOPs) Saturation zone Diminishing returns Penalty zone n_pmt > n_ctx
Compute scaling: performance vs. training FLOPs
Context scaling: performance vs. prompt length
Penalty region: exceeding trained context limit

Penalty Term Visualization

Performance degradation when prompt length exceeds trained context window

0 20 40 60 80 Task Accuracy (%) 0.5x 1.0x 1.5x 2.0x 2.5x n_pmt / n_ctx Ratio Safe zone (n_pmt ≤ n_ctx) Penalty zone (n_pmt > n_ctx) n_ctx limit Graceful degradation Severe penalty

Experimental Pipeline

How the paper generated checkpoints and fitted the context-aware scaling law

Base Models Llama-2-7B Llama-2-13B YaRN Extension 400 steps on PG-19 Context: 8k, 16k, 32k, 64k, 128k Checkpoint Grid 12 models total: 2 base sizes × 6 context limits Downstream Evaluation 3 tasks × 12 models • Arithmetic reasoning • Commonsense reasoning • Machine translation Fit Scaling Law Accuracy(C, n_pmt, n_ctx) Extract exponents α_C, α_n and penalty parameters Predict Performance Extrapolate to unseen compute budgets and context lengths Data Flow Base pretrained models Continued training with YaRN Checkpoints at intervals Downstream task accuracy Fitted parameters Experimental Limitations • Only two model sizes (7B, 13B) — no validation on larger scales (70B, 405B) • Only Llama-2 architecture — no cross-architecture testing (Mistral, GPT, Gemini) • Only YaRN continued training — not tested on models trained with long context from scratch

Missing Citation: Gemini 1.5 Context Scaling

Google DeepMind demonstrated context-length power laws 18 months before this paper

Jan 2020 Kaplan scaling laws Mar 2022 Chinchilla Mar 2024 Gemini 1.5 Tech Report Context power laws! Oct 2025 Montgomery et al. (This paper) 18 months No Gemini citation

References & Sources

  1. Predicting Task Performance with Context-aware Scaling Laws — Montgomery et al., 2025
    arXiv:2510.14919
  2. Scaling Laws for Neural Language Models — Kaplan, McCandlish, Henighan, Brown, et al., 2020
    Google Scholar
  3. Training Compute-Optimal Large Language Models (Chinchilla) — Hoffmann, Borgeaud, Mensch, et al., 2022
    Google Scholar
  4. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context — Google DeepMind, 2024
    arXiv:2403.05530
  5. YaRN: Efficient Context Window Extension of Large Language Models — Peng et al., 2024
    Google Scholar
  6. Inverse Scaling: When Bigger Isn't Better — McKenzie, Lyzhov, Pieler, et al., 2023
    Google Scholar
  7. Predictability and Surprise in Large Generative Models — Ganguli, Hernandez, Lovitt, et al. (Anthropic), 2022
    Google Scholar