AI Post Transformers — Episode Companion

Data Temporality's Hidden Impact on LLM Pretraining

Paper: Understanding Data Temporality Impact on LLMs Pre-training Authors: Pilchen, Fabre, et al. — Kyutai Posted: May 25, 2026 · ICML 2026 arXiv:2605.22769 ↗

Every recent open-weight model loses 11–39% relative accuracy on facts from 2023–2024

Same model, same weights — just facts from closer to the training cutoff score worse than facts from years earlier. Kyutai's sequential-order model breaks that pattern.

KairosQA accuracy by fact year, per model

Cloze-format accuracy (0–1) on Wikidata-derived temporal facts. Warmer = higher accuracy. Watch the right-hand columns (2023–2024): every real open-weight model cools off except the sequentially-trained one.

Relative accuracy loss: 2023–24 vs 2020–21

(old accuracy − recent accuracy) / old accuracy. The paper's headline range is 11–39% for shipping models.

Old facts vs. recent facts, head to head

Sequential-6B matches everyone on 2020–21 knowledge and beats all six models — including 14B Qwen3 — on 2023–24 knowledge.
2020–2021 facts 2023–2024 facts
Scale doesn't fix it: bigger Qwen3 checkpoints (4B → 8B → 14B) just shift the same decaying curve upward — they don't flatten it.

One design choice: shuffle everything, or feed it in order

Standard pretraining pools every Common Crawl snapshot, dedupes and quality-filters each one, then globally shuffles before minibatching — erasing any timestamp signal. Kyutai's sequential model changes only the order.

Every year's snapshot is stripped, filtered, deduped — then dumped into one timeless pool and globally shuffled. A 2013 sentence and a 2024 sentence are equally likely in the same batch.

dactory filtering pipeline

Both runs share the same open-source filtering stack before the ordering choice is applied.

Quality weights by domain

Domain-weighted classifier scores used to weight the 2.5T token budget.

Why order matters: learning-rate decay makes late data sticky

Updates late in training, when the learning rate is low, reshape the weights far more durably than updates from early on. Putting recent facts last exploits the same mechanism that normally causes catastrophic forgetting — pointed at retention instead.

Learning-rate schedule over a training run

Relative LR across the run. The shaded zone is where updates become durable/"sticky" — where sequential ordering places the most recent snapshots.

Branching cooldown: 8 checkpoints from one run

The sequential run marches 2018 → 2025 one year at a time. Each year, a 30,000-step decay ("branching cooldown") carves off a fully-converged checkpoint for evaluation.
Curriculum learning traces back to Bengio et al., 2009 (University of Montreal). Kyutai's twist: apply it along the time axis of the corpus itself, not task difficulty.

Is KairosQA just an easy, gameable benchmark?

Two stress tests the hosts pushed on: does accuracy survive harder multiple-choice, and does it track genuine subject familiarity rather than memorized distractor patterns.

Accuracy vs. number of cloze choices

At 2 choices, 58% is close to a coin flip. Widen to 12 choices — random collapses toward 8%, but the model holds a steady ~15-point margin above chance.
Sequential model Random baseline

F1 by subject popularity

Genuine knowledge should track fame. ~0.40 F1 on the most popular 10% of subjects, down to 0.10–0.15 on the long tail — a curve a gamed benchmark wouldn't produce.
KairosQA is built from 17M Wikidata triplets, filtered to 7,167 subject-relation pairs whose answers changed at least twice between 2018–2025, weighted by Wikipedia page views.

Not a free lunch — three cracks worth sitting with

A data-quality confound in the headline gap, a forgetting cost on older facts, and a result that mostly evaporates on an external benchmark the authors didn't build.

Ablation: which year's data cools the checkpoint?

Cooling the 2021 checkpoint on 2024 data instead of 2021 data adds +1.5 OLMES — some of the "ordering effect" may just be newer Common Crawl being richer text.

Pre-2020 knowledge: what sequential ordering costs

The 2025 sequential checkpoint measurably forgets pre-2020 facts relative to the shuffled baseline. Souping checkpoints and 50/50 replay cooldown (Appendix A.3) both came back inconclusive.
Shuffled baseline Sequential (2025 ckpt)

TAQA (external benchmark): the gap shrinks to noise

On Zhao et al.'s TAQA, every model scores 1.5–10% F1 with only marginal differences — TAQA was designed for 60B+ models, but the headline claim shows up clearly only on the benchmark Kyutai built and tuned itself.
Scorecard: a real, honestly-reported gain on KairosQA · general capability untouched · but a baseline that's more heirloom than matched control · a headline result partly wearing a data-quality costume · external validation that mostly evaporates · a forgetting problem nobody's solved yet.

References