← All episodes Data Temporality's Hidden Impact on LLM Pretraining

Data Temporality's Hidden Impact on LLM Pretraining

Aug 2, 2026
This episode explores why open-weight LLMs like Llama 3.1, Gemma3, Qwen3, and Olmo3 systematically lose 11–39% relative accuracy on facts from 2023–2024 compared to facts from 2020–2021, even though the more recent data falls within their training window. Drawing on Kyutai's paper "Understanding Data Temporality Impact on Large Language Models Pre-training," the discussion traces this "knowledge horizon gap" to a design choice baked into standard pretraining: corpora from many years are pooled and globally shuffled before training, erasing any timestamp signal and letting older, more frequently re-crawled data dominate. The hosts connect this to learning-rate decay schedules, arguing that data seen late in training — when updates are small and durable — gets imprinted far more strongly than data seen early, so chronological ordering (feeding snapshots 2018 through 2025 in sequence) could exploit that same mechanism to anchor recent facts instead of losing them. They situate the work against Zhao et al.'s "Set the Clock" research and Bengio's foundational curriculum-learning ideas, framing chronological training as a strikingly cheap intervention — same tokens, same compute, same architecture — for a problem the field has largely ignored. It's a compelling listen for anyone puzzling over why "knowledge cutoff" claims don't match what models actually seem to know.
Sources:
1. Understanding Data Temporality Impact on Large Language Models Pre-training — Hippolyte Pilchen, Romain Fabre, Franck Signe Talla, Patrick Perez, Edouard Grave, 2026
http://arxiv.org/abs/2605.22769
2. Set the Clock: Temporal Alignment of Pretrained Language Models — Bowen Zhao, Zander Brumbaugh, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, 2024
https://scholar.google.com/scholar?q=Set+the+Clock%3A+Temporal+Alignment+of+Pretrained+Language+Models
3. Time-Aware Language Models as Temporal Knowledge Bases — Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, William W. Cohen, 2022
https://scholar.google.com/scholar?q=Time-Aware+Language+Models+as+Temporal+Knowledge+Bases
4. TemporalWiki: A Lifelong Benchmark for Training and Evaluating Ever-Evolving Language Models — Joel Jang, Seonghyeon Ye, Changho Lee, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Minjoon Seo, 2022
https://scholar.google.com/scholar?q=TemporalWiki%3A+A+Lifelong+Benchmark+for+Training+and+Evaluating+Ever-Evolving+Language+Models
5. RealTime QA: What's the Answer Right Now? — Jungo Kasai, Keisuke Sakaguchi, Yoichi Takahashi, Yutaro Yamada, Deqing Fu, Tushar Khot, Ashish Sabharwal, Rik Koncel-Kedziorski, Yejin Choi, Noah A. Smith, Kentaro Inui, 2022
https://scholar.google.com/scholar?q=RealTime+QA%3A+What%27s+the+Answer+Right+Now%3F
6. Curriculum Learning — Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Jason Weston, 2009
https://scholar.google.com/scholar?q=Curriculum+Learning
7. TimeLMs: Diachronic Language Models from Twitter — Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, Jose Camacho-Collados, 2022
https://scholar.google.com/scholar?q=TimeLMs%3A+Diachronic+Language+Models+from+Twitter
8. In-Context Pretraining: Language Modeling Beyond Document Boundaries — Weijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Wen-tau Yih, Mike Lewis, 2023
https://scholar.google.com/scholar?q=In-Context+Pretraining%3A+Language+Modeling+Beyond+Document+Boundaries
9. Towards Continual Knowledge Learning of Language Models — Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, Minjoon Seo, 2022
https://scholar.google.com/scholar?q=Towards+Continual+Knowledge+Learning+of+Language+Models
10. TiC-LM: A web-scale benchmark for time-continual LLM pretraining — Li, J., Armandpour, M., Mirzadeh, I., Mehta, S., Shankar, V., Vemulapalli, R., Bengio, S., Tuzel, O., Farajtabar, M., Pouransari, H., Faghri, F., 2025
https://scholar.google.com/scholar?q=TiC-LM%3A+A+web-scale+benchmark+for+time-continual+LLM+pretraining
11. How do language models learn facts? Dynamics, curricula and hallucinations — Zucchet, N., Bornschein, J., Chan, S. C., Lampinen, A. K., Pascanu, R., De, S., 2025
https://scholar.google.com/scholar?q=How+do+language+models+learn+facts%3F+Dynamics%2C+curricula+and+hallucinations
12. Data mixing can induce phase transitions in knowledge acquisition — Gu, X., Lyu, K., Li, J., Zhang, J., 2026
https://scholar.google.com/scholar?q=Data+mixing+can+induce+phase+transitions+in+knowledge+acquisition
13. TiMoE: Time-aware mixture of language experts — Faro, R., Fan, D., Alphaidze, T., Jaggi, M., 2025
https://scholar.google.com/scholar?q=TiMoE%3A+Time-aware+mixture+of+language+experts
14. Does your data spark joy? Performance gains from domain upsampling at the end of training — Blakeney, C., Paul, M., Larsen, B. W., Owen, S., Frankle, J., 2024
https://scholar.google.com/scholar?q=Does+your+data+spark+joy%3F+Performance+gains+from+domain+upsampling+at+the+end+of+training
Interactive Visualization: Data Temporality's Hidden Impact on LLM Pretraining