This episode examines TimesFM, Google Research's decoder-only foundation model for time-series forecasting, and its central claim that a single pretrained 200-million-parameter model can forecast unfamiliar datasets zero-shot, without fine-tuning, at accuracy close to models trained specifically on each dataset. The hosts trace the architecture's lineage from patching, borrowed from the Vision Transformer's image-patch approach and specifically from PatchTST's time-series adaptation, to TimesFM's own contribution of pairing patched inputs with a causal, GPT-style autoregressive setup that naturally handles variable context lengths. They contrast this against DeepAR's RNN-based forecasting, which still required target series in training, and against a 2023 NeurIPS trick of feeding raw numbers as text into large language models, which TimesFM claims to beat at a fraction of the cost. A key surprise is the training corpus itself: since real time-series data is far scarcer online than text, the roughly 100 billion timepoints come largely from Google Trends and Wikipedia pageviews, supplemented by synthetic ARMA and seasonal processes engineered to fill coverage gaps. Listeners interested in foundation models, forecasting infrastructure, or how architectural ideas transfer across modalities will find the discussion's skepticism about benchmark claims and evaluation rigor especially engaging as the hosts preview a closer look at the paper's actual scoring methodology.
Sources:
1. TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting
https://arxiv.org/pdf/2310.106882.
https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/ https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/3. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks — David Salinas, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, 2017 (arXiv), 2020 (Intl. J. Forecasting)
https://scholar.google.com/scholar?q=DeepAR%3A+Probabilistic+Forecasting+with+Autoregressive+Recurrent+Networks4. N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting — Boris Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio, 2019 (arXiv), ICLR 2020
https://scholar.google.com/scholar?q=N-BEATS%3A+Neural+Basis+Expansion+Analysis+for+Interpretable+Time+Series+Forecasting5. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting — Bryan Lim, Sercan Arik, Nicolas Loeff, Tomas Pfister, 2019 (arXiv), 2021 (Intl. J. Forecasting)
https://scholar.google.com/scholar?q=Temporal+Fusion+Transformers+for+Interpretable+Multi-horizon+Time+Series+Forecasting6. Chronos: Learning the Language of Time Series — Abdul Fatir Ansari, Lorenzo Stella, et al. (Amazon), 2024
https://scholar.google.com/scholar?q=Chronos%3A+Learning+the+Language+of+Time+Series7. Large Language Models Are Zero-Shot Time Series Forecasters — Nate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon Wilson, 2023 (NeurIPS)
https://scholar.google.com/scholar?q=Large+Language+Models+Are+Zero-Shot+Time+Series+Forecasters8. One Fits All: Power General Time Series Analysis by Pretrained LM — Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, Rong Jin, 2023 (NeurIPS)
https://scholar.google.com/scholar?q=One+Fits+All%3A+Power+General+Time+Series+Analysis+by+Pretrained+LM9. Moirai: Unified Training of Universal Time Series Forecasting Transformers — Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, Doug Arnold (Salesforce), 2024 (ICML)
https://scholar.google.com/scholar?q=Moirai%3A+Unified+Training+of+Universal+Time+Series+Forecasting+Transformers10. Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting — Kashif Rasul, Arjun Ashok, Andrew Robert Williams, et al., 2024
https://scholar.google.com/scholar?q=Lag-Llama%3A+Towards+Foundation+Models+for+Probabilistic+Time+Series+Forecasting11. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. (Google), 2020
https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale12. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2023 (ICLR)
https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-term+Forecasting+with+Transformers13. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting — Haoyi Zhou, Shanghang Zhang, Jieqi Peng, et al., 2021 (AAAI, Best Paper)
https://scholar.google.com/scholar?q=Informer%3A+Beyond+Efficient+Transformer+for+Long+Sequence+Time-Series+Forecasting14. Masked Autoencoders Are Scalable Vision Learners — Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick, 2021
https://scholar.google.com/scholar?q=Masked+Autoencoders+Are+Scalable+Vision+Learners15. A Time Series is Worth 64 Words: Long-Term Forecasting with Transformers (PatchTST) — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2022
https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-Term+Forecasting+with+Transformers+%28PatchTST%2916. TimeGPT-1 — Azul Garza, Max Mergenthaler-Canseco, 2023
https://scholar.google.com/scholar?q=TimeGPT-117. Training Compute-Optimal Large Language Models (Chinchilla) — Jordan Hoffmann et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models+%28Chinchilla%29Interactive Visualization: TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting