← All episodes TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

Sep 3, 2026
This episode examines TimesFM, Google Research's decoder-only foundation model for time-series forecasting, and its central claim that a single pretrained 200-million-parameter model can forecast unfamiliar datasets zero-shot, without fine-tuning, at accuracy close to models trained specifically on each dataset. The hosts trace the architecture's lineage from patching, borrowed from the Vision Transformer's image-patch approach and specifically from PatchTST's time-series adaptation, to TimesFM's own contribution of pairing patched inputs with a causal, GPT-style autoregressive setup that naturally handles variable context lengths. They contrast this against DeepAR's RNN-based forecasting, which still required target series in training, and against a 2023 NeurIPS trick of feeding raw numbers as text into large language models, which TimesFM claims to beat at a fraction of the cost. A key surprise is the training corpus itself: since real time-series data is far scarcer online than text, the roughly 100 billion timepoints come largely from Google Trends and Wikipedia pageviews, supplemented by synthetic ARMA and seasonal processes engineered to fill coverage gaps. Listeners interested in foundation models, forecasting infrastructure, or how architectural ideas transfer across modalities will find the discussion's skepticism about benchmark claims and evaluation rigor especially engaging as the hosts preview a closer look at the paper's actual scoring methodology.
Sources:
1. TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting
https://arxiv.org/pdf/2310.10688
2.
https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/
https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/
3. DeepAR: Probabilistic Forecasting with Autoregressive Recurrent Networks — David Salinas, Valentin Flunkert, Jan Gasthaus, Tim Januschowski, 2017 (arXiv), 2020 (Intl. J. Forecasting)
https://scholar.google.com/scholar?q=DeepAR%3A+Probabilistic+Forecasting+with+Autoregressive+Recurrent+Networks
4. N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting — Boris Oreshkin, Dmitri Carpov, Nicolas Chapados, Yoshua Bengio, 2019 (arXiv), ICLR 2020
https://scholar.google.com/scholar?q=N-BEATS%3A+Neural+Basis+Expansion+Analysis+for+Interpretable+Time+Series+Forecasting
5. Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting — Bryan Lim, Sercan Arik, Nicolas Loeff, Tomas Pfister, 2019 (arXiv), 2021 (Intl. J. Forecasting)
https://scholar.google.com/scholar?q=Temporal+Fusion+Transformers+for+Interpretable+Multi-horizon+Time+Series+Forecasting
6. Chronos: Learning the Language of Time Series — Abdul Fatir Ansari, Lorenzo Stella, et al. (Amazon), 2024
https://scholar.google.com/scholar?q=Chronos%3A+Learning+the+Language+of+Time+Series
7. Large Language Models Are Zero-Shot Time Series Forecasters — Nate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon Wilson, 2023 (NeurIPS)
https://scholar.google.com/scholar?q=Large+Language+Models+Are+Zero-Shot+Time+Series+Forecasters
8. One Fits All: Power General Time Series Analysis by Pretrained LM — Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, Rong Jin, 2023 (NeurIPS)
https://scholar.google.com/scholar?q=One+Fits+All%3A+Power+General+Time+Series+Analysis+by+Pretrained+LM
9. Moirai: Unified Training of Universal Time Series Forecasting Transformers — Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, Doug Arnold (Salesforce), 2024 (ICML)
https://scholar.google.com/scholar?q=Moirai%3A+Unified+Training+of+Universal+Time+Series+Forecasting+Transformers
10. Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting — Kashif Rasul, Arjun Ashok, Andrew Robert Williams, et al., 2024
https://scholar.google.com/scholar?q=Lag-Llama%3A+Towards+Foundation+Models+for+Probabilistic+Time+Series+Forecasting
11. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. (Google), 2020
https://scholar.google.com/scholar?q=An+Image+is+Worth+16x16+Words%3A+Transformers+for+Image+Recognition+at+Scale
12. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2023 (ICLR)
https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-term+Forecasting+with+Transformers
13. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting — Haoyi Zhou, Shanghang Zhang, Jieqi Peng, et al., 2021 (AAAI, Best Paper)
https://scholar.google.com/scholar?q=Informer%3A+Beyond+Efficient+Transformer+for+Long+Sequence+Time-Series+Forecasting
14. Masked Autoencoders Are Scalable Vision Learners — Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, Ross Girshick, 2021
https://scholar.google.com/scholar?q=Masked+Autoencoders+Are+Scalable+Vision+Learners
15. A Time Series is Worth 64 Words: Long-Term Forecasting with Transformers (PatchTST) — Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant Kalagnanam, 2022
https://scholar.google.com/scholar?q=A+Time+Series+is+Worth+64+Words%3A+Long-Term+Forecasting+with+Transformers+%28PatchTST%29
16. TimeGPT-1 — Azul Garza, Max Mergenthaler-Canseco, 2023
https://scholar.google.com/scholar?q=TimeGPT-1
17. Training Compute-Optimal Large Language Models (Chinchilla) — Jordan Hoffmann et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models+%28Chinchilla%29
Interactive Visualization: TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting