1 00:00:01,000 --> 00:00:49,669 [Hal Turing] Alrighty! Thanks for tuning in! Hello AI world! I am your host, Hal Turing, and my co-host is Dr. Ada Shannon. Today we're digging into "A Decoder-Only Foundation Model for Time-Series Forecasting" — that's TimesFM — by Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou, four authors out of Google Research, posted to arXiv in October 2023 and revised through April 2024. And Ada, here's the number that stopped me: a 200-million-parameter model, trained once, forecasting datasets it has never seen, landing close to models trained specifically for each one of those datasets. No fine-tuning, no retraining pipeline per customer. 2 00:00:49,669 --> 00:01:21,944 [Dr. Ada Shannon] Yeah, and what I like is they're upfront about why that's a weird thing to even attempt. In NLP you've got a shared vocabulary — English is English whether it's a tweet or a legal brief. A retail sales series and an ECG signal share basically nothing: different units, different scales, different noise. So the question they're chasing is blunt — can you pretrain one model on a giant pile of time series and have it generalize the way GPT generalizes across writing tasks it was never explicitly trained on? That's the zero-shot bet. 3 00:01:21,944 --> 00:02:12,657 [Hal Turing] Let's actually define that term for anyone new to it, because it's doing a lot of work here. Zero-shot forecasting means you hand the model a dataset it never saw during training — different domain, different scale, different seasonality — and it produces a forecast with zero gradient updates, zero fine-tuning. Compare that to something like DeepAR, from Salinas, Flunkert, Gasthaus and Januschowski, originally 2017 on arXiv and later in the International Journal of Forecasting in 2020. DeepAR was a big deal — it trained one RNN across thousands of related series instead of fitting ARIMA per series — but it still needed the target series baked into training. TimesFM is asking for something stricter. 4 00:02:12,657 --> 00:02:57,889 [Dr. Ada Shannon] Right, and it matters practically, not just academically. Think retail demand forecasting, energy load, traffic, weather — every one of those currently means standing up a training or fine-tuning pipeline per dataset before you get a usable forecast. If a cold-start product with three days of sales history could get a decent forecast with zero training burden, that's real compute and engineering time saved. The paper's other target is sneakier though — there was a 2023 NeurIPS paper by Gruver, Finzi, Qiu and Wilson showing you could feed numbers as literal text strings into GPT-3 or LLaMA-2 and get surprisingly competitive zero-shot forecasts. TimesFM's authors are explicitly saying: a small model built for this purpose beats that trick, at a fraction of the cost. 5 00:02:57,889 --> 00:03:51,249 [Hal Turing] Which is a great contrast, because it tells you something about where foundation-model magic does and doesn't transfer across modalities. Okay, so how do you actually build a model that eats a time series at all? This is where patching comes in — and it's borrowed straight from computer vision. Dosovitskiy and colleagues at Google, the Vision Transformer paper from 2020, chopped images into 16-by-16 pixel patches instead of feeding one pixel at a time into a transformer. TimesFM does the analogous thing: it breaks a time series into fixed-length contiguous chunks — say 32 consecutive points — and each chunk becomes one input token. One token isn't one timestep, it's a little shape: a rising edge, a dip, a seasonal bump. 6 00:03:51,249 --> 00:04:37,828 [Dr. Ada Shannon] Oh wait wait wait — before you move on, I want to flag whose idea this specifically was in the time-series world, because it's not TimesFM's. That's PatchTST, "A Time Series is Worth 64 Words," Nie, Nguyen, Sinthong and Kalagnanam, ICLR 2023 — title's a direct wink at the ViT paper. PatchTST proved patching plus channel-independence crushed the older point-by-point transformers like Informer on long-horizon forecasting. TimesFM's real architectural move is taking that patching idea and marrying it to a decoder-only, GPT-style causal setup instead of PatchTST's encoder-style supervised setup — trained to predict the next patch from past patches, same next-token logic as a language model, just one abstraction level up. 7 00:04:37,828 --> 00:04:49,206 [Hal Turing] And that decoder-only choice is what actually enables the zero-shot part, right? Because a causal, autoregressive-style model naturally handles "I don't know how much context you're going to hand me." 8 00:04:49,206 --> 00:05:38,153 [Dr. Ada Shannon] Exactly — and they add a masking trick during training so the model has seen every context length from one patch up to the max, not just neat multiples. But the piece I think listeners will find most surprising is what they actually trained it on. Real time series are nowhere near as abundant online as text is, so the corpus leans hard on Google Trends search-query volumes and Wikipedia pageview counts — both public, both huge, both free of the usual licensing headaches — topped up with algorithmically generated synthetic series: ARMA processes, seasonal mixtures, trends, step functions, built specifically to plug gaps the real data didn't cover. Roughly O of 100 billion timepoints total, feeding a 200-million-parameter model — tiny by LLM standards. 9 00:05:38,153 --> 00:06:09,872 [Hal Turing] That synthetic-data-as-a-patch idea feels like the part worth sitting with before we go further. In Part 2 we'll get into the actual architecture details — why the output patches are longer than the input patches, how that corpus breaks down number by number, and how TimesFM actually scores against Monash, Darts, and the Informer benchmarks. Then in Part 3, Ada's going to make me defend some of these evaluation claims, because there are some real questions about what "beats the baseline" means here once you look closely. 10 00:06:09,872 --> 00:06:11,079 [Dr. Ada Shannon] Oh, I've got notes. 11 00:06:11,079 --> 00:06:59,980 [Dr. Ada Shannon] Here's the piece worth digging into: the output patches are longer than the input patches, and that's the whole reason zero-shot forecasting stays cheap at inference. Say input patch length is 32 and output patch length is 128. During training, the model learns to take the first 32 points and predict the next 128, then take 64 points and predict points 65 through 192, and so on across every offset. At inference, hand it 256 points of history and ask for a 256-step forecast, and it only needs two autoregressive passes — predict the next 128, feed that back in, predict the following 128. Lock the output patch to the same 32-point length as the input, and that same forecast takes eight passes instead of two. 12 00:06:59,980 --> 00:07:15,120 [Hal Turing] That cuts compounding error, since fewer generation steps means less drift — the same problem that plagues autoregressive language models on long outputs. But there's got to be a ceiling. Why not just make the output patch enormous and get one-shot every time? 13 00:07:15,120 --> 00:07:35,786 [Dr. Ada Shannon] Because plenty of the pretraining data doesn't have room for a giant output window — monthly or yearly granularity series just don't contain enough points to fill it. So there's a practical trade-off baked in. That's also why they need patch masking during training: for every series in a batch, they sample a random cutoff somewhere in that first patch and mask everything before it. 14 00:07:35,786 --> 00:07:52,365 [Hal Turing] So that's what lets you hand it a context of, say, 217 points without it choking. Let's get into what's actually in this training corpus, because the sourcing is where it gets strange — they're not scraping text, they're assembling time-series from some unusual public sources. 15 00:07:52,365 --> 00:08:16,838 [Dr. Ada Shannon] Two real-world pillars. Google Trends gives you about half a billion time-points — search interest for roughly 22,000 head queries going back to 2007, sliced hourly through monthly. Then Wikipedia pageviews, the monster: roughly 300 billion time-points, hourly-through-monthly view counts scraped from 2012 through late 2023 across every Wikimedia page they could clean up and— 16 00:08:16,838 --> 00:08:28,402 [Hal Turing] Oh wait wait wait — 300 billion points just from pageviews, for a 200-million-parameter model? That's a wildly lopsided data-to-parameter ratio next to anything you'd see in LLM scaling. 17 00:08:28,402 --> 00:09:08,572 [Dr. Ada Shannon] Deliberate — time-series don't carry the bits-per-token density language does, so you compensate with volume. But real data alone leaves gaps, which is where synthetic data earns its keep: three million generated series, 2,048 points each, built from ARMA processes, seasonal sine-and-cosine mixtures, trends with change-points, and step functions. That's 20 percent of the training mix, with the rest split evenly across hourly, daily, weekly, and monthly real data. They also fold in the full M4 dataset, hourly and 15-minute Electricity and Traffic, and 10-minute Weather. 18 00:09:08,572 --> 00:09:13,588 [Hal Turing] And with that corpus, how does it score once you point it at data it's never touched? 19 00:09:13,588 --> 00:10:08,201 [Dr. Ada Shannon] Three benchmark groups. Monash starts at 30 datasets, filtered to 18 once you drop anything with missing values — same filter llmtime used. On scaled MAE, geometric mean, TimesFM tops the list, edging past N-BEATS, Oreshkin et al., Element AI, 2019, and beating llmtime by more than 25 percent. On Darts — eight univariate datasets heavy on seasonality — TimesFM lands within statistical significance of the best performers, seasonal ARIMA and llmtime. On the Informer ETT datasets, forecasting 96 and 192 steps of transformer temperature data, TimesFM comes out best, with PatchTST, Nie et al., IBM Research, 2022, close behind — and PatchTST is fully supervised on those exact datasets, while TimesFM has never seen them. 20 00:10:08,201 --> 00:10:10,941 [Hal Turing] What did the ablations actually nail down about these choices? 21 00:10:10,941 --> 00:11:04,579 [Dr. Ada Shannon] Four things. Scaling: 17-million, 70-million, and 200-million parameter versions, scaled MAE plotted against FLOPS — error drops monotonically as compute climbs, same curve shape as an LLM scaling law. Output patch length: longer output patches monotonically improve accuracy on the 512-step ETT task, directly validating that choice. Input patch length: sweeping 8 to 128 shows 16 and 32 are the sweet spot — smaller trains three times slower for marginal gain, bigger drifts toward encoder-decoder territory and hurts accuracy. And the synthetic-data ablation is the most telling: strip it out and Monash performance drops, but on ETT the hourly ETTh datasets barely move, since hourly is well-represented in real data, while the 15-minute ETTm datasets get noticeably worse. 22 00:11:04,579 --> 00:11:13,217 [Hal Turing] Which is the real contrast with llmtime. TimesFM claims better accuracy using a model a tiny fraction of GPT-3's size. 23 00:11:13,217 --> 00:11:28,403 [Dr. Ada Shannon] Right — you don't need a 175-billion-parameter language model doing next-token prediction on text-encoded numbers when a 200-million-parameter model trained from scratch on the right corpus gets there cheaper, and on most of these benchmarks, more accurately too. 24 00:11:28,403 --> 00:12:00,539 [Hal Turing] So let's pull the thread sitting under the whole evaluation section. The Monash comparison starts at 30 datasets and drops to 18 because they filter out anything with missing values. That sounds reasonable on its face, but missing-value series also tend to be the messier, noisier ones in these archives. Does that filter quietly select for the easier eighteen? And are the baseline numbers for those eighteen actually re-run, or borrowed wholesale from earlier papers that scored them on the full thirty? 25 00:12:00,539 --> 00:12:41,499 [Dr. Ada Shannon] Both concerns hold up. They inherited the filter itself from the llmtime paper, Gruver, Finzi, Qiu, and Wilson out of NYU, 2023 — so it's not even TimesFM's own choice, they're matching a prior protocol. But matching the protocol doesn't mean the baselines were re-scored on this exact subset. ETS, ARIMA, CatBoost, DeepAR, WaveNet numbers are pulled from the Monash archive and prior papers, not re-run by these authors. And missing values in these archives usually mean irregular collection, sensor dropout, gaps around holidays — the noisy tail that's genuinely harder to forecast. So there's a real chance the surviving eighteen skew toward the more forecastable end. 26 00:12:41,499 --> 00:12:59,565 [Hal Turing] That same question hits even harder on Darts. TimesFM gets called the 'top model,' but the paper's own text admits the standard errors aren't sharp enough to separate it from ARIMA and llmtime. Eight datasets, one series each — what does 'top' even mean beyond 'statistically tied'? 27 00:12:59,565 --> 00:13:32,165 [Dr. Ada Shannon] It means tied, not ahead. With eight single-series datasets your error bars are enormous, and the text says so directly: no clear ordering among TimesFM, seasonal ARIMA, and llmtime. Calling it the top model in a figure caption is true by point estimate, but it's doing a lot of quiet work for anyone who just skims the bar chart. And the same paragraph flags something bigger — llmtime contamination on Darts 'cannot be ruled out,' because those eight datasets show up constantly in time-series blog posts GPT-3 would've seen during pretraining. 28 00:13:32,165 --> 00:13:56,036 [Hal Turing] Oh wait wait wait — but that logic doesn't stop at llmtime. Google Trends and Wikipedia pageviews are scraped straight off the public web too. If blog-post exposure taints GPT-3's number, where's the independent audit ruling out that Monash, Darts, or Informer series — or near-duplicates — never showed up anywhere in a hundred-billion-point Trends and Wiki corpus? There isn't one in this paper. 29 00:13:56,036 --> 00:14:43,544 [Dr. Ada Shannon] And it's worse than coincidence, because every eval set here is hourly, daily, weekly, or monthly regularly-sampled data — exactly the granularity profile Trends, Wiki, and M4 were built from. So some of what reads as zero-shot generalization could just as easily be pretrained-on-statistically-similar-series under a different label. That caveat connects to PatchTST, the Nie, Nguyen, Sinthong, and Kalagnanam paper out of IBM Research, 2022 — TimesFM's architectural ancestor for patching, and still the strongest supervised baseline on Informer. It's also why TimeGPT-1, Garza and Mergenthaler-Canseco at Nixtla, 2023, is frustrating: the only other contemporaneous zero-shot contender, but closed, so nobody outside Nixtla can even ask this contamination question about it. 30 00:14:43,544 --> 00:15:30,309 [Hal Turing] There's a deeper blind spot underneath that too. The dominant training signal is what people searched for and clicked on Wikipedia — an aggregate attention signal, not a physical or transactional process. The paper never asks whether instincts learned from search behavior transfer cleanly to industrial sensors or financial order books, where the generating process is fundamentally different. And the Impact Statement's bias discussion is thin — it waves away bias risk by noting the model uses no covariates, but the corpus itself is heavily US- and English-skewed, which can quietly degrade accuracy for underrepresented regions without any covariate needed to inherit it. 31 00:15:30,309 --> 00:16:28,916 [Dr. Ada Shannon] Same instinct applies to the scaling ablation — 17, 70, 200 million parameters, one mixture, 1.5 million steps. They cite Hoffmann and colleagues' Chinchilla paper, DeepMind, 2022, as the sharper compute-optimal framing, but never actually run that analysis. Which makes TimesFM-3's jump to 330 million parameters and over a trillion timepoints this year a tacit admission — if 200 million was already compute-optimal, why keep scaling? TimesFM-3 also quietly walks back the original paper's cleanest selling point: zero dataset-specific setup. Now it wants past-future covariates supplied correctly, and its average-rank plots against Chronos-2, Ansari and colleagues at Amazon, 2024, show no confidence intervals or per-task breakdowns — a few big wins could be masking plenty of near-ties, the exact aggregation problem the original paper raised about Monash. 32 00:16:28,916 --> 00:17:20,092 [Hal Turing] So for a practitioner: if your series looks like the pretraining data — regular granularity, retail- or search-adjacent behavior — TimesFM zero-shot is a legitimately cheap first pass before building anything custom. If you're in a genuinely novel domain or need multivariate covariates done right, treat the accuracy claims as a hypothesis to benchmark, not a guarantee. The field is clearly heading toward foundation models absorbing what used to be per-dataset engineering — the honest version of that story just needs contamination audits and confidence intervals alongside the leaderboard rank. That's TimesFM and TimesFM-3 — real progress on zero-shot forecasting, with the evaluation rigor still catching up. Thanks for listening, everyone.