● AI Post Transformers — Episode Companion

TimesFM: A Decoder-Only Foundation Model for Zero-Shot Time-Series Forecasting

A 200M-parameter decoder-only transformer, pretrained once on ~100B timepoints scraped mostly from Google Trends and Wikipedia pageviews, claims to forecast datasets it has never seen — at accuracy close to models trained specifically for each one.

arXiv:2310.10688 ↗ Das, Kong, Sen, Zhou — Google Research Oct 2023, rev. Apr 2024 200M params ~100B timepoints

Zero-Shot Forecasting Pipeline

One pretrained model, no gradient updates at inference. Compare against the classical pattern of training or fine-tuning a bespoke model per dataset.

Why This Is a Weird Bet

Unlike text, time series share no vocabulary across domains.

At a Glance

200M
parameters
~100B
training timepoints
2
autoregressive passes for a 256-step forecast
+25%
beats llmtime on Monash sMAE
20%
of pretraining mix is synthetic
4
authors, Google Research
Central claim: a purpose-built 200M model beats feeding raw numbers as text into a 175B-parameter LLM — at a fraction of the inference cost.

From Pixels to Patches to Time-Series Patches

TimesFM borrows patching from Vision Transformer (2020) via PatchTST's time-series adaptation (ICLR 2023) — a contiguous chunk of points becomes one input token, not one timestep.

Causal Patch Tokenization

Hover a patch to see what it represents — a little shape (rising edge, dip, seasonal bump), not a single value.

Why Output Patches Are Longer Than Input Patches

Toggle to compare inference cost: a 256-step forecast takes 2 passes with a 128-length output patch, or 8 passes if output length is locked to the 32-length input patch. Fewer passes means less compounding autoregressive drift.

Where ~100 Billion Timepoints Come From

Real time-series data is scarce online compared to text, so volume comes from search-interest and pageview logs, not curated datasets.

Real-World Pillars

1

Google Trends~0.5B timepoints — search interest for ~22,000 head queries, back to 2007, hourly through monthly.

2

Wikipedia Pageviews~300B timepoints — hourly-through-monthly view counts, 2012–late 2023, across cleaned Wikimedia pages. The dominant source by volume.

3

Synthetic Series3M generated series × 2,048 points — ARMA processes, seasonal sine/cosine mixtures, trend change-points, step functions. 20% of the training mix, engineered to plug real-data coverage gaps.

4

Curated Benchmarks Folded InFull M4 dataset, hourly + 15-minute Electricity and Traffic, 10-minute Weather.

Deliberate imbalance: 300B points feeding a 200M-parameter model is a wildly lopsided data-to-parameter ratio next to LLM scaling norms — time series carry far less bits-per-token density than text, so volume compensates.

Monash Archive — Scaled MAE (Geometric Mean, lower is better)

18 datasets survive after filtering out series with missing values — TimesFM tops the list, edging past N-BEATS and beating llmtime by >25%.

Informer ETT — Forecast Error at 96 / 192 Steps

TimesFM (zero-shot) vs. PatchTST (fully supervised on these exact datasets) vs. the original Informer transformer.

Scaling Ablation

17M → 70M → 200M parameters. Error drops monotonically with compute — same curve shape as an LLM scaling law. Chinchilla-style compute-optimality was never run.

Input / Output Patch-Length Ablations

Input patch length 8–128: 16 and 32 are the sweet spot. Output patch length: longer monotonically improves 512-step ETT accuracy — the ablation directly validating the long-output-patch design choice.

The Monash Filter: 30 → 18 Datasets

Series with missing values are dropped — inherited from llmtime's protocol, not re-derived. Missing values usually mean irregular collection or sensor dropout: the noisier, harder-to-forecast tail. Does the filter quietly select for the easier eighteen?
Also flagged: baseline numbers (ETS, ARIMA, CatBoost, DeepAR, WaveNet) are pulled from prior papers scored on the full 30 — not re-run on this subset.

Darts: "Top Model" or Statistical Tie?

8 single-series datasets → enormous error bars. The paper's own text admits no clear ordering among TimesFM, seasonal ARIMA, and llmtime.

Contamination Risk, Widened

If llmtime's Darts numbers might be contaminated by GPT-3 pretraining exposure, TimesFM's own corpus — scraped from the same public web — has no independent audit ruling out overlap with Monash/Darts/Informer series either.

Granularity Overlap: Eval Sets vs. Pretraining Sources

Every eval set here is hourly/daily/weekly/monthly regularly-sampled data — exactly the granularity profile Trends, Wiki, and M4 were built from. Some "zero-shot generalization" could be pretrained-on-statistically-similar-series under a different label.
TimesFM-3 tacit admission: the follow-up jumps to 330M params and >1T timepoints, and now wants correctly-supplied past-future covariates — quietly walking back the original paper's cleanest selling point of zero dataset-specific setup.

References