---
title: Time-series foundation models: a practical guide (2026)
description: What a time-series foundation model is, how it differs from ARIMA, Prophet and LLMs, and how to choose, run and trust one.
date: 2026-10-07
updated: 2026-10-07
summary: A time-series foundation model (TSFM) is a neural network pretrained on a large collection of series that forecasts a new series from its history alone, with no training on your data. Most return quantiles, so you get an uncertainty band as well as a line. On GIFT-Eval (97 configurations) the leading open-weights TSFMs score a CRPS of about 0.48 to 0.49 as a ratio to seasonal naive, roughly half the error of that naive baseline. You can run them yourself from Hugging Face or call them through a hosted API.
tags: forecasting, foundation-models, guide
draft: false
---

## Key facts

- A time-series foundation model (TSFM) is pretrained once on many series and then forecasts unseen series "zero-shot": you send history, it returns a forecast, nothing is fitted to your data.
- Chronos-2 (Amazon, 120M parameters), TimesFM 2.5 (Google, 200M), Toto 2 (Datadog, 313M), TiRex-2 (NXAI, 38M univariate), PatchTST-FM r2 (IBM, 385M) and FlowState r1 (IBM, 9M) are open-weights TSFMs published on Hugging Face.
- On GIFT-Eval (97 configurations, CRPS as a ratio to seasonal naive, lower is better), the published leaderboard has TimesFM-3 at 0.4557, Toto-2.0 2.5B at 0.4759, TiRex-2 at 0.4781 and Chronos-2 at 0.4854; the Ephemeris ensemble scores 0.4662 in our run with the official harness.
- TimesFM-3, the best scorer in that table, has a non-commercial licence; five of the six models in the Ephemeris panel are Apache-2.0 and one (PatchTST-FM r2) is OpenMDW-1.0.
- Most TSFMs output quantiles (for example the 10th, 50th and 90th percentiles), so each forecast carries its own uncertainty band.
- Only some TSFMs accept covariates (known future inputs such as a holiday flag or a planned price): in the Ephemeris panel, Chronos-2 and TiRex-2 do.
- Context limits differ by model: from 2,048 points (TiRex-2, FlowState r1) to 16,384 points (TimesFM 2.5).

## What is a time-series foundation model?

A time-series foundation model is a forecasting network trained on a very large and varied set of time series, so that it can forecast a series it has never seen without any further training.

You give it the recent history of one series (its *context*) and the number of steps you want (the *horizon*). It returns the forecast in a single forward pass. There is no fitting step, no hyperparameter search and no per-series model to store.

The idea mirrors language models: pretrain once on a broad corpus, then reuse everywhere. The corpus is numbers instead of text. GIFT-Eval, a common benchmark, ships with a separate pretraining set of about 230 billion data points so that models can be trained without seeing the test series ([Aksu et al., 2024](https://arxiv.org/abs/2410.10393)).

## How is a TSFM different from ARIMA, ETS or Prophet?

Classical models are fitted to each series; a TSFM is fitted once and then only reads your series.

ARIMA and ETS (exponential smoothing) estimate a handful of parameters per series from its own history. Libraries such as [StatsForecast](https://github.com/Nixtla/statsforecast) fit them automatically and fast. [Prophet](https://facebook.github.io/prophet/) fits an additive model with trend, yearly, weekly and daily seasonality and holiday effects.

That per-series fitting has real strengths. It is cheap on a CPU, the parameters are interpretable, and it behaves well on long, clean, regular series. If your series look like that, a well-tuned AutoETS is a strong baseline and you should keep it.

A TSFM trades those strengths for breadth. It has seen patterns from many domains, so in our experience it tends to do better on short histories, irregular seasonality, and series where you do not have time to tune a model each. It also needs a GPU to be fast, and you cannot read its reasoning off a coefficient table.

## How is a TSFM different from an LLM?

A TSFM is trained to predict numbers from numbers; a large language model is trained to predict text.

You can paste a series into a chat model and ask for a forecast, and it will give you one. Research finds this works poorly or inconsistently compared with purpose-built models, for reasons that start with how numbers are split into tokens. We cover the evidence in [why language models are bad at forecasting numbers](/blog/why-language-models-are-bad-at-time).

The useful pattern is to combine them: the LLM gathers the data and reads the context, the forecasting model produces the numbers. See [the split that works](/blog/llm-for-context-forecaster-for-numbers).

## What is a probabilistic or quantile forecast?

A quantile forecast gives several lines instead of one, each with a stated probability of the actual value falling below it.

The 0.9 quantile is a value the actual should fall below 90% of the time. The 0.5 quantile is the median. The band between the 0.1 and 0.9 quantiles should contain the actual value about 80% of the time; when it does, the forecast is *calibrated*.

Accuracy of such forecasts is usually scored with CRPS (continuous ranked probability score), which rewards a band that is both narrow and correct. MASE (mean absolute scaled error) scores the median line against the error of a naive forecast. Benchmarks often report both as ratios to *seasonal naive*, the forecast that repeats last season's values. A ratio of 0.47 means 47% of the naive method's error.

Why the band matters for decisions is the subject of [agents that decide under uncertainty](/blog/agents-that-decide-under-uncertainty).

## Which time-series foundation model should I choose?

Start with the licence and the capabilities your data needs, then look at benchmark scores.

| Model | Publisher | Parameters | Licence | Multivariate | Covariates | Context cap |
|---|---|---|---|---|---|---|
| Chronos-2 | Amazon | 120M | Apache-2.0 | Yes | Yes | 8,192 |
| TimesFM 2.5 | Google Research | 200M | Apache-2.0 | No | No | 16,384 |
| Toto 2 (313M) | Datadog | 313M | Apache-2.0 | Yes | No | 4,096 |
| TiRex-2 | NXAI | 38M / 82M | Apache-2.0 | Yes | Yes | 2,048 |
| PatchTST-FM r2 | IBM Granite | 385M | OpenMDW-1.0 | Yes | No | 8,192 |
| FlowState r1 | IBM Granite | 9M | Apache-2.0 | No | No | 2,048 |

A few practical rules:

- If you have known future inputs (promotions, holidays, a weather forecast), you need a model with covariate support.
- If you have several related series that move together, a multivariate model can use that.
- If you have long histories with slow seasonality, a long context helps.
- If you are building a commercial product, check the licence before the score: the top GIFT-Eval scorer, TimesFM-3, is non-commercial.

The single-model scores are close. In our GIFT-Eval run (97 configurations, CRPS ratio to seasonal naive), Chronos-2 scored 0.4828, Toto 2 0.4842, PatchTST-FM r2 0.4869, TimesFM 2.5 0.4923 and FlowState r1 0.5221. In our experience, the gaps between the top four are smaller than the differences you will see across your own series. Our [head-to-head comparison](/blog/chronos-2-vs-timesfm-vs-toto-vs-tirex) and the [decision guide](/blog/which-forecasting-model-should-i-use) go further.

## Does combining several models help?

Usually a little, and it costs more. In our runs, the Ephemeris ensemble beat its best single member by 3.6% CRPS on GIFT-Eval, 1.6% on TIME and 1.1% on fev-bench, and matched it on BOOM.

Whether that margin is worth paying for several model runs depends on what an error costs you. We go through the numbers in [does ensembling forecasting models help?](/blog/does-ensembling-forecasting-models-help).

## Should I self-host a TSFM or use a hosted API?

Self-host if you have GPUs, steady volume, one or two models you trust, and an engineer to own it. Use an API if you want several models without running them, or if your volume is small or bursty.

Self-hosting is straightforward for a single model. Chronos-2, for example, installs with `pip install "chronos-forecasting>=2.0"` and loads from `amazon/chronos-2` on [Hugging Face](https://huggingface.co/amazon/chronos-2). You control the data path and pay only for compute.

The cost shows up with several models. Each one has its own package and dependency pins, so running a panel side by side usually means separate environments, GPU memory for each, and keeping checkpoints current.

A hosted API removes that work but adds a per-call price, a network round trip and a third party that sees your numbers. Hosted options include Ephemeris (six open-weights models over REST and MCP) and Nixtla's TimeGPT; see [TimeGPT alternatives](/blog/timegpt-alternatives) for a fair comparison.

## How do I run a forecast through an API?

Send the history, the frequency, the horizon and the quantiles you want. With Ephemeris, a minimal request in Python looks like this:

```python
import os
import requests

resp = requests.post(
    "https://ephemeris.cascade.industries/api/v1/forecast",
    headers={"Authorization": f"Bearer {os.environ['EPHEMERIS_API_KEY']}"},
    json={
        "mode": "route",
        "series": [{
            "values": [212, 198, 205, 230, 251, 244, 219, 208, 201, 236, 262, 249],
            "freq": "D",
        }],
        "horizon": 7,
        "quantiles": [0.1, 0.5, 0.9],
    },
    timeout=60,
)
resp.raise_for_status()
forecast = resp.json()["forecasts"][0]["quantiles"]
print(forecast["0.1"], forecast["0.5"], forecast["0.9"], sep="\n")
```

`route` mode picks one suited model, `ensemble` runs every compatible model and combines them, and `explicit` runs the model you name. In practice send more history than this toy example: a few dozen points at least, and more is better.

Each model run costs `rate x slots x ceil(context / 1024) x ceil(horizon / 64)` millicredits, so one series with up to 1,024 points of context and up to 64 steps of horizon costs exactly one model's rate. Rates are on [/pricing](/pricing).

## What are the honest trade-offs?

TSFMs are a good default, not a guarantee. Keep these in mind:

- They know nothing you do not send. A model cannot see a promotion, a price change or an outage unless you pass it as a covariate, and only some models accept covariates.
- They extrapolate history. A structural break (a new product, a pandemic, a pricing change) is invisible until it appears in the data.
- Benchmarks average many datasets. Your series may favour a different model, or a classical one. Backtest on your own data before you rely on a model.
- Bands can be miscalibrated on unusual data. Check how often actuals fall outside the 80% band on a held-out period.
- Short or very noisy series give wide bands. That is the model being honest, not failing.

## The series

This guide is the hub of a series. Read in roughly this order:

- [Why language models are bad at forecasting numbers](/blog/why-language-models-are-bad-at-time): tokenization, calibration and cost, with the research.
- [The split that works: the LLM reads the context, a forecasting model does the numbers](/blog/llm-for-context-forecaster-for-numbers): the architecture for agents.
- [Agents that decide under uncertainty](/blog/agents-that-decide-under-uncertainty): turning quantiles into stock levels, alerts and budgets.
- [A forecasting tool for AI agents over MCP](/blog/forecasting-tool-for-ai-agents-mcp): connecting a forecaster to Claude, Cursor and other MCP clients.
- [Forecasting in agent frameworks](/blog/forecasting-in-agent-frameworks): wiring a forecast tool into common agent frameworks.
- [Chronos-2 vs TimesFM vs Toto vs TiRex](/blog/chronos-2-vs-timesfm-vs-toto-vs-tirex): the open-weights models compared.
- [Does ensembling forecasting models help?](/blog/does-ensembling-forecasting-models-help): what combining models buys you, in numbers.
- [TimeGPT alternatives](/blog/timegpt-alternatives): hosted and self-hosted options compared.
- [Which forecasting model should I use?](/blog/which-forecasting-model-should-i-use): a decision guide by data type and constraint.
- [Forecast store sales with promotions](/blog/forecast-store-sales-with-promotions): retail demand with covariates.
- [Forecast electricity load and solar](/blog/forecast-electricity-load-and-solar): energy series with weather inputs.
- [Crypto volatility ranges](/blog/crypto-volatility-ranges): forecasting ranges, not prices.
- [Capacity planning from ops metrics](/blog/capacity-planning-ops-metrics): infrastructure metrics and alert thresholds.
- [Forecast sensor and IoT readings](/blog/forecast-sensor-iot-readings): high-frequency device data.

## FAQ

### What does zero-shot forecasting mean?

Zero-shot means the model forecasts a series it was never trained on, using only the history you send at request time. There is no fitting or fine-tuning on your data.

### Is a time-series foundation model always better than ARIMA or ETS?

No. On long, clean, regular series a tuned classical model can match or beat a TSFM at a fraction of the compute. Backtest both on your own data.

### Can I use ChatGPT or Claude to forecast a time series?

They will produce a forecast, but research shows language models are inconsistent at numeric forecasting and poorly calibrated. A better pattern is to let the LLM prepare the data and context and call a forecasting model for the numbers.

### Which time-series foundation model is most accurate?

On GIFT-Eval (97 configurations, CRPS ratio to seasonal naive), TimesFM-3 scores best at 0.4557 on the published leaderboard but has a non-commercial licence. Among commercially licensed options, the gaps between the leading open-weights models are small, so test on your data.

### Do time-series foundation models need a GPU?

They run fastest on a GPU, and larger ones are slow on a CPU. Small models such as FlowState r1 (9M parameters) and TiRex-2 (38M univariate) are lighter to run.

## Related

- [Why language models are bad at forecasting numbers](/blog/why-language-models-are-bad-at-time)
- [Chronos-2 vs TimesFM vs Toto vs TiRex](/blog/chronos-2-vs-timesfm-vs-toto-vs-tirex)
- [Which forecasting model should I use?](/blog/which-forecasting-model-should-i-use)
- [Does ensembling forecasting models help?](/blog/does-ensembling-forecasting-models-help)
- [Agents that decide under uncertainty](/blog/agents-that-decide-under-uncertainty)
- [Model panel and licences](/models), [benchmarks](/benchmarks) and [API reference](/llms-full.txt)
