---
title: Which time-series foundation model should I use? A decision guide
description: When to pick Chronos-2, TimesFM 2.5, Toto 2, TiRex-2, PatchTST-FM r2 or FlowState r1, their limits, and when to ensemble instead.
date: 2026-10-07
updated: 2026-10-07
summary: Pick by your data, not by leaderboard rank: the top single models in our GIFT-Eval runs are within about 2% of each other. Use Chronos-2 or TiRex-2 if you have covariates, TimesFM 2.5 for very long histories, Toto 2 for short horizons and multivariate or observability data, and PatchTST-FM r2 for long horizons. If you cannot tell, route or ensemble mode decides for you.
tags: forecasting, models, chronos-2, timesfm, toto, tirex, patchtst, flowstate
draft: false
---

## Key facts

- On GIFT-Eval (97 configurations, CRPS ratio to seasonal naive, lower is better), our single-model runs score Chronos-2 0.4828, Toto 2 (313M) 0.4842, PatchTST-FM r2 0.4869, TimesFM 2.5 0.4923 and FlowState r1 0.5221; TiRex-2 scores 0.4781 on the published leaderboard.
- Only Chronos-2 and TiRex-2 accept covariates in the Ephemeris panel.
- TimesFM 2.5 reads up to 16,384 points of history, the longest context in the panel; TiRex-2 and FlowState r1 read 2,048.
- TimesFM 2.5 and FlowState r1 are univariate only in the panel; the other four forecast multivariate series.
- TiRex-2 is used by route and ensemble mode up to 320 steps; name it in explicit mode for horizons up to 512.
- Five of the six are Apache-2.0; PatchTST-FM r2 is listed as OpenMDW-1.0.
- A different single model was best on each of four benchmarks we ran, which is why route and ensemble mode exist.

## How do I choose a forecasting model?

Start from three questions about your data: do you have covariates, how long is your history, and how far ahead do you forecast. Those three rule most models in or out before accuracy comes into it.

A covariate is an extra series that helps explain the one you forecast, such as price or a holiday flag. Context is the history the model reads. Horizon is how many steps ahead you forecast. Multivariate means forecasting several related series jointly.

All scores below are ratios to a seasonal-naive forecast (repeat the last season), lower is better. CRPS scores the whole uncertainty band; MASE scores the median. They are our own runs with each benchmark's harness, not leaderboard submissions, except where marked. Method: [/benchmarks](/benchmarks).

Every section ends with how to call the model by name. In Ephemeris that is explicit mode: `"mode": "explicit", "model": "<apiName>"`. Call `GET /api/v1/models` (or the MCP tool `list_models`) first: the live panel, horizon limits and prices come from there.

## Chronos-2

Chronos-2 is Amazon's 120M-parameter encoder model for univariate, multivariate and covariate-informed forecasting. It has the best single-model GIFT-Eval scores of the five we ran ourselves: CRPS 0.4828, MASE 0.7037.

Pick it when:

- you have covariates, especially past 320 steps, where route mode sends covariate requests to it;
- you want a strong general default and speed matters: it is the fastest model in the panel;
- your series are multivariate.

Skip it when your history is longer than 8,192 points and the old data matters (TimesFM 2.5 reads twice as much), or when your data is observability metrics, where Toto 2 alone matched the full ensemble on BOOM.

Limits: context cap 8,192 points. Licence Apache-2.0. The [Chronos-2 model card](https://huggingface.co/amazon/chronos-2) lists past-only and known-future covariates, both real-valued and categorical.

Call it with `"mode": "explicit", "model": "chronos2"`. Details: [/models/chronos-2](/models/chronos-2).

## TimesFM 2.5

TimesFM 2.5 is Google Research's 200M-parameter decoder-only forecaster. It reads up to 16,384 points of history, the longest context in the panel, and was the best single model in our panel on fev-bench. On GIFT-Eval it scores CRPS 0.4923, MASE 0.7096.

Pick it when:

- you have long histories with slow seasonality, such as many months of hourly data;
- your data looks more like fev-bench than GIFT-Eval.

Skip it when you have covariates or multivariate series: in the Ephemeris panel it is univariate and takes no covariates. The TimesFM library adds covariate support through XReg, per the [TimesFM README](https://github.com/google-research/timesfm), so self-hosting is the option if you need TimesFM with covariates.

Limits: context cap 16,384. Licence Apache-2.0. The [model card](https://huggingface.co/google/timesfm-2.5-200m-pytorch) notes that the checkpoint "is not an officially supported Google product". Its successor TimesFM-3 scores better on the published GIFT-Eval leaderboard (CRPS 0.4557) but its open weights are licensed for non-commercial, non-production use only; the README says commercial use is permitted through Google Cloud services such as BigQuery ML.

Call it with `"mode": "explicit", "model": "timesfm25"`. Details: [/models/timesfm-2-5](/models/timesfm-2-5).

## Toto 2

Toto 2 (313M) is Datadog's forecaster, trained heavily on observability metrics. In our calibration it is the strongest single member on short horizons, and it was the best single model in our panel on TIME and BOOM. On GIFT-Eval it scores CRPS 0.4842, MASE 0.7038, within 0.3% of Chronos-2 in our runs.

Pick it when:

- you forecast short horizons;
- your series are multivariate: route mode sends multivariate series to it;
- your data is infrastructure or application metrics. On BOOM, Datadog's observability benchmark, it matched the full ensemble on its own.

Skip it when you have covariates (it takes none in the panel) or very long histories (context cap 4,096).

Limits: context cap 4,096. Licence Apache-2.0. The [model card](https://huggingface.co/Datadog/Toto-2.0-313m) lists a family from 4M to 2.5B parameters; Ephemeris serves the 313M size. The 2.5B size scores CRPS 0.4759 on the published GIFT-Eval leaderboard, against 0.4842 for the 313M size in our own run; the two numbers come from different runs, so treat the gap as indicative if you are considering self-hosting the larger size.

Call it with `"mode": "explicit", "model": "toto2-313m"`. Details: [/models/toto-2](/models/toto-2).

## TiRex-2

TiRex-2 is NXAI's xLSTM forecaster: 38M parameters univariate, 82M multivariate, with covariate support. It scores CRPS 0.4781 on the published GIFT-Eval leaderboard, ahead of Chronos-2 (0.4854) on the same leaderboard. We have not run it ourselves on GIFT-Eval, so this number is not directly comparable with our own runs above.

Pick it when:

- you have covariates and a horizon of 320 steps or less;
- you want a small model: it is the second smallest in the panel;
- your series are multivariate.

Skip it when your history is longer than 2,048 points and the old data matters, or your horizon is past 320 steps and you are not naming it explicitly.

Limits: context cap 2,048. Its head emits 320 steps per pass; longer horizons are rolled out with the steps already forecast fed back as missing values. Route and ensemble mode use it up to 320 steps; explicit mode goes up to the API's 512-step limit. Licence Apache-2.0 per the [TiRex-2 model card](https://huggingface.co/NX-AI/TiRex-2), which also lists past and future-known covariates. The original TiRex is under a different licence, the NXAI community licence, per the [TiRex repository](https://github.com/NX-AI/tirex).

Call it with `"mode": "explicit", "model": "tirex2"`. Details: [/models/tirex-2](/models/tirex-2).

## PatchTST-FM r2

PatchTST-FM r2 is IBM Granite's 385M-parameter patch-transformer foundation model, the largest in the panel. It is the strongest single member at GIFT-Eval's long horizons in our run. Over all 97 GIFT-Eval configurations it scores CRPS 0.4869, MASE 0.715.

Pick it when:

- you forecast long horizons;
- your series are multivariate and you have no covariates.

Skip it when you have covariates, or when your horizons are short, where Toto 2 is stronger.

Limits: context cap 8,192. No covariates. Licence listed as OpenMDW-1.0; the [model card](https://huggingface.co/ibm-granite/granite-timeseries-patchtst-fm-r2) describes it as dual-licensed under OpenMDW-1.0 and Apache-2.0. The card also says IBM "will not be maintaining this code going forward", which matters if you plan to self-host.

Call it with `"mode": "explicit", "model": "patchtst-fm-r2"`. Details: [/models/patchtst-fm](/models/patchtst-fm).

## FlowState r1

FlowState r1 is IBM Granite's 9M-parameter state-space forecaster, the smallest in the panel. It adapts to the sampling rate of the series. It has the weakest GIFT-Eval scores of the six: CRPS 0.5221, MASE 0.7506.

Pick it when:

- you want to self-host a very small model and can accept lower accuracy.

Skip it when accuracy is the priority: every other model in the panel scored better on GIFT-Eval.

Limits: context cap 2,048, univariate, no covariates. Licence Apache-2.0. The [model card](https://huggingface.co/ibm-granite/granite-timeseries-flowstate-r1) says it supports zero-shot forecasting only and that the sampling rate is set with a scale factor (for example 0.25 for 15-minute data, 1.0 for hourly) when you run it yourself.

Call it with `"mode": "explicit", "model": "flowstate-r1"`. Details: [/models/flowstate](/models/flowstate).

## Which model fits my situation?

| Your situation | Pick |
|---|---|
| Covariates, horizon up to 320 steps | TiRex-2 or Chronos-2 |
| Covariates, horizon over 320 steps | Chronos-2 |
| History longer than 8,192 points, univariate, no covariates | TimesFM 2.5 |
| Short horizon | Toto 2 |
| Multivariate series, no covariates | Toto 2 |
| Infrastructure or application metrics | Toto 2 |
| Long horizon, no covariates | PatchTST-FM r2 |
| General default, one model, speed matters | Chronos-2 |
| Smallest model to self-host | FlowState r1, then TiRex-2 |
| Not sure | Route mode |
| Calibrated uncertainty matters more than cost | Ensemble mode |

## When should I use route or ensemble mode instead of choosing?

Use them when you cannot answer the three questions above, or your series vary. Choosing one model is a bet that your data resembles where that model is strong; the benchmarks say that bet changes from dataset to dataset.

- Route mode (`"mode": "route"`) picks a model suited to the data and usually bills one model run, like explicit mode. Pass `freq` to help it.
- Ensemble mode (`"mode": "ensemble"`) runs every compatible model and combines them, weighting the more accurate ones more. It costs the sum of the models that ran. In our runs the best single model was 3.6% worse than the ensemble on GIFT-Eval CRPS, 1.6% on TIME, 1.1% on fev-bench, and level on BOOM.
- Explicit mode is for when you have a reason: a constraint above, or you have measured a model on your own data.

```python
import os
import requests

r = requests.post(
    "https://ephemeris.cascade.industries/api/v1/forecast",
    headers={"Authorization": f"Bearer {os.environ['EPHEMERIS_API_KEY']}"},
    json={
        "mode": "explicit",
        "model": "chronos2",
        "series": [{
            "values": [210, 198, 225, 240, 232, 260, 275, 268],
            "freq": "D",
            "covariates": {
                "past": {"promo": [0, 0, 1, 1, 0, 1, 1, 0]},
                "future": {"promo": [1, 1, 0, 0]},
            },
        }],
        "horizon": 4,
        "quantiles": [0.1, 0.5, 0.9],
    },
    timeout=60,
)
r.raise_for_status()
print(r.json()["forecasts"][0]["quantiles"])
```

Naming a model without covariate support while sending covariates is an error in explicit mode, as is a horizon past the model's `max_horizon`.

The honest test is your own data: hold out the last stretch of history, forecast it with two or three candidates, and compare. That beats any guide, including this one.

Ephemeris serves all six models behind one endpoint, so trying two or three on your own holdout is a change of the `model` field. The [models page](/models) lists each one's API name, and new accounts get a small free credit.

## FAQ

### What is the best time-series foundation model?

None of the six wins on every benchmark: Chronos-2 led our GIFT-Eval runs, Toto 2 led TIME and BOOM, and TimesFM 2.5 led fev-bench. TimesFM-3 scores better than all of them on the published GIFT-Eval leaderboard (CRPS 0.4557) but its open weights are non-commercial.

### Which forecasting model supports covariates?

In the Ephemeris panel, Chronos-2 and TiRex-2. Route and ensemble mode narrow to those two when you send covariates.

### Which model handles the longest history?

TimesFM 2.5, with a context cap of 16,384 points. Chronos-2 and PatchTST-FM r2 read 8,192, Toto 2 4,096, and TiRex-2 and FlowState r1 2,048.

### Which model is best for long forecast horizons?

In our run, PatchTST-FM r2 was the strongest single model at GIFT-Eval's long horizons. TiRex-2 emits 320 steps per pass, so route and ensemble mode leave it out past 320 steps.

### Should I pick a model or use an ensemble?

Pick a model if a constraint (covariates, history length, horizon) decides it or you have measured it on your data. Otherwise use route mode for the lowest cost or ensemble mode for the best calibrated forecast.

## Related

- [Chronos-2 vs TimesFM 2.5 vs Toto 2 vs TiRex-2](/blog/chronos-2-vs-timesfm-vs-toto-vs-tirex)
- [Does ensembling forecasting foundation models help?](/blog/does-ensembling-forecasting-models-help)
- [A guide to time-series foundation models](/blog/time-series-foundation-models-guide)
- [Forecast store sales with promotions](/blog/forecast-store-sales-with-promotions)
- [TimeGPT alternatives in 2026](/blog/timegpt-alternatives)
- [Model pages](/models)
