---
title: Chronos-2 vs TimesFM 2.5 vs Toto 2 vs TiRex-2: one harness, same data
description: Six open-weights forecasting models compared on GIFT-Eval and TIME with one harness: size, licence, context, covariates and CRPS/MASE scores.
date: 2026-10-07
updated: 2026-10-07
summary: On GIFT-Eval (97 configurations, our runs with the official harness), Chronos-2 has the best single-model score of those we ran at CRPS 0.4828, with Toto 2 (313M) close behind at 0.4842; PatchTST-FM r2 scores 0.4869 and TimesFM 2.5 0.4923. The gaps are small, and no single model leads on every benchmark we ran. TimesFM-3 scores better than all of them (0.4557) but its open weights carry a non-commercial licence.
tags: forecasting, benchmarks, chronos-2, timesfm, toto, tirex
draft: false
---

## Key facts

- On GIFT-Eval (97 configurations, CRPS ratio to seasonal naive, lower is better), our single-model runs score Chronos-2 0.4828, Toto 2 (313M) 0.4842, PatchTST-FM r2 0.4869, TimesFM 2.5 0.4923 and FlowState r1 0.5221.
- On the same GIFT-Eval MASE ratio, Chronos-2 scores 0.7037, Toto 2 (313M) 0.7038, TimesFM 2.5 0.7096, PatchTST-FM r2 0.715 and FlowState r1 0.7506.
- TiRex-2 scores CRPS 0.4781 on the published GIFT-Eval leaderboard; we have no single-model run of our own for it.
- TimesFM-3 scores CRPS 0.4557 / MASE 0.6668 on GIFT-Eval, ahead of every model here, but its open weights are licensed for non-commercial use only.
- Chronos-2, TimesFM 2.5, Toto 2, TiRex-2 and FlowState r1 are Apache-2.0; PatchTST-FM r2 is listed as OpenMDW-1.0.
- Context caps as served by Ephemeris: TimesFM 2.5 16,384 points, Chronos-2 and PatchTST-FM r2 8,192, Toto 2 4,096, TiRex-2 and FlowState r1 2,048.
- Only Chronos-2 and TiRex-2 take covariates in the Ephemeris panel.

## What are these models?

They are time-series foundation models: neural networks pretrained on large collections of series so they can forecast a new series zero-shot, with no training on your data. You send history, you get a forecast back.

All six return quantiles, not just one line. A quantile is a level the future value should fall below with a given probability: the 0.9 quantile should be exceeded about 10% of the time. A set of quantiles gives you an uncertainty band.

| Model | Publisher | Parameters | Licence | Context cap | Multivariate | Covariates |
|---|---|---|---|---|---|---|
| Chronos-2 | Amazon | 120M | Apache-2.0 | 8,192 | Yes | Yes |
| TimesFM 2.5 | Google Research | 200M | Apache-2.0 | 16,384 | No | No (in Ephemeris) |
| Toto 2 (313M) | Datadog | 313M | Apache-2.0 | 4,096 | Yes | No |
| TiRex-2 | NXAI | 38M univariate, 82M multivariate | Apache-2.0 | 2,048 | Yes | Yes |
| PatchTST-FM r2 | IBM Granite | 385M | OpenMDW-1.0 | 8,192 | Yes | No |
| FlowState r1 | IBM Granite | 9M | Apache-2.0 | 2,048 | No | No |

Context cap is the number of most recent points the model reads; longer histories are truncated. Multivariate means the model forecasts several related series (variates) jointly. Covariates are extra input series that help explain the target, such as a price or a holiday flag; "future" covariates are ones you know in advance over the forecast period.

Two notes on the table. The TimesFM library itself has covariate support through XReg, added for TimesFM 2.5 according to the [TimesFM README](https://github.com/google-research/timesfm); Ephemeris does not serve that path. And the [PatchTST-FM r2 model card](https://huggingface.co/ibm-granite/granite-timeseries-patchtst-fm-r2) describes it as dual-licensed under OpenMDW-1.0 and Apache-2.0.

## How do you read CRPS and MASE ratios?

Both scores are ratios to seasonal naive, and lower is better. Seasonal naive is the simplest sensible forecast: repeat the values from the last season (yesterday's hours for hourly data with a daily cycle, for example). A ratio of 0.70 means the model's error is 30% lower than that baseline.

MASE (mean absolute scaled error) measures the point forecast. Here it scores the median: how far, on average, the median forecast lands from what actually happened.

CRPS (continuous ranked probability score) measures the whole forecast distribution. It rewards bands that are both narrow and correct, and penalises bands that are too wide or miss. For a forecast with no spread it reduces to absolute error. Our CRPS is computed from nine quantiles.

Each benchmark has many configurations (a dataset at a frequency and horizon). Each configuration's score is divided by seasonal naive on the same window, then the ratios are averaged geometrically, the way the benchmarks report them. These are our own runs with each benchmark's harness, not leaderboard submissions. The full method is on [/benchmarks](/benchmarks).

## Which model scores best on GIFT-Eval?

Of the five models we ran ourselves, Chronos-2 has the lowest single-model CRPS and MASE on GIFT-Eval, with Toto 2 (313M) almost tied. TiRex-2, the sixth, scores better still on the published leaderboard. The spread between the top four is about 2% of CRPS.

| Model | GIFT-Eval CRPS ratio | GIFT-Eval MASE ratio | Source |
|---|---|---|---|
| TimesFM-3 | 0.4557 | 0.6668 | Published leaderboard; non-commercial licence |
| Toto-2.0 2.5B | 0.4759 | not reported | Published leaderboard; not served by Ephemeris |
| TiRex-2 | 0.4781 | not reported | Published leaderboard |
| Chronos-2 | 0.4828 | 0.7037 | Our run |
| Toto 2 (313M) | 0.4842 | 0.7038 | Our run |
| PatchTST-FM r2 | 0.4869 | 0.715 | Our run |
| TimesFM 2.5 | 0.4923 | 0.7096 | Our run |
| FlowState r1 | 0.5221 | 0.7506 | Our run |

All rows use the same aggregation over 97 configurations. Mixing our runs with published rows needs care: the published leaderboard lists Chronos-2 at 0.4854, against 0.4828 in our run, so do not over-read differences of a few thousandths between the two sources.

On those published numbers, TiRex-2 (0.4781) and the larger Toto-2.0 2.5B (0.4759) score better than any single model in our runs. TiRex-2 is also the smallest model Ephemeris serves apart from FlowState r1.

## How do they compare on TIME?

TIME is a second benchmark with 31 models on its board. We scored it with its own runner against the published model outputs; two of the models we serve have rows there.

| Model | TIME MASE ratio | TIME CRPS ratio |
|---|---|---|
| TimesFM-3 | 0.640 | 0.536 |
| Toto-2.0 2.5B | 0.642 | 0.539 |
| Toto-2.0 313M | 0.644 | 0.542 |
| Chronos-2 | 0.662 | 0.556 |

On TIME, Toto 2 (313M) is ahead of Chronos-2 on both scores, the reverse of the near-tie on GIFT-Eval. That is a fair summary of the whole comparison: rankings move between benchmarks.

## Which model wins where?

No model wins everywhere. These are the patterns we measured; each is from our own calibration runs.

- Short horizons: Toto 2 (313M) is the strongest single member in our calibration.
- Long horizons: PatchTST-FM r2 is the strongest single member at GIFT-Eval's long horizons.
- Covariates: Chronos-2 and TiRex-2 accept them. Past 320 steps, route mode sends covariate requests to Chronos-2, because TiRex-2's head emits 320 steps per pass and longer horizons are rolled out.
- Long histories: TimesFM 2.5 reads up to 16,384 points, the longest context in the panel.
- Speed and size: Chronos-2 is the fastest member of the panel; FlowState r1, at 9M parameters, is the smallest and adapts to the sampling rate of the series, but scores lowest here.

Across four benchmarks, the best single model was Chronos-2 on GIFT-Eval, Toto 2 on TIME, TimesFM 2.5 on fev-bench and Toto 2 on BOOM. We cover what that means for combining them in [Does ensembling forecasting foundation models help?](/blog/does-ensembling-forecasting-models-help).

## What about TimesFM-3?

TimesFM-3 is the strongest model on GIFT-Eval that we know of: CRPS 0.4557, about 2% better than our ensemble (0.4662) and about 6% better than Chronos-2, the best single model in our runs. On TIME it is level with the top (MASE 0.640, CRPS 0.536).

Its open weights are under a separate non-commercial licence, restricted to "non-commercial, non-production use", per the [TimesFM README](https://github.com/google-research/timesfm). The same README says commercial use is permitted through Google Cloud services such as BigQuery ML. If you can use it through Google Cloud, or your use is non-commercial research, it is a strong choice. Ephemeris does not serve it.

## How do I run the same comparison myself?

Run each model on the same history and horizon and score against held-out data. You can self-host all six from their Hugging Face checkpoints, or call each by name through one API. With Ephemeris, explicit mode runs exactly one named model:

```python
import os
import requests

history = [112, 118, 132, 129, 121, 135, 148, 148, 136, 119, 104, 118,
           115, 126, 141, 135, 125, 149, 170, 170, 158, 133, 114, 140]

for model in ["chronos2", "timesfm25", "toto2-313m", "tirex2", "patchtst-fm-r2", "flowstate-r1"]:
    r = requests.post(
        "https://ephemeris.cascade.industries/api/v1/forecast",
        headers={"Authorization": f"Bearer {os.environ['EPHEMERIS_API_KEY']}"},
        json={
            "mode": "explicit",
            "model": model,
            "series": [{"values": history, "freq": "M"}],
            "horizon": 12,
            "quantiles": [0.1, 0.5, 0.9],
        },
        timeout=60,
    )
    r.raise_for_status()
    print(model, r.json()["forecasts"][0]["quantiles"]["0.5"])
```

Each run is billed separately, so six explicit calls cost six model runs. Check `GET /api/v1/models` first for the live panel and horizon limits.

## FAQ

### Is Chronos-2 better than TimesFM 2.5?

On GIFT-Eval in our runs, yes, by a small margin: CRPS 0.4828 against 0.4923, and MASE 0.7037 against 0.7096 (ratios to seasonal naive, 97 configurations). TimesFM 2.5 reads longer histories (16,384 points against 8,192), and on fev-bench it was the best single model in our panel.

### Which of these models supports covariates?

In the Ephemeris panel, Chronos-2 and TiRex-2. The TimesFM library also has covariate support through XReg, but Ephemeris does not serve it.

### Is Toto 2 only for observability data?

No. Datadog trained it heavily on observability metrics, but it scores CRPS 0.4842 on the general GIFT-Eval benchmark in our run, essentially level with Chronos-2, and was the best single model in our panel on TIME and BOOM.

### Which model is best overall?

None of the six wins on every benchmark: Chronos-2 led on GIFT-Eval, Toto 2 on TIME and BOOM, TimesFM 2.5 on fev-bench. TimesFM-3 scores better than all of them on GIFT-Eval but its open weights are non-commercial.

### Are these leaderboard results?

No. Single-model scores marked "our run" come from our own runs of each benchmark's official harness. Rows marked published come from the benchmark's public leaderboard on the same aggregation.

## Related

- [Which time-series foundation model should I use?](/blog/which-forecasting-model-should-i-use)
- [Does ensembling forecasting foundation models help?](/blog/does-ensembling-forecasting-models-help)
- [TimeGPT alternatives in 2026](/blog/timegpt-alternatives)
- [A guide to time-series foundation models](/blog/time-series-foundation-models-guide)
- [Model pages](/models) and [benchmark method](/benchmarks)
