ephemerisBlog

Blog · Updated

Does ensembling forecasting foundation models help? Results on four benchmarks

Short answer

Usually, by a small margin. Across four benchmarks, the best single model in our panel was 3.6% worse than the ensemble on GIFT-Eval CRPS, 1.6% worse on TIME and 1.1% worse on fev-bench. On BOOM there was no gain at all. A different model was best on each benchmark, which is the main argument for ensembling; the cost is running every compatible model.

forecasting · ensembles · benchmarksRead as markdown

Key facts

  • The Ephemeris ensemble scores CRPS 0.4662 and MASE 0.6841 on GIFT-Eval (97 configurations, ratios to seasonal naive, lower is better).
  • On GIFT-Eval the best single model in the panel, Chronos-2, is 3.6% worse than the ensemble on CRPS and 2.9% worse on MASE.
  • On TIME (134 configurations) the best single model, Toto 2, is 1.6% worse on CRPS and 2.5% worse on MASE.
  • On fev-bench (96 configurations) the best single model, TimesFM 2.5, is 1.1% worse on CRPS and 1.9% worse on MASE.
  • On BOOM (1,073 configurations) the ensemble gives no gain: Toto 2 alone matches it on CRPS (0.0%) and is 0.2% better on MASE.
  • An ensemble forecast runs and bills every compatible model, so it costs the sum of their per-model rates; route mode usually runs one.

Does an ensemble beat the best single model?

On three of four benchmarks it did, by 1% to 4%; on the fourth it tied. The table shows how much worse the best individual model in the panel was than the ensemble, per benchmark. A positive number means the single model was worse.

BenchmarkConfigurationsBest single modelIts CRPS vs ensembleIts MASE vs ensemble
GIFT-Eval97Chronos-2+3.6%+2.9%
TIME134Toto 2+1.6%+2.5%
fev-bench96TimesFM 2.5+1.1%+1.9%
BOOM1,073Toto 2+0.0%-0.2%

CRPS (continuous ranked probability score) measures the whole forecast distribution: it rewards uncertainty bands that are narrow and correct. MASE (mean absolute scaled error) measures the point forecast, here the median. Both are divided by the error of a seasonal-naive forecast (repeat the last season) on the same window and averaged geometrically across configurations, as the benchmarks report them.

These are our own runs with each benchmark's harness, not leaderboard submissions. The TIME row (134 configurations) uses a different aggregation from the 31-model TIME leaderboard table on the benchmarks page, so the two TIME numbers are not directly comparable. The GIFT-Eval row uses our own run of Chronos-2 (CRPS 0.4828), not its published leaderboard score (0.4854). The method is on /benchmarks.

Why is a different model best on each benchmark?

Each model was pretrained on different data with a different architecture, so each is strongest on different kinds of series. Chronos-2 was best on GIFT-Eval, Toto 2 on TIME and BOOM, and TimesFM 2.5 on fev-bench.

This matters more than the size of the gain. If you pick one model on one benchmark, you are betting that your data looks like that benchmark. An ensemble hedges that bet: on every benchmark here it was ahead of, or within 0.2% of, whichever single model happened to be best there.

The spread inside one benchmark is small, too. On GIFT-Eval the four best single models in our runs sit within about 2% of each other on CRPS (Chronos-2 0.4828, Toto 2 0.4842, PatchTST-FM r2 0.4869, TimesFM 2.5 0.4923). Picking "the best model" from a leaderboard is choosing between near-ties. The per-model comparison is in Chronos-2 vs TimesFM 2.5 vs Toto 2 vs TiRex-2.

Why does the ensemble not help on BOOM?

On BOOM, Toto 2 alone is level with the ensemble on CRPS and 0.2% better on MASE. Ensembling bought nothing there.

BOOM is Datadog's observability benchmark, per the Toto 2 model card, and Toto 2 was trained heavily on observability metrics. Our reading is that when one specialist model matches the data well, mixing in generalists does not add accuracy. If your data is infrastructure metrics, Toto 2 on its own is a reasonable choice and costs one model run.

How are the models combined?

The Ephemeris ensemble is an accuracy-weighted mixture: models that were more accurate at the requested horizon get more weight. The weights are reported per horizon by GET /api/v1/models.

There are two ways to combine quantile forecasts, selected with the combine field in ensemble mode.

  • mixture (the default) pools the models' predictive distributions. Think of it as drawing each outcome from model A with probability equal to A's weight, from model B with B's weight, and so on, then reading quantiles off the pooled result. When the models disagree, the band widens to cover the disagreement.
  • vincentize averages the quantiles level by level: the combined 0.9 quantile is the weighted average of each model's 0.9 quantile. The band stays about as wide as an average member's band, even when the members disagree about where the series is going.

The API guidance is to prefer mixture. In our view that is the right default when you care about calibration (the 0.9 quantile being exceeded about 10% of the time), because disagreement between models is real uncertainty. vincentize gives a narrower, smoother band if that is what a downstream system expects.

You can also cap the panel with top_k (1 to 16 models).

python
import os
import requests

r = requests.post(
    "https://ephemeris.cascade.industries/api/v1/forecast",
    headers={"Authorization": f"Bearer {os.environ['EPHEMERIS_API_KEY']}"},
    json={
        "mode": "ensemble",
        "combine": "mixture",
        "series": [{"values": [42, 45, 44, 50, 53, 51, 49, 55, 58, 57, 54, 60], "freq": "W"}],
        "horizon": 8,
        "quantiles": [0.1, 0.5, 0.9],
    },
    timeout=60,
)
r.raise_for_status()
body = r.json()
print(body["forecasts"][0]["quantiles"])
print(body["meta"]["models_used"], body["meta"]["billing"]["settled_mc"])

meta.models_used lists the models that contributed and meta.notes says why any were skipped.

How much more does an ensemble cost?

An ensemble costs the sum of every model that ran; route and explicit mode usually cost one model. Each model run is priced as:

price_per_kslot_mc x slots x ceil(context / 1024) x ceil(horizon / 64) millicredits (1 credit = 1,000 millicredits).

A slot is one variate of one series. For one univariate series with up to 1,024 points of context and a horizon up to 64 steps, each model run costs exactly that model's rate, so an ensemble of six costs the six rates added together. Per-model rates are live on /pricing and in GET /api/v1/models.

The panel is not always six. Sending covariates narrows it to models that support them (Chronos-2 and TiRex-2), and past 320 steps TiRex-2 sits out of automatic selection. You pay only for models that actually ran.

When should I use route mode instead?

Use route mode when cost or volume matters more than the last 1% to 4% of accuracy. Route mode picks one model suited to the data and usually bills one run; it may fall back to a small ensemble.

Ensemble mode is worth it when:

  • the forecast feeds a decision where calibrated uncertainty matters (stock levels, capacity, risk);
  • you do not know in advance what your data resembles;
  • the number of forecasts is small enough that several model runs per forecast is affordable.

Route or a single model is the better choice when:

  • you forecast thousands of series and the budget scales with runs;
  • your data closely matches one model's strength, as BOOM-like observability data matches Toto 2;
  • you need one model's specific behaviour, such as TimesFM 2.5's 16,384-point context, in explicit mode.

Can I build an ensemble myself?

Yes. All six models in the panel have open weights on Hugging Face, so you can run them yourself and combine their quantiles. The work is in the weights: fit them per horizon on data you will not evaluate on, or your scores will be optimistic. In our view, equal weights are a safer starting point than weights fitted on your test data.

If you want to test the ensemble against a single model on your own data, Ephemeris runs both from the same request by switching mode; new accounts get a small free credit (sign up).

FAQ

Does ensembling time-series foundation models improve accuracy?

On three of four benchmarks we ran, yes: the best single model was 3.6% (GIFT-Eval), 1.6% (TIME) and 1.1% (fev-bench) worse than the ensemble on CRPS. On BOOM there was no gain.

What is the difference between mixture and vincentization?

A mixture pools the models' predictive distributions, so the band widens when models disagree. Vincentization averages each quantile level across models, so the band stays about as wide as one model's band.

Is the ensemble better than TimesFM-3?

No. On GIFT-Eval, TimesFM-3 scores CRPS 0.4557 against the ensemble's 0.4662. On TIME the ensemble is 0.001 ahead on MASE (0.639 against 0.640) and 0.002 behind on CRPS (0.538 against 0.536). TimesFM-3's open weights carry a non-commercial licence.

Why does an ensemble cost more?

It runs every compatible model and bills each run, so its cost is the sum of their rates. Route mode usually runs one model.