Benchmarks · ensemble 0.7.5
Level with the best forecasting models, under an open licence.
The Ephemeris ensemble blends six open-weights foundation models, weighting each by its measured accuracy at the requested horizon. Scored with each benchmark's own harness, it is level with the top of TIME and beats every open-licence model on GIFT-Eval. The only model ahead of it on GIFT-Eval is TimesFM-3, which cannot be used commercially.
31 models
TIME
Lower is better. The ensemble has the best average MASE rank of all 31 models (5.684), and is within 0.002 of the top on both scores.
| Model | CRPS ratio | MASE ratio | Note |
|---|---|---|---|
| QiYao-M | 0.536 | 0.638 | |
| Ephemeris ensemble | 0.538 | 0.639 | |
| TimesFM-3 | 0.536 | 0.640 | non-commercial licence |
| Toto-2.0 2.5B | 0.539 | 0.642 | |
| Toto-2.0 313M | 0.542 | 0.644 | |
| Chronos-2 | 0.556 | 0.662 |
97 configurations
GIFT-Eval
Lower is better. Other rows are the published leaderboard on the same aggregation; "—" is not reported there.
| Model | CRPS ratio | MASE ratio | Note |
|---|---|---|---|
| TimesFM-3 | 0.4557 | 0.6668 | non-commercial licence |
| Ephemeris ensemble | 0.4662 | 0.6841 | |
| Toto-2.0 2.5B | 0.4759 | — | |
| TiRex-2 | 0.4781 | — | |
| Chronos-2 | 0.4854 | — |
Why an ensemble
Ensemble vs its best single model
How much worse the best individual model in the panel is than the ensemble, per benchmark. A different model is best on each, which is why no single model is the right default.
| Benchmark | Configurations | Best single model | Its CRPS vs ensemble | Its MASE vs ensemble |
|---|---|---|---|---|
| GIFT-Eval | 97 | Chronos-2 | +3.6% | +2.9% |
| TIME | 134 | Toto 2 | +1.6% | +2.5% |
| fev-bench | 96 | TimesFM 2.5 | +1.1% | +1.9% |
| BOOM | 1,073 | Toto 2 | +0.0% | -0.2% |
Reproducibility
Method
Scores
CRPS (probabilistic accuracy, from nine quantiles) and MASE (point accuracy of the median), each divided by a seasonal-naive forecast of the same window and averaged geometrically across configurations, as the benchmarks themselves report them.
Harnesses
TIME: the benchmark's own runner and compute_local_leaderboard.py, ranked against its published model outputs. GIFT-Eval: the official harness and published seasonal naive. These are our runs, not leaderboard submissions. The ensemble weights were fixed on separate data before any of these benchmarks were scored.
Same ensemble, one request