Blog · Topic
Benchmarks
4 posts. See all posts.
4 posts
Chronos-2 vs TimesFM 2.5 vs Toto 2 vs TiRex-2: one harness, same data
Six open-weights forecasting models compared on GIFT-Eval and TIME with one harness: size, licence, context, covariates and CRPS/MASE scores.
Do reasoning models forecast time series better?
What 2025-2026 papers found when "thinking" LLMs and RL-trained time-series reasoners were asked to forecast, and where reasoning actually helps.
Does ensembling forecasting foundation models help? Results on four benchmarks
An accuracy-weighted ensemble of six forecasting models against its best single member on GIFT-Eval, TIME, fev-bench and BOOM, and what it costs.
What 2025-2026 research says about LLMs reasoning over time series
A review of recent papers on how well language models reason about time series, where they fail on numbers and calibration, and what hybrids do better.