ephemerisBlog

Blog · Updated · 13 min read

What 2025-2026 research says about LLMs reasoning over time series

Short answer

Recent benchmarks show LLMs do better at perceiving and describing time series than at forecasting them, and that skill does not reliably carry over to accurate numbers: in QuantSightBench, none of 11 models reached 90% coverage on 90% prediction intervals. Most of the stronger results keep a forecasting model in the loop.

LLMsAgentsBenchmarksRead as markdown

On this page
  1. Key facts
  2. What does "time-series reasoning" mean?
  3. What are LLMs good at with time series?
  4. Where do LLMs still fall short?
  5. Can LLMs give calibrated forecast ranges?
  6. Does giving an LLM text context reliably help?
  7. What works best: hybrids and agents?
  8. Which papers are in this review?
  9. How much should you trust this research?
  10. What should you do when building a forecasting agent?
  11. FAQ
  12. Related

Key facts

  • TSRBench (Yu et al., ICML 2026) tested over 30 LLMs, vision-language models and time-series LLMs on 4125 problems from 14 domains, and found "scaling laws hold for perception and reasoning but break down for prediction" (arXiv 2601.18744).
  • The same paper reports that "strong reasoning does not guarantee accurate context-aware forecasting", which the authors describe as a decoupling between semantic understanding and numerical prediction.
  • On QuantSightBench (April 2026), none of 11 frontier and open-weight models reached the 90% coverage target for 90% prediction intervals; the best were Gemini 3.1 Pro at 79.1%, Grok 4 at 76.4% and GPT-5.4 at 75.3% (arXiv 2604.15859).
  • HEARTS (February 2026) evaluated 16 LLMs on more than 20,000 health time-series test samples and found they "substantially underperform specialized models", with performance only weakly related to general reasoning scores (arXiv 2603.06638).
  • Ashok et al. (TMLR) name an "Execution Gap": models correctly explain how context should change a forecast but fail to apply that reasoning to the numbers (arXiv 2508.09904).
  • LLM coding agents that query the data with Python beat models reading raw numbers by up to 10% on two time-series understanding benchmarks, but the best agent still got about 22-34% of questions wrong (arXiv 2606.16545).
  • On the TimesX multimodal forecasting benchmark, the best performers were simple average ensembles, with the top method averaging TimesFM-2.5 and Gemini-2.0-Flash forecasts (arXiv 2607.06973).

What does "time-series reasoning" mean?

It means answering questions about a series that need more than continuing the curve: describing its shape, spotting an anomaly, explaining a change, linking it to an event in text, or deciding what to do next.

The 2026 benchmarks split it into layers. TSRBench uses four dimensions: Perception (what is in the series), Reasoning (what follows from it), Prediction (what comes next) and Decision-Making (what to do). HEARTS, a health benchmark, uses Perception, Inference, Generation and Deduction. TSCognition (Qiu et al., June 2026) uses Decoding, Grounding, Inferring, Extrapolating and Acting.

Forecasting is only one of these layers. That matters, because a recurring finding of the past year is that LLMs do better on perception than on forecasting itself. TSRBench still found models "struggle significantly with complex reasoning, forecasting, and decision-making tasks", so the gap is not simply "good at reasoning, bad at numbers". Our earlier post on why language models are bad at forecasting numbers covers the older evidence (tokenization, LLMTime, ablations); this post covers what came after.

What are LLMs good at with time series?

Describing and reasoning about simple patterns, and connecting a series to the text around it. Several recent papers show real skill here, with caveats.

  • Simple concepts. In TimeSeriesExam, a multiple-choice exam on synthetic series covering pattern recognition, noise, similarity, anomalies and causality, closed models such as GPT-4 and Gemini understood simple concepts better than open models (Cai et al., NeurIPS 2024 workshop).
  • Perception and reasoning scale. TSRBench found perception and reasoning scores rise with model size, as they do on other tasks.
  • Explaining context. The Execution Gap finding has a positive half: models do explain correctly how a piece of context should affect a forecast.
  • Analysis with tools. Rechtorík et al. (June 2026) found coding agents "can select appropriate statistical tests, but often miss important nuances", and that agents with code access beat raw-number reading by up to 10%.
  • Interpretable classification. TimeMaster (Zhang et al., arXiv 2506.13705) trained a 3B vision-language model with reinforcement learning to classify plotted series, beating classical time-series models by over 14.6% and few-shot GPT-4o by 7.3% on six TimerBed tasks, while writing explanations.

There is also a dissenting result on the older question of whether LLMs help forecasting at all. Qiu et al. (arXiv 2602.14744, February 2026) evaluated LLM-based forecasting models (language models aligned to time-series inputs) across 8 billion observations and concluded that they do improve forecasting, with especially large gains across domains, in contrast to earlier studies that reported comparable performance without the LLM.

Where do LLMs still fall short?

On turning understanding into accurate numbers, on hard multi-step questions, and on honest uncertainty.

Prediction does not follow reasoning. TSRBench reports that Perception, Reasoning and Decision scores are highly correlated with each other but only weakly correlated with Prediction. Bigger models did not forecast better. TSRBench framed forecasting as multiple choice "to reduce the difficulty" for generalist models; in an appendix the authors also scored open-ended numeric forecasts and found the same lack of improvement with model size.

Hard questions remain hard. TimeSeriesExam found all models struggled with causality analysis. Its 2026 follow-up, TimeSeriesExamAgent (Gwiazda et al., April 2026), generated exams from real healthcare, finance and weather data and concluded that "LLM performance remains limited in both abstract time series reasoning and domain-specific applications". HEARTS found LLMs "often rely on simple heuristics and struggle with multi-step temporal reasoning", and that performance falls as temporal complexity rises.

Benchmarks that only score forecasts hide failures. TemporalBench (Weng et al., February 2026) separates historical interpretation, context-free forecasting, contextual reasoning and event-conditioned prediction across retail, healthcare, energy and physical systems. It found that "strong numerical forecasting accuracy does not reliably translate into robust contextual or event-aware temporal reasoning", and that agent frameworks show "systematic failure modes that remain largely hidden under forecasting-only benchmarks". MTBench (Chen et al., arXiv 2503.16858), on finance news and weather reports, reported difficulty with long-term dependencies and with fusing text and series.

Knowing is not doing. The Execution Gap (Ashok et al.) is the clearest statement of the problem: the model's explanation is right, but its forecast does not reflect it.

Can LLMs give calibrated forecast ranges?

Not reliably yet. Calibration means a stated range holds the truth as often as it claims: a 90% interval should contain the outcome 90% of the time. That fraction is called coverage.

QuantSightBench (Qin and Andriushchenko, April 2026) asked 11 frontier and open-weight models for prediction intervals on 1,000 numeric forecasting questions drawn from the web, covering topics such as business and finance, politics, infrastructure and sports. None reached 90% coverage; the top three were 10 or more points short. The authors report that calibration "degrades sharply at extreme magnitudes, revealing systematic overconfidence across all evaluated models."

QuantSightBench is about one-off numeric questions, not long time series, but the pattern matches older findings on series. The TimeReasoner study of reasoning models on forecasting (Cheng et al., arXiv 2505.24511) lists principled uncertainty quantification as open work.

There are signs this is fixable by training. Baldelli et al. (arXiv 2605.11845, May 2026) fine-tuned 12 models to sample from target distributions and found "probabilistic calibration is a trainable capability", though the gains sometimes reduced arithmetic reasoning. That work tests sampling fidelity, not forecast intervals.

Does giving an LLM text context reliably help?

Sometimes. The benefit depends on the data, the text and the model.

Zhang et al. (arXiv 2506.21611) tested 16 forecasting tasks across 7 domains and found the benefits of adding text "highly condition-dependent". Gains were more likely with a high-capacity text model, a comparatively weak time-series model, enough training data, and text that carries signal the series does not already contain.

TimesX (Liu et al., July 2026), a benchmark built from real series with automatically gathered text context and controls for data leakage, found that many approaches that do well on existing benchmarks "may fail" on it. Zero-shot LLMs kept an edge over unimodal foundation models, but a smaller one than on synthetic benchmarks.

Zheng et al. (arXiv 2603.12451) hypothesise that multimodal models often fail to beat unimodal ones because the context in existing datasets is low quality. They built CAF-7M, 7 million semi-synthetic context-augmented windows with a verified test set, showed that pre-training on it transfers to real-world evaluation with "clear evidence of context utilization", and conclude that dataset quality, rather than model architecture, has been the main bottleneck.

What works best: hybrids and agents?

Most of the stronger results in this set keep a numeric forecasting model in the loop and give the LLM a job around it: correcting, weighting, choosing context, or explaining.

  • Correct a forecast instead of writing one. Ashok et al. gave the LLM a probabilistic forecast from a quantitative model (Lag-Llama, Chronos or ARIMA) and asked it to correct the forecast using the context. On the Context is Key benchmark this correction approach beat direct prompting for 13 of 17 LLMs, with improvements of up to 50% over asking the LLM to forecast directly.
  • Average the LLM with a foundation model. In the main TimesX comparison, simple equal-weight averages beat the LLM revision methods tested, with TimesFM-2.5 plus Gemini-2.0-Flash ranked first. The authors note this is not meant as the optimal design, and a tuned text-revision variant later edged past it on a subset.
  • Let the LLM weight an ensemble. TSOrchestra (Cao et al., arXiv 2512.16022) fine-tunes an LLM to act as a judge that weights and explains an ensemble of time-series foundation models; the authors report it "significantly outperforms" leading foundation models on GIFT-Eval's 97 settings (code, leaderboard) on both CRPS and MASE.
  • Inject reasoning into the forecaster. STRIDE (Ahamed et al., arXiv 2605.08625) feeds a small LLM's reasoning into a foundation model as an embedding, trained jointly with cross-entropy and quantile losses, and reports 0.674 MASE and 0.454 CRPS on GIFT-Eval. TS-Reasoner (Yu et al., TMLR 2026) goes the other way, feeding a frozen foundation model's representations to an LLM for question answering.
  • Agents on top of a forecaster. Liao et al. (arXiv 2606.02497) call the step of adjusting a baseline for holidays, campaigns and expert feedback "last-mile forecasting", and build an agent that turns its reasoning into explicit, auditable revisions under safety constraints. They show case studies, not benchmark scores.

One counter-example is worth noting. Nexus (Das et al., arXiv 2605.14389), a multi-agent LLM system with no external statistical model, matched or beat foundation models on Zillow and stock data dated after the LLMs' knowledge cutoffs. The authors argue LLMs have "substantially stronger intrinsic forecasting ability than previously recognized" when the reasoning is organised well. It covers two data sources, so it does not yet show the result holds broadly.

Which papers are in this review?

PaperVenue, dateWhat it testsHeadline findingLink
TSRBench (Yu et al.)ICML 2026, Jan 20264125 problems, 14 domains, 30+ modelsReasoning scales with size, prediction does not2601.18744
TimeSeriesExamAgent (Gwiazda et al.)arXiv, Apr 2026Auto-generated exams on real dataLLM time-series reasoning "remains limited"2604.10291
HEARTS (Li et al.)arXiv, Feb 2026110 health tasks, 16 LLMsLLMs substantially underperform specialised models2603.06638
TemporalBench (Weng et al.)arXiv, Feb 2026Four tiers from history to eventsForecast accuracy does not imply contextual reasoning2602.13272
QuantSightBench (Qin, Andriushchenko)arXiv, Apr 202690% prediction intervals, 11 modelsBest coverage 79.1%; systematic overconfidence2604.15859
Beyond Naive Prompting (Ashok et al.)TMLR, Aug 2025 (rev. 2026)Prompting strategies on Context is KeyExecution Gap; correcting a model forecast up to 50% better than direct prompting2508.09904
Can LLM Coding Agents Reason About Time Series? (Rechtorík et al.)arXiv, Jun 2026Raw data vs code accessCode helps by up to 10%; 22-34% still wrong2606.16545
TimesX (Liu et al.)arXiv, Jul 2026Real multimodal forecasting with leakage controlSimple LLM + foundation-model averages rank best2607.06973
When Does Multimodality Lead to Better TSF? (Zhang et al.)arXiv, Jun 202516 tasks, 7 domainsText helps only under specific conditions2506.21611
Overcoming the Modality Gap (Zheng et al.)arXiv, Mar 2026CAF-7M context corpusSuggests dataset quality, not architecture, was the main bottleneck2603.12451
TSOrchestra (Cao et al.)arXiv, Dec 2025LLM judge weighting foundation modelsBeats leading foundation models on GIFT-Eval2512.16022
STRIDE (Ahamed et al.)arXiv, May 2026LLM reasoning injected into a foundation model0.674 MASE, 0.454 CRPS on GIFT-Eval2605.08625
Nexus (Das et al.)arXiv, May 2026Multi-agent LLM forecastingMatches or beats foundation models on 2 post-cutoff sources2605.14389
Last-mile forecasting (Liao et al.)arXiv, Jun 2026Agent revising a baseline forecastAuditable revisions; case studies only2606.02497
LLM-based forecasting agents survey (Xu et al.)arXiv, Aug 2026Review of methods and evaluationMeasurement is the central limitation2608.23058

How much should you trust this research?

Treat it as early. Most of these papers are preprints, the benchmarks are new, and they do not share a scoring standard.

A survey of LLM-based forecasting agents (Xu et al., August 2026) reviews both positive and negative evidence, including sensitivity to small input changes, ablations where the LLM adds nothing, and benchmark gains that "may reflect contamination instead of temporal reasoning". Contamination means the test data, or news about it, was in the model's training set. The survey concludes that "measurement is a central limitation" and calls for calibration under distribution shift, live evaluation and reporting cost alongside accuracy.

Several results above score point forecasts or multiple-choice answers, not full probabilistic forecasts. Results on reasoning models specifically are covered in do reasoning models forecast better?.

What should you do when building a forecasting agent?

Use the LLM for the layers the research says it handles, and keep a forecasting model for the numbers.

  1. Give the LLM the reading and reasoning. Describing the series, linking it to events in text, spotting anomalies and explaining results are where the evidence is most positive.
  2. Do not let the LLM write the numbers alone. TSRBench and the Execution Gap point the same way: understanding does not become accurate forecasts on its own. HEARTS also found LLMs well behind specialised models on health series.
  3. Do not trust a range the LLM writes. QuantSightBench found overconfidence across every model it tested. Take intervals from a model trained to produce quantiles.
  4. Let the LLM adjust, weight or correct a numeric forecast. Correction, simple averaging and LLM-weighted ensembles have some of the strongest results in the papers reviewed here, though Nexus shows a pure-LLM design can also compete.
  5. Give it tools. Code access helped by up to 10%, and turning known future events into numeric inputs (covariates) keeps them auditable.
  6. Check the result, not the explanation. A good explanation is not evidence of a good forecast; score forecasts against outcomes.

This is the split described in the LLM reads the context, a forecasting model does the numbers. Ephemeris is one way to give an agent the forecasting-model half as a tool.

FAQ

Are LLMs good at time-series reasoning?

Partly. Recent benchmarks show they handle perception and simple pattern questions and can explain how context affects a series, but they struggle with complex and multi-step reasoning, causality and specialised domains, and they lag specialised models on health data in HEARTS.

Does better reasoning make an LLM a better forecaster?

Not on current evidence. TSRBench found reasoning scores rise with model size while prediction scores do not, and only weakly correlate with them.

Can an LLM produce calibrated prediction intervals?

Not reliably. On QuantSightBench none of 11 frontier and open-weight models reached 90% coverage on 90% intervals, with the best at 79.1%, and all showed systematic overconfidence.

What is the "Execution Gap" in LLM forecasting?

It is the finding by Ashok et al. that LLMs often explain correctly how context should change a forecast but then fail to apply that reasoning to the numbers they output.

What is the best way to combine an LLM and a forecasting model?

Most of the papers reviewed here favour keeping the forecasting model in charge of the numbers and letting the LLM correct, weight or average its output using context. Simple averages of an LLM and TimesFM-2.5 ranked best on TimesX.