Blog · Updated · 9 min read
Do reasoning models forecast time series better?
Short answer
Mostly no. Off-the-shelf reasoning models show mixed results, and an early TimesX test (July 2026) found GPT-5, Gemini-2.5-Flash and DeepSeek-R1 "do not provide a clear advantage". RL-trained time-series reasoners report gains, mostly on point accuracy or reasoning tasks. The strongest forecasting results use reasoning to correct, weight or guide a forecasting model.
On this page
- Key facts
- What is a reasoning model?
- Do off-the-shelf reasoning models forecast better?
- Does thinking longer help?
- Does training a model to reason about time series help?
- Where does reasoning help forecasting most?
- What is still missing from this research?
- Which papers are in this review?
- What does this mean for a forecasting agent?
- FAQ
- Related
Key facts
- In the ReC4TS benchmark (Liu et al., February 2025), DeepSeek-R1 was the only effective reasoning model for zero-shot forecasting; o1-mini and Gemini-2.0-Flash-Thinking were not, and self-consistency was the most effective test-time strategy (arXiv 2503.01895).
- On the TimesX benchmark (Liu et al., July 2026), an early experiment evaluated GPT-5, Gemini-2.5-Flash and DeepSeek-R1 on samples after January 2025 and found "these reasoning models do not provide a clear advantage" (arXiv 2607.06973).
- The TimeReasoner study found slow-thinking LLMs show "non-trivial zero-shot forecasting capabilities", but longer chains of thought were "disproportionately concentrated in the worst MSE rankings" (arXiv 2505.24511).
- TSRBench (ICML 2026) found that switching GPT-5 and Qwen3 models from reasoning to non-reasoning mode left perception tasks intact but caused "a sharp degradation" on reasoning, prediction and decision tasks; its prediction tasks are multiple choice (arXiv 2601.18744).
- Time-R1 trains Qwen2.5-7B-Instruct with supervised fine-tuning then reinforcement learning to reason before forecasting, and reports that it "significantly improves forecast performance across diverse datasets", scored with MSE and MAE (arXiv 2506.10630).
- VeriTime (2026) trains small models on verifiable reasoning traces and reports 3B and 4B models matching or exceeding larger proprietary LLMs on time-series reasoning tasks (arXiv 2602.07830).
What is a reasoning model?
A reasoning model (also called a thinking or slow-thinking model) writes out intermediate steps, a chain of thought, before it gives a final answer. OpenAI's o-series, DeepSeek-R1 and the thinking modes of Gemini, GPT-5 and Qwen3 are examples.
Many are trained with reinforcement learning (RL): the model tries an answer, gets a reward for getting it right, and is updated to make rewarded behaviour more likely. DeepSeek-R1 popularised a method called GRPO (Group Relative Policy Optimization), which compares several attempts at the same problem and rewards the better ones.
The question for forecasting is simple: if the model thinks before writing numbers, are the numbers better? Two kinds of evidence exist. One prompts off-the-shelf reasoning models. The other trains a model with RL specifically to reason about series.
Do off-the-shelf reasoning models forecast better?
The evidence is mixed, and leans towards no clear gain.
ReC4TS (Liu et al., arXiv 2503.01895) was an early benchmark of reasoning strategies for zero-shot forecasting across eight domains, scored with MSE (mean squared error). DeepSeek-R1 was the only effective reasoning model, with significant improvements in three of four settings and beating its non-reasoning sibling DeepSeek-V3 in 60% of cases. o1-mini and Gemini-2.0-Flash-Thinking were ineffective. The most effective strategy was self-consistency: sample several answers and pick the most consistent, a test-time technique that works with ordinary models too.
TimeReasoner (Cheng et al., arXiv 2505.24511) prompted slow-thinking models, mainly DeepSeek-R1, on standard forecasting datasets. The authors found "non-trivial zero-shot forecasting capabilities, especially in capturing high-level trends and contextual shifts". They also catalogued four failure patterns:
- Peak clipping: flattening peaks and troughs.
- Phase shift: the right shape, at the wrong time.
- Copy-paste repeat: replicating a past segment without adapting it.
- Constant collapse: degenerating into a nearly flat line.
TimesX (Liu et al., arXiv 2607.06973) is the most recent test. Its main benchmark includes non-reasoning LLMs such as GPT-4o, Gemini-2.0-Flash and DeepSeek-V3. "As an early experiment", it also evaluated GPT-5, Gemini-2.5-Flash and DeepSeek-R1, using only real-world samples after January 2025 to reduce data contamination, scored with MASE (mean absolute scaled error). Its conclusion: "these reasoning models do not provide a clear advantage."
Does thinking longer help?
Not reliably. Some thinking helps reasoning tasks; more thinking did not help forecasting in the one study that measured it.
TimeReasoner found that longer chains of thought clustered in the worst-scoring forecasts. The authors suggest extended reasoning can bring in irrelevant detail.
TSRBench (Yu et al., ICML 2026) ran GPT-5, Qwen3-32B and Qwen3-VL-32B with reasoning on and off. Perception stayed robust without reasoning, while reasoning, prediction and decision-making degraded sharply. The authors conclude that models "can intuitively perceive temporal patterns through fast, heuristic processing", but drawing conclusions needs deliberate reasoning. Their prediction tasks are multiple choice rather than written-out forecasts. The same paper found scaling laws "break down for prediction", and describes "a decoupling between semantic understanding and numerical prediction".
Read together: reasoning helps a model work out what a series means. It is weaker evidence that reasoning produces better numbers.
Does training a model to reason about time series help?
Training helps on the tasks it targets, but most results are scored on point accuracy or classification, and few test calibrated ranges.
- Time-R1 (Zhou et al., arXiv 2506.10630) fine-tunes Qwen2.5-7B-Instruct in two stages: supervised fine-tuning on synthetic reasoning traces, then RL with a reward built for forecasting and a GRPO-style method called GRIP. The traces are generated with DeepSeek-R1. It reports significantly better forecasts across datasets, measured with MSE and MAE. The authors argue training is better than prompting large reasoning models because prompting has "high computational cost, privacy risks" and limited domain depth.
- VeriTime (Zhou et al., arXiv 2602.07830) builds training data whose reasoning steps can be checked, orders it by difficulty, and trains with RL. It reports small 3B and 4B models reaching or beating larger proprietary models on time-series reasoning tasks.
- TimeMaster (Zhang et al., arXiv 2506.13705) uses supervised fine-tuning then GRPO on a 3B vision-language model (Qwen2.5-VL-3B-Instruct) reading plotted series. It is a classification result, not forecasting: over 14.6% better than classical time-series models and 7.3% better than few-shot GPT-4o on six TimerBed tasks.
- SenTSR-Bench (He et al., AISTATS 2026) injects insights from a fine-tuned time-series LLM into a general reasoning model's chain of thought, using RL with verifiable rewards. On diagnostic reasoning it reports gains of 9.1%-26.1% over time-series LLMs and 7.9%-22.4% over general reasoning models.
These are real gains in reasoning about series. But point metrics such as MSE say nothing about whether a model's uncertainty is honest, which is what decisions need.
Where does reasoning help forecasting most?
When it works around a numeric forecaster rather than replacing it.
TSOrchestra (Cao et al., arXiv 2512.16022) applies R1-style training (supervised fine-tuning then GRPO) to a small LLM (Qwen-2.5-3B-Instruct), but its job is to weight and explain an ensemble of time-series foundation models, not to write forecasts. The authors note that direct use of LLMs for forecasting "has proven ineffective", and report that the orchestrated ensemble outperforms every individual foundation model on GIFT-Eval (code, leaderboard) by average rank on CRPS and MASE.
STRIDE (Ahamed et al., arXiv 2605.08625) distills reasoning traces into a small LLM and feeds its hidden state into a foundation model, trained jointly with a quantile loss. It reports 0.674 MASE and 0.454 CRPS on GIFT-Eval and improvements to Chronos-2 and Timer-S1.
Prompting work points the same way. Ashok et al. (arXiv 2508.09904) found an "Execution Gap", where models "explain correctly but fail to improve forecasts", which "affects even the largest models". Asking the LLM to correct a quantitative model's probabilistic forecast worked best for 13 of 17 LLMs tested, with improvements of up to 50% over asking it to forecast directly.
What is still missing from this research?
Three things: probabilistic scoring, cost, and contamination control.
- Calibration. Most reasoning-model forecasting papers report MSE or MAE. TimeReasoner names uncertainty quantification as future work. Separately, QuantSightBench (arXiv 2604.15859) found that none of the 11 frontier and open-weight models it tested reached 90% coverage on 90% prediction intervals; the best, Gemini 3.1 Pro, reached 79.1%.
- Cost. Reasoning models write many tokens before answering. Few papers report cost next to accuracy, which the 2026 survey of forecasting agents (Xu et al., arXiv 2608.23058) lists as a gap.
- Contamination. Older benchmark series and the events around them may be in a model's training data. TimesX built its evaluation so forecast horizons start after the pretraining cutoff, and the 2026 survey warns that some benchmark gains "may reflect contamination instead of temporal reasoning".
Which papers are in this review?
| Paper | Venue, date | Setup | Headline finding | Link |
|---|---|---|---|---|
| ReC4TS (Liu et al.) | arXiv, Feb 2025 | Reasoning strategies, zero-shot, 8 domains | Only DeepSeek-R1 effective; self-consistency best | 2503.01895 |
| TimeReasoner (Cheng et al.) | arXiv, May 2025 | Prompted slow-thinking LLMs | Non-trivial zero-shot skill; long chains score worst | 2505.24511 |
| TimesX (Liu et al.) | arXiv, Jul 2026 | Early test of reasoning LLMs, post-Jan-2025 data | No clear advantage from reasoning | 2607.06973 |
| TSRBench (Yu et al.) | ICML 2026, Jan 2026 | Reasoning on vs off | Turning reasoning off hurts reasoning and prediction tasks | 2601.18744 |
| Time-R1 (Zhou et al.) | arXiv, Jun 2025 (rev. Aug 2026) | SFT + RL on Qwen2.5-7B | Better MSE/MAE across datasets | 2506.10630 |
| VeriTime (Zhou et al.) | arXiv, Feb 2026 | Verifiable reasoning data + RL | 3B-4B models match larger proprietary LLMs on reasoning | 2602.07830 |
| TimeMaster (Zhang et al.) | arXiv, Jun 2025 | GRPO on plotted series | Classification gains over GPT-4o and classical models | 2506.13705 |
| SenTSR-Bench (He et al.) | AISTATS 2026, Feb 2026 | Knowledge injection into reasoning traces | 7.9%-26.1% gains on diagnostic reasoning | 2602.19455 |
| TSOrchestra (Cao et al.) | arXiv, Dec 2025 | RL-trained LLM weights foundation models | Beats every single foundation model on GIFT-Eval (average rank) | 2512.16022 |
| STRIDE (Ahamed et al.) | arXiv, May 2026 | Reasoning injected into a foundation model | 0.674 MASE, 0.454 CRPS on GIFT-Eval | 2605.08625 |
What does this mean for a forecasting agent?
Use reasoning for interpretation and for steering a forecaster, not as the forecaster.
A reasoning model is a good choice for the agent's planning: reading the request, deciding which context matters, checking a forecast against what it knows, and explaining it. For the numbers and the uncertainty band, in our reading the 2025-2026 evidence still favours a model trained to produce quantiles, with the LLM correcting or weighting it where text adds information. The broader picture is in what 2025-2026 research says about LLMs reasoning over time series, and the architecture in the split that works.
FAQ
Do reasoning models like o1, DeepSeek-R1 or GPT-5 forecast better than normal LLMs?
Not consistently. ReC4TS found only DeepSeek-R1 helped among the reasoning models it tested, and an early TimesX test found GPT-5, Gemini-2.5-Flash and DeepSeek-R1 "do not provide a clear advantage" on recent data.
Does a longer chain of thought improve a forecast?
There is no evidence that it does. In the TimeReasoner study, longer chains of thought were concentrated among the worst-scoring forecasts.
What is Time-R1?
Time-R1 is a method that fine-tunes Qwen2.5-7B-Instruct with supervised learning and then reinforcement learning so it reasons step by step before forecasting. Its authors report better MSE and MAE across datasets; it is evaluated on point accuracy, not calibrated ranges.
Where does LLM reasoning help forecasting most?
Around a numeric model: correcting its forecast with context, or weighting an ensemble of forecasting models. TSOrchestra and STRIDE, which use reasoning this way, report strong GIFT-Eval results.
Related
- What 2025-2026 research says about LLMs reasoning over time series
- Why language models are bad at forecasting numbers
- The split that works: the LLM reads the context, a forecasting model does the numbers
- Time-series foundation models: a practical guide
- Does ensembling forecasting models help?
- Model catalogue