Blog · Updated
Why language models are bad at forecasting numbers
Short answer
Language models read numbers as text fragments, have no built-in sense of what is normal for a series, and give uncertainty that is not calibrated. Research finds that removing the LLM from three LLM-based forecasters does not hurt accuracy, and that GPT-4 was calibrated worse than GPT-3, likely because of alignment such as RLHF. LLMs do help when the forecast depends on written context. Let the LLM handle context and reasoning, and give the numbers to a forecasting model.
Key facts
- Tokenizers split numbers inconsistently: GPT-3 tokenizes 42235630 as [422, 35, 630], so a one-digit change can produce a completely different token sequence (Gruver et al., NeurIPS 2023).
- In the same study GPT-4 had worse uncertainty calibration than GPT-3, which the authors attribute to alignment such as RLHF (reinforcement learning from human feedback).
- Ablating three LLM-based forecasters (OneFitsAll, Time-LLM, CALF), removing the LLM or replacing it with a single attention layer did not degrade accuracy and in most cases improved it (Tan et al., NeurIPS 2024).
- Time-LLM has 6,642M parameters and took 3,003 minutes to train on the Weather dataset; the ablations had 0.245M parameters and took 2.17 minutes on average (Tan et al.).
- On the Context is Key benchmark (71 tasks where the text context is essential), prompted LLMs beat statistical models and time-series foundation models, but without the context they "no longer dominate" (Williams et al., ICML 2025).
- Context is Key also reports LLM forecasts that miss by at least 500% of the ground-truth range, and notes that many LLM-based forecasters are Pareto-dominated by quantitative forecasters such as Lag-Llama and Chronos.
Can a language model forecast a time series at all?
Yes, somewhat. A chat model will return a plausible continuation of a series, and with careful formatting it can be competitive.
The clearest evidence is LLMTime (Gruver et al., 2023). The authors encoded series as strings of digits and found that GPT-3 and LLaMA-2 could forecast zero-shot at a level comparable to, or better than, purpose-built models of the time. They credit the models' preference for simple and repeating patterns, which many real series have.
That result came with conditions. It depended on reformatting numbers so the tokenizer handled them well, on rescaling the series first, and on sampling many continuations to build a distribution. A model asked casually in a chat window gets none of that help.
Why do numbers confuse a language model?
Because a language model never sees a number. It sees tokens, and tokenizers were built for words.
Byte-pair encoding, the common scheme, splits digits into chunks based on how often those chunks appeared in training text. Gruver et al. show GPT-3 splitting 42235630 into [422, 35, 630]. Change one digit and the split can change completely, so two nearly equal values can look unrelated to the model.
LLMTime's fix was to put a space between every digit, so each digit becomes its own token, and to separate time steps with commas. It worked for GPT-3. GPT-4's tokenizer handled digits differently, and the trick did not carry over.
Why do errors creep in over a long forecast?
A language model writes its forecast one token at a time, and each token is conditioned on the ones it has already written.
There is no arithmetic unit doing the sums. The model is predicting which digit usually comes next. Gruver et al. list weak arithmetic as a limitation of the approach. In our reading, this is why long horizons are hard: a small slip early on, a misplaced digit or a drifting level, becomes part of the context for every step after it.
Forecasting models avoid this by producing the whole horizon as numbers, in one pass or in fixed-size patches, with no text in between.
Does a language model know what is normal for a series?
Not in the way a forecasting model does. It has read a great deal about the world, but not been trained to predict values from long runs of other values.
Tan et al. tested this directly (NeurIPS 2024). They took three published methods that put a pretrained LLM inside a forecasting architecture (OneFitsAll, Time-LLM and CALF) and removed the LLM, or swapped it for a basic attention layer. Accuracy did not drop, and in most cases it improved. They also report that pretrained LLMs did no better than models trained from scratch and did not represent the sequential dependencies in the series.
Note the scope: this study is about LLMs built into trained forecasting pipelines, not about prompting a chat model. It shows that the language pretraining was not what made those methods work.
Is a language model's uncertainty trustworthy?
Not by default. A forecast is only useful for decisions if its stated uncertainty matches how often it is wrong, which is called calibration.
Gruver et al. found GPT-4 had worse uncertainty calibration than GPT-3 and suggest alignment (RLHF) as the likely cause. Alignment teaches a model to give confident, helpful answers; that is a different goal from spreading probability honestly across outcomes.
When you ask a chat model for "a 90% range", it writes numbers that look like a range. Nothing in its training checks that 90% of outcomes fall inside it. A forecasting model that outputs quantiles is trained against a loss that does check this. Why that matters is covered in agents that decide under uncertainty.
What about context windows, cost and latency?
Numbers are expensive in tokens. A digit-per-token encoding turns a few thousand data points into many thousands of tokens, which fills the context window and the bill.
Gruver et al. list the context window as a limit on how much history fits, and note multivariate series are especially hard for that reason. They removed decimal points partly to save context.
The Context is Key authors put the cost point plainly: "LLMs require significant computational power, making them unsuitable for real-world practical forecasting at scale where speed and cost matter." They also found many LLM forecasters were Pareto-dominated, meaning another method was both cheaper and at least as accurate, by quantitative models such as Lag-Llama and Chronos.
Tan et al. give a training example: Time-LLM has 6,642M parameters and took 3,003 minutes to train on one dataset, against 0.245M parameters and 2.17 minutes for the ablations that matched it. For comparison, the time-series foundation models in the Ephemeris panel range from 9M parameters (FlowState r1) to 385M (PatchTST-FM r2).
Where do language models genuinely help?
When the forecast depends on information that exists only as text. Here the evidence is in the LLM's favour.
The Context is Key benchmark (Williams et al., ICML 2025) has 71 tasks across seven domains. Each one is built so that the written context is essential: a note that a sensor will be offline, that a policy changes next week, that a value cannot go above a cap. On these tasks, carefully prompted LLMs beat both statistical models and time-series foundation models; the top method was Llama-3.1-405B-Instruct with a direct prompt.
Two findings from the same paper keep this in proportion. First, with the context removed, LLM forecasters "no longer dominate" the quantitative models; some stayed competitive, while the top models were significantly weaker. Second, LLMs sometimes failed badly, overshooting or undershooting by at least 500% of the ground-truth range.
So the LLM's advantage is reading the context, not doing the numbers.
So what should you do?
Split the work. Let the LLM do what it is good at, and hand the numbers to a model built for them.
The language model gathers the history, works out the frequency and horizon, and reads the context. Where it can, it turns that context into numeric inputs a forecasting model accepts, such as a holiday flag or a planned price as a covariate (an extra series known in advance). The forecasting model returns quantiles. The language model then explains the band in words and helps the user decide.
The pattern, with a full request, is in the split that works. If the context cannot be expressed as numbers, a context-aware LLM forecast, checked against a numeric baseline, is a reasonable option per Context is Key.
FAQ
Why can't ChatGPT or Claude forecast numbers accurately?
They read numbers as text tokens that split digits inconsistently, they produce forecasts digit by digit without arithmetic, and their stated uncertainty is not calibrated. Purpose-built forecasting models read numbers directly, produce the horizon as numbers, and are trained with a loss that scores their quantiles.
Did any research show LLMs can forecast?
Yes. LLMTime (Gruver et al., NeurIPS 2023) showed GPT-3 and LLaMA-2 forecasting zero-shot at a level comparable to specialised models of the time, using digit-by-digit encoding and rescaling. GPT-4 did worse in that setup, partly due to tokenization and calibration.
When is an LLM better than a forecasting model?
When essential information exists only as text, such as a planned outage or a known cap. On the Context is Key benchmark, prompted LLMs beat statistical and foundation models on such tasks, but lost that lead when the text was removed.
What is the best way to combine an LLM and a forecasting model?
Let the LLM collect the data, turn known future events into covariates, and call a forecasting model for quantiles. The LLM then explains the result and the uncertainty band.