---
title: The split that works: the LLM reads the context, a forecasting model does the numbers
description: An architecture for AI agents that forecast: the LLM gathers data and turns context into covariates, a forecasting model returns quantiles.
date: 2026-10-07
updated: 2026-10-07
summary: An agent should not write forecasts itself. It should gather the history, work out the frequency and horizon, turn known context (holidays, planned promotions, prices, outages) into covariate arrays, call a forecasting model, and explain the returned quantile band in words. This post shows the pattern with a complete covariate request and an illustrative agent walk-through.
tags: agents, llm, covariates
draft: false
---

## Key facts

- A *covariate* is an extra series that helps explain the target, such as a promotion flag or a price. Ephemeris accepts up to 16 covariates per series.
- `past` covariates are aligned with the history (same length as `values`); `future` covariates hold exactly `horizon` values known in advance. Every `future` key must also appear in `past`.
- In the Ephemeris panel only Chronos-2 and TiRex-2 accept covariates; in `route` and `ensemble` mode a request with covariates is narrowed to those models.
- A request can carry 1 to 64 series, a horizon of 1 to 512 steps and up to 21 quantile levels strictly between 0 and 1.
- On the Context is Key benchmark (71 tasks), prompted LLMs beat purely numeric models when the written context was essential, and lost that lead when it was removed ([Williams et al., ICML 2025](https://arxiv.org/abs/2410.18959)).
- The agent must send only the history it actually has: no invented, padded or interpolated values.

## Why split the work between an LLM and a forecasting model?

Because each is good at the part the other is bad at.

Language models read numbers as text fragments and their stated uncertainty is not calibrated; we cover the research in [why language models are bad at forecasting numbers](/blog/why-language-models-are-bad-at-time). Forecasting models are trained to turn history into quantiles, but they cannot read an email saying the store runs a promotion next Thursday.

The [Context is Key benchmark](https://arxiv.org/abs/2410.18959) shows the same division from the other side: LLMs led when essential information was in text, and lost the lead without it. The split keeps the LLM on the text and gives the numbers to a forecaster.

## What does the agent do, step by step?

Six steps, in this order:

1. **Gather the history.** Pull the series from the user, a file or a database tool. Keep it oldest first and note the timestamps.
2. **Work out frequency and horizon.** Read the frequency from the timestamps (daily is `"D"`, hourly `"H"`). Convert the user's question ("next two weeks") into a number of steps (14 daily steps).
3. **Turn context into covariates.** For each known factor, build one array over the history and one over the horizon: 1 on promotion days and 0 otherwise, the planned price per day, 1 during a known outage.
4. **Check the request.** Lengths must match, the horizon must be within limits, and covariates need a model that supports them.
5. **Call the forecaster.** Ask for the quantiles the decision needs, for example 0.1, 0.5 and 0.9.
6. **Explain the band.** Report the median and the range, say which context was used, and say what was not.

## What can be a covariate, and what cannot?

A covariate must be a number for every past step and, if it is a `future` covariate, for every forecast step.

Good candidates:

- Calendar flags: public holidays, school holidays, paydays, a store closure.
- Planned actions: promotion days, a scheduled price, a marketing send.
- Forecasts of drivers: a weather forecast for temperature.
- Known disruptions: an announced maintenance window, as a 0/1 flag.

Things that are not covariates: a vague sense that "demand may pick up", a competitor rumour, or anything without a value for each step. The agent should mention these to the user in words, not encode them as guesses.

A covariate only helps if the history shows its effect. A holiday flag that is 0 for every past day gives the model nothing to learn from. Past covariates must be what actually happened, so record the promotions that ran, not the ones that were planned.

## What does a request with covariates look like?

Here is a complete request for four weeks of daily store sales, starting on a Monday, with a promotion flag and a holiday flag, forecasting the next 7 days.

History: 28 values. The promotion ran Thursday to Saturday of week 2, and the Monday of week 3 was a public holiday. In the coming week, Monday is a public holiday and a promotion runs Thursday to Saturday.

```bash
curl --request POST \
  --url https://ephemeris.cascade.industries/api/v1/forecast \
  --header "Authorization: Bearer $EPHEMERIS_API_KEY" \
  --header "Content-Type: application/json" \
  --header "Idempotency-Key: store-12-week5-001" \
  --data '{
    "mode": "route",
    "series": [{
      "values": [142, 138, 145, 151, 176, 214, 198,
                 140, 136, 149, 188, 205, 236, 201,
                 96, 141, 147, 155, 179, 219, 203,
                 144, 140, 150, 158, 181, 222, 206],
      "freq": "D",
      "covariates": {
        "past": {
          "promo":   [0, 0, 0, 0, 0, 0, 0,
                      0, 0, 0, 1, 1, 1, 0,
                      0, 0, 0, 0, 0, 0, 0,
                      0, 0, 0, 0, 0, 0, 0],
          "holiday": [0, 0, 0, 0, 0, 0, 0,
                      0, 0, 0, 0, 0, 0, 0,
                      1, 0, 0, 0, 0, 0, 0,
                      0, 0, 0, 0, 0, 0, 0]
        },
        "future": {
          "promo":   [0, 0, 0, 1, 1, 1, 0],
          "holiday": [1, 0, 0, 0, 0, 0, 0]
        }
      }
    }],
    "horizon": 7,
    "quantiles": [0.1, 0.5, 0.9]
  }'
```

Check it against the rules: `values` and each `past` array have 28 entries, each `future` array has 7 entries (the horizon), and every `future` key also appears in `past`.

Four weeks is short and is used here so the example fits on screen. In practice send as much real history as you have: covariate effects are only learnable when the history contains several examples of them.

## What comes back?

One entry per input series, mapping each quantile level to `horizon` values, plus metadata:

```json
{
  "forecasts": [
    {
      "quantiles": {
        "0.1": ["...7 values..."],
        "0.5": ["...7 values..."],
        "0.9": ["...7 values..."]
      }
    }
  ],
  "meta": {
    "models_used": ["chronos2"],
    "billing": { "settled_mc": "...", "balance_mc": "..." }
  }
}
```

The placeholders stand in for numbers. `meta.models_used` tells you which model ran; with covariates in `route` mode it will be a covariate-capable model. `meta.notes`, when present, explains any model that was skipped.

## What does this look like inside an agent?

The walk-through below is **illustrative**: it shows the shape of a good exchange, not a recorded session or real output.

1. **User:** "How many units will store 12 sell next week? We're running the Thursday-to-Saturday promo again, and Monday is a bank holiday."
2. **Agent (tool call):** queries the sales table for store 12's daily totals. Gets 28 days, Monday to Sunday, no gaps.
3. **Agent (reasoning):** daily data, so `freq` is `"D"`; next week is 7 steps. Two known factors: promotion on days 4 to 6 of the horizon, holiday on day 1. History shows one past promotion (days 11 to 13) and one past holiday (day 15), so both can be covariates.
4. **Agent (tool call):** `list_models` once for the session, sees covariate support on Chronos-2 and TiRex-2. Then `forecast` with the request above.
5. **Agent (to user):** "Median forecast for next week is X units in total, about Y per day. On Monday expect fewer sales than a normal Monday; on the promotion days, more. There is a 1 in 10 chance any given day comes in below its 0.1 line and a 1 in 10 chance above its 0.9 line. This uses only four weeks of history with one past promotion, so treat the promotion uplift as rough."

In that last message the agent reports the band, not just the median; says what context it used; and flags the weakness in the data. Note that adding up the daily 0.9 quantiles does not give a 0.9 quantile for the week; if the user needs a weekly total range, forecast weekly totals instead. [Agents that decide under uncertainty](/blog/agents-that-decide-under-uncertainty) explains why.

## What must the agent never do?

These rules keep the forecaster's output honest:

- **Do not invent or pad history.** If 3 days are missing, do not fill them with guesses to make the series look complete. Send what exists, or ask the user.
- **Do not write the forecast itself.** If the tool call fails, say so; do not produce numbers "in the meantime".
- **Do not encode guesses as covariates.** A covariate is a known value, not a hunch about the future.
- **Do not let covariate arrays drift out of alignment.** Off-by-one errors shift a promotion onto the wrong day, silently.
- **Do not forecast in a loop without the user's intent.** Every successful forecast spends credits; use an `Idempotency-Key` when retrying so a retry is not charged twice.

## Do I need Ephemeris for this pattern?

No. The pattern works with any forecasting model that accepts covariates.

You can self-host Chronos-2 from [Hugging Face](https://huggingface.co/amazon/chronos-2), which supports past-only and known-future covariates, and expose it to your agent as a tool. A classical model with regressors, such as [Prophet](https://facebook.github.io/prophet/) with holiday effects, also fits the same slot. Ephemeris is the example here because it exposes several models behind one request format and a remote MCP server; setup is in [a forecasting tool for AI agents over MCP](/blog/forecasting-tool-for-ai-agents-mcp).

The quickest way to try this pattern is to give your agent the Ephemeris MCP server ([setup for each client](/blog/forecasting-tool-for-ai-agents-mcp)) and ask it to forecast a series you already have.

## FAQ

### What is a covariate in time-series forecasting?

A covariate is an extra series that helps explain the one you are forecasting, such as a promotion flag, a price or a temperature. A past covariate is observed alongside the history; a future covariate is known in advance for each forecast step.

### How long must future covariates be?

In the Ephemeris API each `future` covariate array must have exactly `horizon` values, and each `past` array must be the same length as the history. Every `future` key must also appear in `past`.

### Which models accept covariates?

In the Ephemeris panel, Chronos-2 and TiRex-2. In `route` and `ensemble` mode a request with covariates uses only those models; in `explicit` mode naming a model without covariate support is an error.

### Should the agent fill gaps in the history?

No. Send only the values you have, or ask the user. Invented values become part of what the model treats as real.

## Related

- [Why language models are bad at forecasting numbers](/blog/why-language-models-are-bad-at-time)
- [Agents that decide under uncertainty](/blog/agents-that-decide-under-uncertainty)
- [A forecasting tool for AI agents over MCP](/blog/forecasting-tool-for-ai-agents-mcp)
- [Forecasting in agent frameworks](/blog/forecasting-in-agent-frameworks)
- [Forecast store sales with promotions](/blog/forecast-store-sales-with-promotions)
- [API reference: covariates and limits](/llms-full.txt)
