Blog · Updated
How to forecast store sales with promotions and holidays as covariates
Short answer
Store sales move with the calendar you already know: promotions, holidays and opening days. Send those as covariates, with history in `past` and the planned values over the forecast window in `future`, and the forecast can react to them. This tutorial does it for 1,115 Rossmann stores, 64 stores per request, then uses a high quantile instead of the median to set stock.
Key facts
- A covariate is an extra series that helps explain the target; a known-future covariate, such as a planned promotion or a public holiday, has values you know in advance over the forecast window.
- In the Ephemeris API,
covariates.pastmust be the same length asvaluesandcovariates.futuremust be exactlyhorizonlong; everyfuturekey must also appear inpast, with up to 16 covariates per series. - Two models in the Ephemeris panel use covariates: Chronos-2 (Amazon, 120M parameters) and TiRex-2 (NXAI);
routeandensemblemode narrow to those models when a request carries covariates. - One Ephemeris request takes at most 64 series and 256 slots, so 1,115 stores need 18 requests.
- The Rossmann Store Sales data on Kaggle covers 1,115 stores, daily, from 2013-01-01 to 2015-07-31, with
Promo,StateHoliday,SchoolHolidayandOpencolumns. - For stock, the right quantile is the one that balances the cost of running out against the cost of holding extra: cost of a lost sale ÷ (cost of a lost sale + cost of a leftover unit).
What makes store sales hard to forecast?
Store sales are driven mostly by things you already know about: the day of the week, whether the store is open, whether a promotion is running and whether it is a holiday. The hard part is not finding the pattern, it is telling the model about the calendar ahead.
A model that sees only past sales will learn the weekly cycle. It cannot know that next Tuesday has a promotion unless you tell it. That is what covariates are for.
The other difficulties are practical:
- Closed days. A closed store sells nothing. Those zeros are not demand, and a model that treats them as demand learns a wrong level.
- Gaps. Some stores have months missing from their history.
- Scale. A chain has hundreds or thousands of stores, each its own series.
- The decision is not the median. You stock to a level that covers most days, not the typical day.
Which dataset does this tutorial use?
It uses Rossmann Store Sales, a 2015 Kaggle competition from a German drugstore chain. You need a Kaggle account and must accept the competition rules to download it; those rules govern how you may use the data.
train.csv has 1,017,209 rows with the columns Store, DayOfWeek, Date, Sales, Customers, Open, Promo, StateHoliday and SchoolHoliday. StateHoliday is a (public holiday), b (Easter), c (Christmas) or 0 (none). test.csv holds the 48 days after training (2015-08-01 to 2015-09-17) with the same calendar columns but no sales.
Sales is daily turnover per store, not units per product. The stock example below is therefore illustrative: in practice you would forecast units per product per store, but the method is the same.
About 180 stores have no rows for the second half of 2014. The code below handles that by forecasting from each store's latest unbroken stretch of days.
How do I send promotions and holidays as covariates?
Put each covariate's history in covariates.past, aligned with values, and its planned values over the forecast window in covariates.future. The rules in the API are strict, so check them before you build the request:
past.<name>has the same length asvalues.future.<name>has exactlyhorizonvalues.- Every key in
futurealso appears inpast. - Covariates are numbers, so encode flags as 0 and 1.
For Rossmann we use four covariates, all known in advance: Open, Promo, StateHoliday (as a 0/1 flag) and SchoolHoliday. To score the forecast honestly, the code holds back the last 48 days of train.csv and treats their calendar columns as the known future. In production you would use test.csv (or your own promotion plan) as future and send all of the history.
import os
import numpy as np
import pandas as pd
import requests
API = "https://ephemeris.cascade.industries/api/v1/forecast"
KEY = os.environ["EPHEMERIS_API_KEY"]
HORIZON = 48 # days held back and forecast
CONTEXT = 728 # at most 104 weeks of history per store
QUANTILES = [0.1, 0.5, 0.9, 0.95]
COVARIATES = ["Open", "Promo", "StateHoliday", "SchoolHoliday"]
train = pd.read_csv("train.csv", parse_dates=["Date"], dtype={"StateHoliday": str})
train["StateHoliday"] = (train["StateHoliday"] != "0").astype(int)
train = train.sort_values(["Store", "Date"])
def latest_stretch(store_df):
"""Rows after the store's last missing day, so the series has no gaps."""
store_df = store_df.reset_index(drop=True)
starts = store_df["Date"].diff().dt.days.ne(1) # True at row 0 and after every gap
return store_df.loc[starts[starts].index[-1]:]
series, holdout = [], []
for store, g in train.groupby("Store"):
g = latest_stretch(g)
if len(g) < HORIZON + 8 * 7: # need at least 8 weeks of history
continue
hist = g.iloc[:-HORIZON].tail(CONTEXT)
fut = g.iloc[-HORIZON:]
series.append({
"values": hist["Sales"].astype(float).tolist(),
"freq": "D",
"covariates": {
"past": {c: hist[c].astype(float).tolist() for c in COVARIATES},
"future": {c: fut[c].astype(float).tolist() for c in COVARIATES},
},
})
holdout.append((store, fut))Closed days stay in the series as zeros, with Open telling the model why. Dropping them would break the regular daily spacing that freq: "D" promises.
How do I forecast hundreds of stores at once?
Send them in batches of up to 64 series per request; the API limit is 64 series and 256 slots, and a univariate series is one slot. The response has one forecast per input series, in the same order, so you can zip results back to stores.
def forecast(body, idempotency_key):
r = requests.post(
API,
headers={
"Authorization": f"Bearer {KEY}",
"Content-Type": "application/json",
"Idempotency-Key": idempotency_key,
},
json=body,
timeout=300,
)
r.raise_for_status()
return r.json()
results = []
for i in range(0, len(series), 64):
body = {
"mode": "route",
"series": series[i:i + 64],
"horizon": HORIZON,
"quantiles": QUANTILES,
"context_len": CONTEXT,
}
out = forecast(body, f"rossmann-holdout-batch-{i // 64:03d}")
results.extend(out["forecasts"])
print(i, out["meta"]["models_used"], out["meta"]["billing"]["settled_mc"], "mc")The idempotency key matters for batch jobs. If a request times out and you retry with the same key and body, the API replays the stored result instead of charging again.
context_len defaults to 256 points. For daily retail data that is only about eight months, which cuts off last year's holiday season, so the code sets it to 728 days. Cost scales with ceil(context / 1024) and ceil(horizon / 64), so 728 days and 48 steps sit in the cheapest bucket of each; see /pricing for the per-model rates.
Because these requests carry covariates, route mode only considers the covariate-capable models (Chronos-2 and TiRex-2). meta.models_used tells you which one ran.
How do I check whether the forecast is any good?
Score the held-back 48 days with interval coverage: the share of actual days that fall inside the forecast band. For a 0.1 to 0.9 band, about 80% of days should land inside. Much lower means the bands are too narrow; much higher means they are wider than they need to be.
rows = []
for (store, fut), fc in zip(holdout, results):
q = {k: np.array(v) for k, v in fc["quantiles"].items()}
is_open = fut["Open"].to_numpy() == 1
actual = fut["Sales"].to_numpy()[is_open]
lo, mid, hi, p95 = (np.clip(q[k][is_open], 0, None) for k in ("0.1", "0.5", "0.9", "0.95"))
rows.append({
"store": store,
"inside_80": np.mean((actual >= lo) & (actual <= hi)),
"at_or_below_p95": np.mean(actual <= p95),
"abs_pct_error_median": np.mean(np.abs(actual - mid) / np.maximum(actual, 1)),
})
scores = pd.DataFrame(rows)
print(scores[["inside_80", "at_or_below_p95", "abs_pct_error_median"]].describe())Only open days are scored. On a closed day the answer is zero and you already know it, so after forecasting you should set closed days to zero rather than trust the model's number.
We have not run this backtest, so we do not quote a result for Rossmann. Run it yourself, and run a simple baseline next to it (for example, the same weekday's sales from the last four weeks) so the numbers have something to be compared with.
Which quantile should I use for stock?
Use the quantile that matches your costs, not the median. If running out costs you the margin on a lost sale and a leftover unit costs you holding or waste, the stock level that minimises expected cost is the forecast quantile at the critical ratio: cost of a lost sale ÷ (cost of a lost sale + cost of a leftover unit).
This is the classic newsvendor result. A lost sale that costs 4 times more than a leftover unit gives 4 / (4 + 1) = 0.8, so you stock to the 0.8 quantile. Perishable goods with expensive waste push the ratio down; high-margin goods with cheap storage push it up. Ask for that quantile directly in quantiles rather than interpolating between others.
Two cautions:
- Sums of quantiles are not quantiles of sums. If you replenish weekly, the 0.95 quantile of weekly demand is usually lower than the sum of seven daily 0.95 quantiles, because bad days rarely all happen together. Summing daily upper quantiles over-stocks. If the decision is weekly, forecast weekly totals (
freq: "W") instead. - Check calibration first. A 0.95 quantile is only useful if actual sales fall at or below it about 95% of the time. The
at_or_below_p95column above measures exactly that.
What are the alternatives?
Several approaches work well for retail sales, and some will beat a zero-shot model when you have rich data:
- Gradient-boosted trees (LightGBM, XGBoost) with lag, calendar and promotion features. They need feature engineering and a training pipeline, but they use store attributes (
store.csvhas store type, assortment and competitor distance), which the Ephemeris API has no field for. - Classical models such as ETS and ARIMA with regressors, for example in Nixtla's open-source statsforecast. Fast, well understood, one model per series.
- Self-hosting Chronos-2: it is Apache-2.0 on Hugging Face (amazon/chronos-2) and supports covariates. If you already run GPUs, this avoids an API altogether.
- A zero-shot API like Ephemeris, when you want probabilistic forecasts across many stores without training or hosting anything.
In our view, the right choice depends on how much you want to own. Trees with good features are hard to beat on a well-studied dataset like Rossmann; a pretrained model is quicker to stand up for a new chain or a new product line with short history.
If you would rather not host the models, the requests above run as written against Ephemeris: create an account (new accounts get a small free credit), make a key, and set EPHEMERIS_API_KEY.
FAQ
Can I send promotions that are planned for the future?
Yes. Put the promotion flag's history in covariates.past and the planned flags for the forecast window in covariates.future. The future array must be exactly horizon values long and its key must also appear in past.
Should I remove days when the store was closed?
No, keep them so the series stays evenly spaced, and add an Open flag as a covariate. After forecasting, set the forecast to zero on days you know the store is closed.
How many stores can I forecast in one request?
Up to 64 series and 256 slots per request; a univariate store series is one slot. For 1,115 stores, send 18 requests and match results to stores by position.
Which models use covariates?
In the Ephemeris panel, Chronos-2 and TiRex-2. In route and ensemble mode the panel narrows to them automatically; in explicit mode, naming a model without covariate support is an error.
Which quantile should I stock to?
The critical ratio: cost of a lost sale divided by the sum of that cost and the cost of a leftover unit. A ratio of 0.8 means stocking to the 0.8 quantile, provided your backtest shows the quantiles are calibrated.