ephemerisBlog

Blog · Updated

Forecasting electricity load and solar output with prediction intervals

Short answer

Electricity load repeats every day and every week, and moves with temperature; solar output follows the sun and is zero at night. A good forecast gives a band, not a line, because grid and trading decisions are about the risk of being short. This tutorial forecasts German hourly load and solar from Open Power System Data with temperature and irradiance as covariates, and shows how to handle 15-minute data and night-time zeros.

forecasting · energy · electricity · solar · covariates · tutorialRead as markdown

Key facts

  • Hourly load has a daily cycle of 24 steps and a weekly cycle of 168 steps; at 15-minute resolution those cycles are 96 and 672 steps.
  • The Ephemeris horizon is counted in steps of the series' frequency, from 1 to 512: 48 hours is 48 steps at freq: "H" and 192 steps at freq: "15min".
  • context_len defaults to 256 points, which is under 11 days of hourly data; for weekly seasonality, set it to several weeks (2,016 hourly points is 12 weeks).
  • Temperature and irradiance are known-future covariates only when you have a weather forecast for the horizon; using measured weather in a backtest assumes a perfect weather forecast.
  • In the Ephemeris panel, Chronos-2 (context cap 8,192 points) and TiRex-2 (context cap 2,048 points) accept covariates.
  • Open Power System Data's time series package (version 2020-10-06) has load and solar in MW for 32 European countries from 2015 to mid-2020, at 60-minute resolution and, for some countries including Germany, 15-minute resolution.

What matters when forecasting load and solar?

Seasonality, weather and the cost of being wrong in one direction. Load has three overlapping cycles: the day (people wake, work, cook), the week (weekends are lower) and the year (heating and cooling). Solar has the day and the year, set by the sun, scaled by cloud cover.

Weather explains much of what the cycles do not. Load in Germany rises in cold spells; solar collapses under cloud. A forecaster that sees only past load cannot know tomorrow is colder than today.

The decisions that use these forecasts are asymmetric. A grid operator needs reserves for the high end of load. A solar trader who sells power they do not produce pays an imbalance penalty. That is why the useful output is a prediction interval: a range that should contain the outcome a stated share of the time. Quantiles give you that: the 0.05 and 0.95 quantiles bound a 90% interval.

Which data does this tutorial use?

It uses two packages from Open Power System Data (OPSD):

  • Time series, version 2020-10-06: load, wind and solar per country from the ENTSO-E Transparency Platform, in MW. The columns used here are DE_load_actual_entsoe_transparency, DE_solar_generation_actual and DE_load_forecast_entsoe_transparency (the published day-ahead load forecast, useful as a baseline).
  • Weather data, version 2020-09-16: hourly, population-weighted country averages from NASA's MERRA-2 reanalysis, through 2019. The columns used are DE_temperature (°C), DE_radiation_direct_horizontal and DE_radiation_diffuse_horizontal (W/m²).

Both use a utc_timestamp column. The weather data ends in 2019, so the examples work inside 2019.

On licensing: OPSD asks you to cite each package, and its legal page explains that the underlying ENTSO-E data come without an open licence; ENTSO-E's own terms apply. That is fine for learning on your own machine. Check those terms before you republish the data or build on it commercially.

The files are large (the package pages list the hourly time series at 124 MB and the weather file at 223 MB), so download them once and read them locally.

How do I set up the request?

Choose a forecast origin, take the history just before it as values, and pass the weather over the forecast window as covariates.future. Here is a function that builds one series for any origin, so the same code serves one forecast or a backtest over many origins.

python
import os

import numpy as np
import pandas as pd
import requests

API = "https://ephemeris.cascade.industries/api/v1/forecast"
KEY = os.environ["EPHEMERIS_API_KEY"]

TS_URL = "https://data.open-power-system-data.org/time_series/2020-10-06/time_series_60min_singleindex.csv"
WX_URL = "https://data.open-power-system-data.org/weather_data/2020-09-16/weather_data.csv"

ts = pd.read_csv(
    TS_URL,
    usecols=["utc_timestamp", "DE_load_actual_entsoe_transparency",
             "DE_solar_generation_actual", "DE_load_forecast_entsoe_transparency"],
    parse_dates=["utc_timestamp"], index_col="utc_timestamp",
).rename(columns={
    "DE_load_actual_entsoe_transparency": "load",
    "DE_solar_generation_actual": "solar",
    "DE_load_forecast_entsoe_transparency": "tso_forecast",
})
wx = pd.read_csv(
    WX_URL,
    usecols=["utc_timestamp", "DE_temperature",
             "DE_radiation_direct_horizontal", "DE_radiation_diffuse_horizontal"],
    parse_dates=["utc_timestamp"], index_col="utc_timestamp",
)
wx["irradiance"] = wx["DE_radiation_direct_horizontal"] + wx["DE_radiation_diffuse_horizontal"]
wx = wx.rename(columns={"DE_temperature": "temperature"})[["temperature", "irradiance"]]

df = ts.join(wx, how="inner").loc["2019-01-01":"2019-12-31"]

HORIZON = 48       # hours
CONTEXT = 2016     # 12 weeks of hourly history
QUANTILES = [0.05, 0.1, 0.5, 0.9, 0.95]


def make_series(origin, target, covariate):
    hist = df.loc[df.index < origin].tail(CONTEXT)
    fut = df.loc[df.index >= origin].head(HORIZON)
    if len(fut) != HORIZON or hist[[target, covariate]].isna().any().any():
        raise ValueError(f"incomplete data around {origin}")
    if fut[covariate].isna().any():
        raise ValueError(f"missing future {covariate} at {origin}")
    return {
        "values": hist[target].tolist(),
        "freq": "H",
        "covariates": {
            "past": {covariate: hist[covariate].tolist()},
            "future": {covariate: fut[covariate].tolist()},
        },
    }, fut


def forecast(body, idempotency_key):
    r = requests.post(
        API,
        headers={
            "Authorization": f"Bearer {KEY}",
            "Content-Type": "application/json",
            "Idempotency-Key": idempotency_key,
        },
        json=body,
        timeout=300,
    )
    r.raise_for_status()
    return r.json()


origin = pd.Timestamp("2019-06-17 00:00", tz="UTC")
load_series, load_fut = make_series(origin, "load", "temperature")
solar_series, solar_fut = make_series(origin, "solar", "irradiance")

out = forecast(
    {
        "mode": "route",
        "series": [load_series, solar_series],
        "horizon": HORIZON,
        "quantiles": QUANTILES,
        "context_len": CONTEXT,
    },
    "opsd-de-2019-06-17",
)
load_fc, solar_fc = out["forecasts"]
print(out["meta"]["models_used"], out["meta"].get("notes"))

Each series carries its own covariates, so load (with temperature) and solar (with irradiance) can share one request. The response returns forecasts in the order you sent them.

Why does context length matter so much here?

Because the model can only reuse a pattern it can see. With the default context_len of 256 hourly points, the model sees less than 11 days, which is barely one full weekly cycle. Setting it to 2,016 points shows it 12 weeks: twelve examples of a Monday morning ramp, twelve weekends.

Longer is not free. Cost scales with ceil(context / 1024), so 2,016 points costs twice what 1,024 does, and each model truncates to its own cap (2,048 points for TiRex-2, 8,192 for Chronos-2). Annual seasonality at hourly resolution would need 8,760 points, more than most models read; weather covariates carry most of the seasonal signal instead.

What are the units of the horizon?

Steps, not hours. The horizon is the number of future points at the series' frequency, from 1 to 512. At hourly data, 48 steps is two days. At 15-minute data, two days is 192 steps and the 512-step limit is about 5.3 days.

To work at 15 minutes with OPSD, read time_series_15min_singleindex.csv instead, set freq to "15min", and bring the hourly weather onto the same grid. Interpolating a smooth variable like temperature between hourly points is reasonable; the result is still a forecast input, not an invented observation of your target.

python
TS15_URL = "https://data.open-power-system-data.org/time_series/2020-10-06/time_series_15min_singleindex.csv"
# Read it as above, then align the hourly weather to 15-minute steps:
wx15 = wx.resample("15min").interpolate()
HORIZON_15 = 192   # two days of 15-minute steps
CONTEXT_15 = 2688  # four weeks: 4 x 672

How do I handle solar's night-time zeros?

Post-process them with what you know: the sun is down. A statistical model may put a small positive (or even negative) quantile on a night hour, because it smooths across the daily cycle. Physics says output is zero, so set it to zero and clip every quantile at zero.

python
solar_q = {k: np.clip(np.array(v), 0, None) for k, v in solar_fc["quantiles"].items()}
night = solar_fut["irradiance"].to_numpy() <= 0
for k in solar_q:
    solar_q[k][night] = 0.0

Here the night mask comes from the irradiance covariate. Without one, compute sunrise and sunset or a clear-sky profile for the site; the open-source pvlib library does both.

Two more solar-specific points:

  • Installed capacity grows. German solar output in 2019 is not comparable to 2015 because more panels were built. The OPSD time series package has DE_solar_capacity and DE_solar_profile (generation as a share of capacity) columns; forecasting the profile and multiplying by current capacity removes that drift.
  • Score day hours only. Night hours are trivially correct and inflate any accuracy metric. Evaluate coverage and error on hours with the sun up.

Where do the weather covariates come from in production?

From a weather forecast, not from measurements. In this tutorial, covariates.future holds MERRA-2 reanalysis values: what the weather actually was. That is a perfect weather forecast, which you will never have, so a backtest built this way overstates accuracy, especially for solar and beyond a day or two.

In production, fill future with a numerical weather prediction for the horizon (your national weather service or a commercial provider) and keep past as observed weather. If you can, backtest with archived weather forecasts instead of reanalysis, so the backtest sees the same quality of input that production will.

How do I check the intervals are calibrated?

Run many origins and count how often the actual value lands inside each interval. A 90% interval (0.05 to 0.95) should contain about 90% of actual hours; an 80% interval (0.1 to 0.9) about 80%.

python
origins = pd.date_range("2019-07-01", "2019-12-23", freq="7D", tz="UTC")  # 26 Mondays
batch, futures = [], []
for o in origins:
    s, f = make_series(o, "load", "temperature")
    batch.append(s)
    futures.append(f)

out = forecast(
    {"mode": "route", "series": batch, "horizon": HORIZON,
     "quantiles": QUANTILES, "context_len": CONTEXT},
    "opsd-de-load-backtest-2019h2",
)

inside_90, inside_80, ae_model, ae_tso = [], [], [], []
for fc, fut in zip(out["forecasts"], futures):
    q = {k: np.array(v) for k, v in fc["quantiles"].items()}
    y = fut["load"].to_numpy()
    inside_90.append(np.mean((y >= q["0.05"]) & (y <= q["0.95"])))
    inside_80.append(np.mean((y >= q["0.1"]) & (y <= q["0.9"])))
    ae_model.append(np.mean(np.abs(y - q["0.5"])))
    ae_tso.append(np.mean(np.abs(y - fut["tso_forecast"].to_numpy())))

print("90% interval coverage:", np.mean(inside_90))
print("80% interval coverage:", np.mean(inside_80))
print("MAE median vs published day-ahead forecast (MW):", np.mean(ae_model), np.mean(ae_tso))

Twenty-six origins times 48 hours is 1,248 hourly checks, enough to see a badly miscalibrated band but not enough to separate small differences. We have not run this, so we do not quote numbers for it.

Read the comparison with the published day-ahead forecast carefully. That forecast was made a day ahead with the weather forecasts available then; the code above gives the model the actual weather. To compare fairly, either use archived weather forecasts or drop the covariates and compare history-only forecasts.

What are the alternatives?

Load forecasting is a mature field and the established methods are strong:

  • Regression with weather and calendar features, often gradient-boosted trees, is the standard in utilities. It captures holidays and temperature effects explicitly and is easy to explain.
  • Multiple-seasonality statistical models such as MSTL, TBATS or Prophet handle the daily and weekly cycles; Nixtla's statsforecast includes MSTL.
  • Physical models for solar: irradiance forecast times a PV system model (pvlib), which needs site and panel data but extrapolates well.
  • Self-hosting Chronos-2 or TiRex-2, both on Hugging Face, if you want the pretrained-model approach without an API.
  • A hosted zero-shot API like Ephemeris, when you want probabilistic forecasts for many meters, feeders or sites without training a model for each.

In our experience, the biggest gains in this domain come from the quality of the weather forecast and the holiday calendar, whichever model sits on top.

The code above runs against Ephemeris as written. To try it on your own meters, create an account; new accounts get a small free credit, enough for a first backtest on a handful of series.

FAQ

How long should the history be for hourly load?

Long enough to contain several weekly cycles: at least 4 weeks (672 points), and 8 to 12 weeks if the model reads it. Set context_len explicitly, because the default of 256 points is under 11 days.

Can I use temperature as a covariate?

Yes. Send observed temperature in covariates.past and the temperature forecast for the horizon in covariates.future, exactly horizon values long. Only covariate-capable models (Chronos-2 and TiRex-2 in the Ephemeris panel) use it.

Why does my solar forecast show output at night?

Statistical models smooth across the daily cycle and can put small values on night hours. Clip quantiles at zero and set night hours to zero using sunrise and sunset times or an irradiance forecast.

How far ahead can I forecast 15-minute data?

The Ephemeris API allows up to 512 steps, which is about 5.3 days of 15-minute data or 21 days of hourly data. Accuracy falls with distance, and weather covariates past a few days are themselves uncertain.

Should I forecast in UTC or local time?

Send the series in UTC so there are no duplicated or missing hours at daylight-saving changes. Be aware that the daily load pattern follows local time, so in UTC it shifts by one hour between summer and winter.