ephemerisBlog

Blog · Updated

Forecasting sensor and IoT readings: missing data, sampling rates and many series at once

Short answer

Sensor data has gaps, uneven sampling and many channels. Do not fill gaps with invented values; forecast from the latest unbroken stretch of real readings, and skip a sensor whose recent data is too short or too old. Put each series on a regular grid with a `freq` string that matches it, and send many sensors in one request, either as separate series or as one multivariate series per device. This tutorial does all three on the UCI Air Quality dataset.

forecasting · iot · sensors · missing-data · tutorialRead as markdown

Key facts

  • The Ephemeris API takes values as numbers on a regular grid, oldest first, with no timestamps; a gap cannot be marked, so send only the latest stretch with no missing points.
  • Filling gaps by interpolation or imputation makes the model treat guesses as observations, which narrows the forecast band; the API's guidance is to send only the values you have.
  • freq is a pandas-style string per series, such as "5min", "15min", "H", "D", "W" or "M"; horizon is counted in steps of that frequency, up to 512.
  • One request holds up to 64 series and 256 slots; a multivariate series of 8 sensor channels uses 8 slots.
  • The UCI Air Quality dataset (CC BY 4.0) has about 9,350 hourly rows from one gas multisensor device in an Italian city, March 2004 to April 2005, with missing values coded as -200.
  • In that file the device's own channels are all missing for the same 366 hours, while the reference analyser's NMHC(GT) column is missing for 8,443 of 9,357 hours.

Why is sensor data harder to forecast than business data?

Because the collection is part of the signal. Sensors drop out, gateways lose connectivity, batteries die and devices are serviced. A sales ledger rarely has a missing day; a field sensor often has a missing week.

Three problems come with that:

  • Gaps, short (a missed reading) or long (a device offline).
  • Sampling rates that differ by device, drift over time or are irregular.
  • Scale: thousands of sensors, each a series, often several channels per device.

There is also drift: a sensor's response changes as it ages. The UCI page for the dataset used here notes evidence of drift and cross-sensitivity in its sensors. A forecast reproduces the pattern in the readings, drift included; it does not correct the instrument.

Which dataset does this tutorial use?

It uses UCI Air Quality (DOI 10.24432/C59K5F), licensed CC BY 4.0. It holds hourly averaged responses of five metal-oxide sensors on one device, with temperature and humidity, next to reference concentrations from a certified analyser.

The CSV inside the zip uses semicolons as separators and commas as decimal points. The columns are Date (day/month/year), Time (hours.minutes.seconds), the reference values CO(GT), NMHC(GT), C6H6(GT), NOx(GT) and NO2(GT), the sensor responses PT08.S1(CO) to PT08.S5(O3), and T, RH and AH (temperature, relative humidity and absolute humidity). Missing values are written as -200.

python
import io
import os
import zipfile

import numpy as np
import pandas as pd
import requests

URL = "https://archive.ics.uci.edu/static/public/360/air+quality.zip"
z = zipfile.ZipFile(io.BytesIO(requests.get(URL, timeout=60).content))

raw = pd.read_csv(z.open("AirQualityUCI.csv"), sep=";", decimal=",")
raw = raw.dropna(axis=1, how="all").dropna(subset=["Date"])   # empty trailing columns and rows
raw.index = pd.to_datetime(raw["Date"] + " " + raw["Time"], format="%d/%m/%Y %H.%M.%S")

data = (
    raw.drop(columns=["Date", "Time"])
    .replace(-200, np.nan)      # the file's missing-value code
    .asfreq("h")                # strict hourly grid: any missing hour becomes NaN
)
print(data.isna().sum())

The zip also contains an Excel copy, so open the CSV by name rather than letting pandas guess.

How should I handle gaps in sensor data?

Forecast from the latest unbroken stretch of real readings, and say so. Everything before the last gap is dropped; nothing is invented. This is the approach the rest of the tutorial uses.

python
def latest_block(df):
    """Rows after the last row with any missing value, so the block has no gaps."""
    ok = df.notna().all(axis=1)
    last_ok = ok[ok].index[-1]
    df, ok = df.loc[:last_ok], ok.loc[:last_ok]
    bad = ok[~ok]
    return df if bad.empty else df.loc[bad.index[-1]:].iloc[1:]

Two checks keep this honest:

  • Is the stretch long enough? A forecaster needs at least a few dozen points, and enough to show the cycle you care about; for hourly data, 3 days (72 points) is a sensible floor.
  • Is it recent? If a sensor's last reading was hours or days ago, its forecast starts then, not now, and part of the horizon is already in the past. Skip it, or report the origin with the forecast.

Here is how the other options compare:

OptionWhat it doesWhen it is reasonable
Latest unbroken stretchUses only real readings after the last gapDefault; costs older history
Aggregate to a coarser rateAverages, say, 5-minute readings into hourly means, keeping only complete hoursShort, frequent dropouts; the question tolerates the coarser rate
Interpolate or imputeFills gaps with estimatesNot as forecast input: the model treats the estimates as data and its band comes out too narrow
Skip the sensorReports "not enough recent data"Stretch too short or too old

Aggregation is not invention, but it changes the question: an hourly mean forecast says nothing about a 5-minute spike.

How do I forecast many sensors in one request?

Put each sensor in its own entry of series, up to 64 series and 256 slots per request. The response returns forecasts in the same order, so keep a list of names alongside.

python
API = "https://ephemeris.cascade.industries/api/v1/forecast"
KEY = os.environ["EPHEMERIS_API_KEY"]

HORIZON = 24                      # hours
CONTEXT = 336                     # up to 2 weeks of history
MIN_POINTS = 72                   # need at least 3 days since the last gap
NOW = data.index[-1]              # in production: the current time, floored to the hour
MAX_AGE = pd.Timedelta(hours=2)   # skip sensors whose last reading is older than this


def forecast(body, idempotency_key):
    r = requests.post(
        API,
        headers={
            "Authorization": f"Bearer {KEY}",
            "Content-Type": "application/json",
            "Idempotency-Key": idempotency_key,
        },
        json=body,
        timeout=300,
    )
    r.raise_for_status()
    return r.json()


series, names = [], []
for col in data.columns:
    if data[col].notna().sum() == 0:
        continue
    block = latest_block(data[[col]])[col]
    if NOW - block.index[-1] > MAX_AGE:
        print(f"skip {col}: last reading at {block.index[-1]}")
        continue
    if len(block) < MIN_POINTS:
        print(f"skip {col}: only {len(block)} hours since the last gap")
        continue
    series.append({"values": block.tail(CONTEXT).tolist(), "freq": "H"})
    names.append(col)

out = forecast(
    {"mode": "route", "series": series, "horizon": HORIZON,
     "quantiles": [0.05, 0.5, 0.95], "context_len": CONTEXT},
    f"airquality-univariate-{NOW:%Y%m%dT%H}",
)
for name, fc in zip(names, out["forecasts"]):
    print(name, "next hour 90% band:", fc["quantiles"]["0.05"][0], "to", fc["quantiles"]["0.95"][0])

On this file the two checks do real work. NMHC(GT) is skipped because its last reading is on 1 May 2004, eleven months before the data ends. CO(GT) is skipped because its last gap ends at 04:00 on the final day, which leaves only 10 hours of history. The remaining channels go out in one request.

When should sensors be one multivariate series?

When they are channels of the same device or process and share a clock. Then values is a list of equal-length lists, one per channel, and a multivariate model can use how the channels move together: here, how the gas sensors respond to temperature and humidity.

python
device = ["PT08.S1(CO)", "PT08.S2(NMHC)", "PT08.S3(NOx)", "PT08.S4(NO2)",
          "PT08.S5(O3)", "T", "RH", "AH"]
block = latest_block(data[device]).tail(CONTEXT)

out = forecast(
    {"mode": "route",
     "series": [{"values": [block[c].tolist() for c in device], "freq": "H"}],
     "horizon": HORIZON, "quantiles": [0.05, 0.5, 0.95], "context_len": CONTEXT},
    f"airquality-device-{block.index[-1]:%Y%m%dT%H}",
)
medians = dict(zip(device, out["forecasts"][0]["quantiles"]["0.5"]))   # one list per channel

This request uses 8 slots, one per channel. In this file the device's channels drop out together (the same 366 hours), so the joint stretch is as long as each channel's own: it starts after the last outage on 11 March 2005, a little over three weeks before the data ends.

Devices from different sites should stay separate series, even if they measure the same thing. In the Ephemeris panel, Chronos-2, Toto 2, TiRex-2 and PatchTST-FM r2 accept multivariate series, and route mode sends them to Toto 2.

Which freq string should I use?

The one that matches the grid you put the data on, written the pandas way. The table shows what the 512-step horizon limit means at common rates:

SamplingfreqSteps per day512 steps covers
Every 5 minutes5min288about 42.7 hours
Every 15 minutes15min96about 5.3 days
HourlyH24about 21.3 days
DailyD1512 days

A few practical rules:

  • Resample first. If readings arrive irregularly, put them on a grid with resample(...).mean() (or .last() for states and counters), and treat empty bins as gaps.
  • Choose the coarsest rate that answers the question. Hourly forecasts of a slow temperature are easier and cheaper than 1-minute forecasts.
  • Group requests by rate. freq is set per series, but horizon is set per request, so 24 steps means 2 hours for a 5-minute sensor and a day for an hourly one. Send each sampling rate in its own request.
  • Spelling. The API's examples use H for hourly. Recent pandas versions spell hourly h in their own functions, as in asfreq("h") above.

What are the forecast quantiles useful for with sensors?

Three things, mostly:

  • Threshold alerts. Warn when the 0.95 quantile of a reading crosses a safety or comfort limit within the horizon, the same way ops teams alert on CPU.
  • Anomaly flags. When a new reading lands outside a wide band, such as 0.01 to 0.99, flag it for a look.
  • Drift and fault checks. If a sensor's readings keep landing outside its own forecast band day after day, the sensor may need recalibration or repair.

Check coverage on your own history before relying on any of these: about 90% of readings should fall inside a 0.05 to 0.95 band. We have not run this on the Air Quality data, so we do not quote a coverage number.

What are the alternatives?

For sensor fleets, the common options are:

  • Classical per-series models (ETS, ARIMA, seasonal naive) in Nixtla's statsforecast, fast enough for many thousands of series on one machine.
  • Your time-series database's built-in functions, for simple extrapolation and smoothing where the data already lives.
  • Self-hosted pretrained models such as Chronos-2 or TiRex-2 from Hugging Face, which suits edge or on-premises deployments where data cannot leave the site.
  • A hosted API like Ephemeris, when you want probabilistic forecasts for many sensors without running models.

If the data must stay on the device or the network is unreliable, self-hosting or a small classical model is the better choice; a remote API needs connectivity at forecast time.

The batching code above works as written against Ephemeris. Create an account to try it on your own sensors; new accounts get a small free credit.

FAQ

Should I interpolate missing sensor readings before forecasting?

No, not as forecast input. The model treats interpolated values as real observations and its uncertainty band comes out too narrow. Forecast from the latest stretch with no gaps, or aggregate to a coarser rate where bins are complete.

What if a sensor has a gap right before the forecast time?

Then its latest unbroken stretch is short, and the forecast should either use that short stretch, if it still holds a few dozen points, or be skipped. Report the forecast origin so nobody reads a stale forecast as current.

Can I send sensors with different sampling rates in one request?

Each series has its own freq, but horizon applies to the whole request in steps, so the same horizon means different durations. Group sensors by sampling rate and send one request per rate.

How many sensors can I forecast in one request?

Up to 64 series and 256 slots, where a slot is one variate of one series. Sixty-four single-channel sensors fit, or 32 devices with 8 channels each.

Which frequency string should I use for hourly data?

Use "H" in the request, as in the API examples. The horizon is then counted in hours, up to 512.