Blog · Updated
Capacity planning from ops metrics: forecasting traffic, CPU and queue depth
Short answer
For capacity planning, the question is not what CPU will typically be tomorrow but how likely it is to cross the limit. Forecast the metric with quantiles and alert when a high quantile, such as the 0.95, crosses the limit inside the horizon; that gives you lead time and a stated risk instead of a page at the peak. This tutorial does it on 5-minute AWS CloudWatch metrics from the Numenta Anomaly Benchmark, as single metrics and as one multivariate series.
Key facts
- Alert on a high quantile, not the median: if the 0.95 quantile of CPU crosses 80% within the horizon, there is roughly a 5% or greater chance of crossing at that step, if the forecast is calibrated.
- At 5-minute resolution a day is 288 steps and a week is 2,016 steps; the Ephemeris API allows horizons up to 512 steps (about 42 hours at 5 minutes).
context_lendefaults to 256 points, under 22 hours of 5-minute data, so set it to cover at least a few daily cycles.- Toto 2 (Datadog, 313M parameters, Apache-2.0) was trained heavily on observability metrics, and it is the model
routemode picks for multivariate series. - A multivariate series sends several aligned metrics together as one series; each metric is one slot, and a request holds at most 64 series and 256 slots.
- The Numenta Anomaly Benchmark (NAB, MIT licence) includes 17 AWS CloudWatch metric files, such as EC2 CPU utilisation and ELB request count, each with
timestamp,valuecolumns at 5-minute intervals.
What does capacity planning need from a forecast?
A probability of running out, with enough lead time to act. Autoscaling reacts within minutes; buying hardware, raising a quota or resharding a database takes hours to weeks. A forecast tells you which of those you need to start now.
The median forecast is the wrong number for this. Actual load lands above the median about half the time, so a plan built on the median runs short on every busy day. Capacity decisions use the upper quantiles: the 0.9 or 0.95 quantile is the level the metric should stay under 90% or 95% of the time.
Ops metrics have a few traits that shape the forecast:
- Daily and weekly cycles driven by users, plus batch jobs at fixed times.
- Hard bounds: CPU utilisation is between 0 and 100%, queue depth is never negative.
- Spikes and incidents in the history, which widen the bands the model learns from.
- Coupled metrics: request rate, CPU and latency move together on the same service.
Which data does this tutorial use?
It uses the realAWSCloudwatch folder of the Numenta Anomaly Benchmark (NAB), released under the MIT licence. NAB was built to test anomaly detectors, and its files contain labelled anomalies; that is a realistic feature of ops data, not a problem.
Each CSV has two columns, timestamp and value, at 5-minute intervals. The files used here are ec2_cpu_utilization_825cc2.csv (EC2 CPU, percent), elb_request_count_8c0756.csv (load balancer request count) and ec2_network_in_257a54.csv (network bytes in). Each has 4,032 rows, which is 14 days, and all three start on 2014-04-10. A few intervals are 10 minutes rather than 5, so the code puts everything on a strict 5-minute grid and checks for gaps.
NAB does not say these three metrics come from the same service. We group them to show the multivariate request shape; on your own systems, group metrics that belong to one service.
If you use NAB in published work, cite Ahmad, Lavin, Purdy and Agha, "Unsupervised real-time anomaly detection for streaming data", Neurocomputing, 2017.
How do I forecast CPU and alert on a quantile?
Forecast the next 24 hours with the 0.5, 0.9 and 0.95 quantiles, clip to the metric's bounds, and find the first step where the 0.95 quantile crosses the limit.
import os
import numpy as np
import pandas as pd
import requests
API = "https://ephemeris.cascade.industries/api/v1/forecast"
KEY = os.environ["EPHEMERIS_API_KEY"]
NAB = "https://raw.githubusercontent.com/numenta/NAB/master/data/realAWSCloudwatch/"
def load_metric(name):
s = pd.read_csv(NAB + name + ".csv", parse_dates=["timestamp"], index_col="timestamp")["value"]
return s.resample("5min").mean() # strict 5-minute grid; empty bins are NaN
def latest_block(df):
"""Rows after the last row with any missing value, so the block has no gaps."""
ok = df.notna().all(axis=1)
df, ok = df.loc[:ok[ok].index[-1]], ok.loc[:ok[ok].index[-1]]
bad = ok[~ok]
return df if bad.empty else df.loc[bad.index[-1]:].iloc[1:]
def forecast(body, idempotency_key):
r = requests.post(
API,
headers={
"Authorization": f"Bearer {KEY}",
"Content-Type": "application/json",
"Idempotency-Key": idempotency_key,
},
json=body,
timeout=300,
)
r.raise_for_status()
return r.json()
HORIZON = 288 # 24 hours of 5-minute steps
CONTEXT = 2016 # 7 days of history
LIMIT = 80.0 # CPU percent you do not want to exceed
QUANTILES = [0.5, 0.9, 0.95]
cpu = latest_block(load_metric("ec2_cpu_utilization_825cc2").to_frame())["value"]
out = forecast(
{
"mode": "route",
"series": [{"values": cpu.tail(CONTEXT).tolist(), "freq": "5min"}],
"horizon": HORIZON,
"quantiles": QUANTILES,
"context_len": CONTEXT,
},
"nab-cpu-825cc2-24h",
)
q = {k: np.clip(np.array(v), 0, 100) for k, v in out["forecasts"][0]["quantiles"].items()}
future_index = pd.date_range(cpu.index[-1], periods=HORIZON + 1, freq="5min")[1:]
over = np.flatnonzero(q["0.95"] > LIMIT)
if over.size:
print(f"0.95 quantile crosses {LIMIT}% at {future_index[over[0]]} "
f"({(over[0] + 1) * 5} minutes ahead); {over.size} of {HORIZON} steps above")
else:
print(f"0.95 quantile stays under {LIMIT}% for the next 24 hours")
print(out["meta"]["models_used"], out["meta"]["billing"]["settled_mc"], "mc")The forecast starts from the last real timestamp in the data, so future_index is built from it. If your data ends earlier than now, part of the horizon is already in the past.
Should the alert use the 0.9 or the 0.95 quantile?
Use the quantile that matches the cost of being caught short. The 0.95 quantile says "about a 1 in 20 chance of being above this at this step"; the 0.9 says about 1 in 10. A database that falls over when full deserves 0.95 or 0.99; a web tier that autoscales in minutes can plan on 0.9.
Three refinements make quantile alerts less noisy:
- Require persistence. Alert when the quantile stays above the limit for several consecutive steps (say 3, which is 15 minutes), not on a single step.
- Alert on lead time. "Crosses in 6 hours" and "crosses in 20 minutes" need different responses; route them differently.
- Check calibration. On your own history, count how often the actual value exceeds the forecast 0.95 quantile. It should be about 5% of steps. If it is 20%, the alert is understating risk.
How do I forecast several metrics of one service together?
Send them as one multivariate series: values is a list of equal-length lists, one per metric. A model that reads them together can use the link between request rate and CPU; the response returns one inner list per metric, in the same order.
names = ["ec2_cpu_utilization_825cc2", "elb_request_count_8c0756", "ec2_network_in_257a54"]
frame = pd.concat({n: load_metric(n) for n in names}, axis=1)
block = latest_block(frame).tail(CONTEXT)
out = forecast(
{
"mode": "route",
"series": [{"values": [block[n].tolist() for n in names], "freq": "5min"}],
"horizon": HORIZON,
"quantiles": QUANTILES,
"context_len": CONTEXT,
},
"nab-service-multivariate-24h",
)
p95 = out["forecasts"][0]["quantiles"]["0.95"] # one list per metric
cpu_p95, requests_p95, net_p95 = (np.array(v) for v in p95)
print(out["meta"]["models_used"])latest_block matters more here: a multivariate series needs every metric present at every step, so the usable history is the stretch where all of them are. In these files the request-count metric misses a scrape on 2014-04-20, so the joint block starts after that and holds under 4 days, against about 10 days for CPU alone. That loss is the honest cost of not inventing data; if it happens often on your systems, fix the collection rather than the forecast.
The request uses 3 slots, one per metric, and context_len applies per variate.
In route mode, Ephemeris picks Toto 2 for multivariate series. Datadog trained it heavily on observability metrics, which is the closest match in the panel to this kind of data. Other multivariate-capable models in the panel are Chronos-2, TiRex-2 and PatchTST-FM r2.
To forecast many hosts at once, put each host's metric (or each host's multivariate block) in its own entry of series, up to 64 series and 256 slots per request.
What about queue depth?
Queue depth is harder than CPU because it is mostly zero and then grows quickly when arrivals outpace processing. A forecast of the queue itself has wide, skewed bands and needs clipping at zero.
Often a better plan is to forecast the drivers, arrival rate and processing rate, and reason about the queue from them. If the 0.95 quantile of arrivals exceeds your measured processing capacity for a sustained period, the queue will grow; Little's law (average number of items in the system = arrival rate × average time in the system) turns a target wait time into a capacity requirement. NAB does not include a queue-depth series, so this section has no code.
How often should I re-forecast?
As often as the decision changes, not on every scrape. Re-forecasting a 24-hour horizon every 15 to 60 minutes is enough for most capacity alerts, and each call spends credits. Cost per model run scales with ceil(context / 1024) and ceil(horizon / 64), so a 2,016-point context and a 288-step horizon cost 2 × 5 = 10 times a minimal request; see /pricing for the rates.
If you call the API from a scheduler, send an Idempotency-Key built from the metric and the forecast time, so a retried job is not charged twice.
What are the alternatives?
Most monitoring stacks already have simple forecasting, and for some jobs it is enough:
- Prometheus
predict_linearfits a straight line to a range and extrapolates. It is free and fine for slowly filling disks; it has no daily cycle and no uncertainty band. - Built-in forecasts in monitoring products, such as Datadog's
forecast()function in monitors, are convenient where your metrics already live. - Classical models like Holt-Winters or MSTL (in Nixtla's statsforecast) handle daily and weekly cycles and give intervals.
- Self-hosting Toto 2 from Hugging Face (Datadog/Toto-2.0-313m, Apache-2.0) if you want the same model next to your metrics store.
- A hosted API like Ephemeris, when you want quantile forecasts across many hosts and metrics without running models yourself.
For a disk that fills at a steady rate, a linear fit is the right tool. Quantile forecasts earn their cost when load is cyclical and the risk at the peak is what you are planning for.
To run the alert above without serving a model yourself, point it at Ephemeris: create an account (a small free credit comes with it) and set EPHEMERIS_API_KEY.
FAQ
Why alert on the 0.95 quantile instead of the forecast?
The median forecast is exceeded about half the time, so it hides the risk at peaks. The 0.95 quantile is the level the metric should stay under about 95% of the time, so it crossing the limit means a real chance of exceeding it.
How much history should I send for 5-minute metrics?
Enough for several daily cycles: at least 3 days (864 points), and a week (2,016 points) if you want weekly patterns. Set context_len explicitly, because the default of 256 points is under 22 hours.
Can I forecast CPU and request rate together?
Yes. Send them as one multivariate series, with values as a list of equal-length lists, one per metric. Each metric counts as one slot, and route mode picks a multivariate-capable model.
Which model is suited to observability metrics?
In the Ephemeris panel, Toto 2 from Datadog was trained heavily on observability metrics and is the route-mode choice for multivariate series. Check its fit on your own metrics with a calibration backtest.
What should I do about gaps in the metrics?
Do not fill them with invented values. Put the data on a strict time grid, find the latest stretch with no missing points, and forecast from that stretch.