# Ephemeris: full reference for LLMs and agents > Ephemeris is a time-series forecasting API. Send numeric history, get probabilistic (quantile) forecasts back from a panel of open-weights zero-shot foundation models. Use it over REST or as a remote MCP server. Base URL: https://ephemeris.cascade.industries ## What it does Ephemeris runs a panel of pretrained time-series foundation models behind one endpoint: - Chronos-2 (Amazon) - TimesFM 2.5 (Google) - Toto 2 (Datadog) - TiRex-2 (NX-AI) - IBM Granite PatchTST-FM and FlowState The exact panel changes over time; `GET /api/v1/models` (or the MCP tool `list_models`) is the source of truth for model names, health, capabilities, horizon limits and prices. Per-model pages (size, licence, capabilities, how to call each by name): https://ephemeris.cascade.industries/models Nothing is trained on your data. You send history, you get forecasts back at the quantile levels you ask for (for example the 10th, 50th and 90th percentiles), so every forecast carries its own uncertainty band. ## When to use it, and when not to Use Ephemeris when: - the task needs a forecast of a numeric series: demand, sales, web traffic, energy load, prices, sensor readings, operational metrics; - uncertainty matters (planning ranges, risk, capacity), not just a single line; - the user can supply the history (at least a few dozen points; more is better). Do not use it: - to fill in or invent missing history: send only the values you have; - for non-numeric or categorical prediction, classification, or text; - speculatively or in a loop without the user's intent: every successful forecast spends credits. ## Authentication Every endpoint takes an API key as a bearer token: Authorization: Bearer pc_live_... Create keys in the dashboard after signing up at https://ephemeris.cascade.industries/sign-up. Keys are shown once. Use a separate key per agent or integration and revoke keys you no longer use: anyone holding a key can spend its account's credits. ## Endpoints - POST /api/v1/forecast: forecast one or more series - GET /api/v1/models: live panel, capabilities, horizon limits, ensemble weights, prices - GET /api/v1/balance: spendable credits (excludes credits reserved by in-flight forecasts) - GET /api/v1/usage: request history OpenAPI: https://ephemeris.cascade.industries/openapi-m1.json ## Forecast request curl --request POST \ --url https://ephemeris.cascade.industries/api/v1/forecast \ --header "Authorization: Bearer pc_live_your_key" \ --header "Content-Type: application/json" \ --header "Idempotency-Key: forecast-2026-10-02-001" \ --data '{ "mode": "route", "series": [{ "values": [100.2, 101.4, 99.8, 103.1, 104.6, 102.2], "freq": "H" }], "horizon": 24, "quantiles": [0.1, 0.5, 0.9] }' Fields: - `mode` (required): `route`, `ensemble` or `explicit` - `route`: Ephemeris picks the best-suited model for the data. The cheapest good default when you are unsure. - `ensemble`: runs every compatible model and combines them into one forecast, weighting the more accurate models more. Best calibrated, highest cost. - `explicit`: runs the single model named in `model`. - `model`: required in explicit mode; must be a name from `/api/v1/models`. - `series` (required): 1 to 64 entries, each with: - `values`: oldest first. A flat array of numbers is one univariate series. An array of equal-length arrays is one multivariate series (one inner array per variate). - `freq` (optional, recommended): pandas-style frequency such as "5min", "H", "D", "W", "M". Improves routing and seasonal handling. - `covariates` (optional): see Covariates below. - `horizon`: steps to forecast, 1 to 512 (default 64). - `quantiles`: up to 21 levels strictly between 0 and 1 (for example [0.1, 0.5, 0.9]). - `context_len`: most recent points per variate to use and bill for, 1 to 16384 (default 256). Longer inputs are truncated to a model's context cap. - `combine` (ensemble only): `mixture` (default; pools the models' predictive distributions) or `vincentize` (averages quantiles). Prefer `mixture`. - `top_k` (ensemble only): cap on how many models participate, 1 to 16. Limits: at most 64 series and 256 slots per request, where a slot is one variate of one series. ## Forecast response { "forecasts": [ { "quantiles": { "0.1": [103.8, 104.1, 104.5], "0.5": [105.2, 105.7, 106.1], "0.9": [106.9, 107.4, 108.0] } } ], "meta": { "gateway_request_id": "req_01J...", "request_id": "model_01J...", "models_used": ["chronos2"], "model_revisions": { "chronos2": "3f9c1e0a..." }, "billing": { "settled_mc": "12", "balance_mc": "4988" } } } - `forecasts` has one entry per input series, in order. Each maps a quantile level (as a decimal string) to `horizon` values; a multivariate series returns one inner array per variate. - `meta.models_used` lists the models that contributed; `meta.notes` explains any model that was skipped or dropped. - `meta.billing.settled_mc` is what this call cost; `balance_mc` is what remains. ## Covariates Each series may carry covariates: "covariates": { "past": { "temperature": [ ...same length as values... ] }, "future": { "temperature": [ ...exactly horizon values... ] } } - `past` holds observed history, aligned with `values`. `future` holds values known in advance over the forecast horizon (a planned price, a holiday flag, a weather forecast). Every `future` key must also appear in `past`. Up to 16 covariates per series. - Only models with `covariates: true` in `/api/v1/models` use them. In `route` and `ensemble` mode the panel narrows to those models. In `explicit` mode, naming a model without covariate support is an error. ## Horizon limits Horizons go up to 512 steps. `/api/v1/models` reports two limits per model: - `max_horizon`: the longest horizon the model can forecast at all (null means no limit). Past it the model is skipped in every mode, and naming it in `explicit` mode is an error. - `auto_max_horizon`: the longest horizon at which `route` and `ensemble` use the model (null means no limit beyond `max_horizon`). Past it the model sits out of automatic selection because the panel is more accurate without it there, but you can still name it in `explicit` mode. TiRex-2 has `auto_max_horizon` 320: its head emits 320 steps per pass and longer horizons are rolled out the way TiRex does it. When a model is skipped for horizon, `meta.notes` says so. ## Accuracy Scored with each benchmark's own harness (our runs, not leaderboard submissions), as ratios to seasonal naive, lower is better: - TIME: ensemble MASE 0.639 / CRPS 0.538, against 0.638 / 0.536 for the leading model (QiYao-M) and 0.640 / 0.536 for TimesFM-3; the best average MASE rank of the 31 models. - GIFT-Eval (97 configurations): ensemble CRPS 0.4662 / MASE 0.6841. Only TimesFM-3 (0.4557), which has a non-commercial licence, scores better; Toto-2 2.5B 0.4759, TiRex-2 0.4781, Chronos-2 0.4854. - The ensemble beats its best single member on GIFT-Eval (+3.6% CRPS for the best member), TIME (+1.6%) and fev-bench (+1.1%), and matches it on BOOM. Full tables: https://ephemeris.cascade.industries/benchmarks ## Retries and idempotency Send an `Idempotency-Key` header (8 to 128 letters, digits, dot, underscore, colon or hyphen). Retrying an identical request with the same key replays the stored result without charging again; reusing a key with a different body is a 409. ## Pricing Prepaid credits; 1 credit = 1,000 millicredits ("mc"). A forecast's cost is, per model run: price_per_kslot_mc x slots x ceil(context / 1024) x ceil(horizon / 64) where the price per model comes from `/api/v1/models`. You pay for the models that actually ran: one in explicit mode, usually one in route mode (more if it falls back to a small ensemble), every compatible model in ensemble mode. Credits are reserved when a forecast starts and settled from what ran. Top-ups: $10 buys 10,000 credits, $25 buys 30,000, $100 buys 200,000. Per-model rates and dollar costs: https://ephemeris.cascade.industries/pricing ## Errors - 400: invalid request (the message says which field and why) - 401: missing, invalid or revoked API key - 402: insufficient credits; top up using the returned topup_url, then retry - 409: idempotency key in progress, or reused with a different body - 429: rate limited; retry after the number of seconds in Retry-After - 503: forecasting temporarily unavailable; retry with backoff ## MCP server (for AI agents) Ephemeris is also a remote MCP server (Streamable HTTP): URL: https://ephemeris.cascade.industries/api/mcp Header: Authorization: Bearer pc_live_your_key Tools: - `forecast`: the same pipeline as POST /api/v1/forecast; accepts an optional `idempotency_key`; returns the forecast plus models used and credits charged. - `list_models`: the live panel with capabilities (multivariate, covariates, max_horizon), ensemble weights by horizon, served weight revisions and prices. Call it before using explicit mode. - `get_balance`: spendable credits. - `get_usage`: recent request history. Claude Code: claude mcp add --transport http ephemeris https://ephemeris.cascade.industries/api/mcp \ --header "Authorization: Bearer pc_live_your_key" Cursor (~/.cursor/mcp.json): { "mcpServers": { "ephemeris": { "url": "https://ephemeris.cascade.industries/api/mcp", "headers": { "Authorization": "Bearer pc_live_your_key" } } } } VS Code (.vscode/mcp.json, prompts for the key so it never lands in the repository): { "inputs": [{ "type": "promptString", "id": "ephemeris-key", "description": "Ephemeris API key", "password": true }], "servers": { "ephemeris": { "type": "http", "url": "https://ephemeris.cascade.industries/api/mcp", "headers": { "Authorization": "Bearer ${input:ephemeris-key}" } } } } Claude Desktop (claude_desktop_config.json, through the mcp-remote bridge): { "mcpServers": { "ephemeris": { "command": "npx", "args": ["-y", "mcp-remote", "https://ephemeris.cascade.industries/api/mcp", "--header", "Authorization:${EPHEMERIS_AUTH}"], "env": { "EPHEMERIS_AUTH": "Bearer pc_live_your_key" } } } } Claude API (Messages API MCP connector, beta header `mcp-client-2025-11-20`): { "model": "claude-opus-5-5", "max_tokens": 16000, "mcp_servers": [{ "type": "url", "url": "https://ephemeris.cascade.industries/api/mcp", "name": "ephemeris", "authorization_token": "pc_live_your_key" }], "tools": [{ "type": "mcp_toolset", "mcp_server_name": "ephemeris" }], "messages": [{ "role": "user", "content": "Forecast the next 12 points of 3,5,4,6,7,6,8,9,8,10" }] } Clients that only speak stdio can use the same bridge: `npx mcp-remote --header "Authorization: Bearer "`. ## Guidance for agents 1. Call `list_models` once per session; model names and limits come from the live panel, not from memory. 2. Prefer `route` when unsure, `ensemble` when calibrated uncertainty matters more than cost, `explicit` only when the user names a model. 3. In explicit mode, check the model supports the request: its `covariates` flag if you send covariates, its `max_horizon` against your horizon, and `multivariate` for multivariate series. 4. Send only the history the user gave you. Pass `freq` when you know it. 5. Report the quantile band, not only the median: the spread is the uncertainty. 6. Use an idempotency key on retries; never loop forecasts without the user's intent.