ephemerisBlog

Blog · Updated

How to give Claude, Cursor or any AI agent a forecasting tool (MCP)

Short answer

Add a remote MCP server to your agent's config and it gains a forecast tool that returns quantile bands instead of guessed numbers. For Ephemeris that is one URL (https://ephemeris.cascade.industries/api/mcp) and one bearer header, and the setup differs only in where each client keeps its config. Keep the key in an environment variable or a password prompt, and check it works by asking the agent to call list_models.

mcp · agents · claude · cursor · forecastingRead as markdown

Key facts

  • The Ephemeris MCP server is at https://ephemeris.cascade.industries/api/mcp, uses the Streamable HTTP transport, and authenticates with the header Authorization: Bearer pc_live_....
  • It exposes four tools: forecast, list_models, get_balance and get_usage. The forecast tool runs the same pipeline as POST /api/v1/forecast.
  • One forecast call takes 1 to 64 series, a horizon of 1 to 512 steps and up to 21 quantile levels strictly between 0 and 1.
  • A model run costs price_per_kslot_mc x slots x ceil(context / 1024) x ceil(horizon / 64) millicredits, where 1 credit = 1,000 millicredits.
  • Credits are prepaid: $10 buys 10,000 credits, $25 buys 30,000 and $100 buys 200,000.
  • Retrying a forecast with the same idempotency key replays the stored result without charging again; reusing a key with a different request body returns HTTP 409.
  • The Claude API reaches remote MCP servers through the MCP connector, which needs the beta header mcp-client-2025-11-20 and currently supports tool calls only.

What is MCP?

The Model Context Protocol (MCP) is an open standard for connecting AI applications to external tools and data. An MCP server publishes tools with a name, a description and an input schema; an MCP client (Claude Code, Cursor, VS Code and many others) lists those tools, shows them to the model, and runs the calls the model makes. A remote MCP server runs on the internet and the client reaches it over HTTP, so there is nothing to install locally. The specification and a list of clients are at modelcontextprotocol.io.

For forecasting this matters because language models are poor at extrapolating numbers (see Why language models are bad at forecasting numbers). A forecasting tool lets the model hand the numbers to a model built for time series and spend its own effort on reading the request and explaining the result.

What do the four forecasting tools do?

The forecast tool produces the forecast; the other three tell the agent what is available and what it has spent.

ToolWhat it doesWhen an agent should call it
forecastForecasts one or more series and returns quantiles, the models used and the credits charged. Accepts an optional idempotency_key.When the user asks for a forecast of numbers they supplied.
list_modelsReturns the live model panel: capabilities (multivariate, covariates, max_horizon), ensemble weights by horizon, weight revisions and prices.Once per session, and before naming a model in explicit mode.
get_balanceReturns spendable credits.Before a large batch.
get_usageReturns recent request history.When the user asks what was spent.

forecast takes the same fields as the REST API. The ones an agent uses most are mode (route lets Ephemeris pick a model, ensemble combines every compatible model, explicit runs the one named in model), series (each with values, oldest first, and an optional freq such as "H" or "D"), horizon (steps to forecast) and quantiles (for example [0.1, 0.5, 0.9]). A quantile forecast at 0.9 is a value the outcome should fall below 90% of the time, so the 0.1 and 0.9 lines together form an 80% band.

How do I set it up in each client?

Every client below needs the same two things: the URL https://ephemeris.cascade.industries/api/mcp and the header Authorization: Bearer <your key>. Create a key in the Ephemeris dashboard after signing up; keys start with pc_live_ and are shown once. The examples read the key from an environment variable called EPHEMERIS_API_KEY wherever the client supports it.

Claude Code

Add the server with one command. The shell expands $EPHEMERIS_API_KEY when you run it:

bash
claude mcp add --transport http ephemeris https://ephemeris.cascade.industries/api/mcp \
  --header "Authorization: Bearer $EPHEMERIS_API_KEY"

The default scope is local: the server is available in the current project and private to you. Add --scope user to make it available in every project. To share the server with a team through a project's .mcp.json without committing the key, write the header with Claude Code's ${VAR} expansion so each person's own environment supplies the key:

json
{
  "mcpServers": {
    "ephemeris": {
      "type": "http",
      "url": "https://ephemeris.cascade.industries/api/mcp",
      "headers": { "Authorization": "Bearer ${EPHEMERIS_API_KEY}" }
    }
  }
}

Check it with claude mcp list in a terminal or /mcp inside a session. Reference: Claude Code MCP docs.

Cursor

Add the server to ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project). Cursor resolves ${env:NAME} in url and headers, so the key can stay in your environment:

json
{
  "mcpServers": {
    "ephemeris": {
      "url": "https://ephemeris.cascade.industries/api/mcp",
      "headers": { "Authorization": "Bearer ${env:EPHEMERIS_API_KEY}" }
    }
  }
}

If you paste the key in place of ${env:EPHEMERIS_API_KEY}, keep that file out of version control. Reference: Cursor MCP docs.

VS Code

Put this in .vscode/mcp.json. The inputs entry makes VS Code prompt for the key once and store it, so the key never lands in the repository:

json
{
  "inputs": [{ "type": "promptString", "id": "ephemeris-key", "description": "Ephemeris API key", "password": true }],
  "servers": {
    "ephemeris": {
      "type": "http",
      "url": "https://ephemeris.cascade.industries/api/mcp",
      "headers": { "Authorization": "Bearer ${input:ephemeris-key}" }
    }
  }
}

Run MCP: List Servers from the Command Palette to start the server and see its status. Reference: VS Code MCP configuration reference.

Claude Desktop

Claude Desktop has two routes, and which one works depends on your account.

The first is a custom connector: Customize > Connectors > Add custom connector, enter the URL, choose No sign-in, and add an authorization header under Request headers with the value Bearer pc_live_... (the word Bearer and a space included, because Claude sends the value exactly as typed). Anthropic's documentation says request-header authentication is in beta and available to a limited set of organizations, so you may not see that section. Reference: Add a connector that isn't in the directory.

The second works everywhere: run the server through the mcp-remote bridge, which turns a remote server into a local one. It needs Node.js. Open Settings > Developer > Edit Config to edit claude_desktop_config.json (~/Library/Application Support/Claude/ on macOS, %APPDATA%\Claude\ on Windows):

json
{
  "mcpServers": {
    "ephemeris": {
      "command": "npx",
      "args": ["-y", "mcp-remote", "https://ephemeris.cascade.industries/api/mcp",
               "--header", "Authorization:${EPHEMERIS_AUTH}"],
      "env": { "EPHEMERIS_AUTH": "Bearer pc_live_your_key" }
    }
  }
}

There is no space after Authorization: on purpose: Claude Desktop on Windows (and Cursor) do not escape spaces inside args, so the space lives in the environment variable instead. Quit and restart Claude Desktop after saving. References: mcp-remote README, Connect to local MCP servers.

Claude API (Messages API MCP connector)

The Messages API can call a remote MCP server for you, so your code does not need an MCP client. You declare the server in mcp_servers, enable its tools with an mcp_toolset entry in tools, and send the beta header mcp-client-2025-11-20. With the anthropic Python SDK:

python
import os

import anthropic

client = anthropic.Anthropic()

response = client.beta.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    betas=["mcp-client-2025-11-20"],
    mcp_servers=[{
        "type": "url",
        "url": "https://ephemeris.cascade.industries/api/mcp",
        "name": "ephemeris",
        "authorization_token": os.environ["EPHEMERIS_API_KEY"],
    }],
    tools=[{"type": "mcp_toolset", "mcp_server_name": "ephemeris"}],
    messages=[{"role": "user", "content": "Forecast the next 12 points of 3,5,4,6,7,6,8,9,8,10"}],
)

for block in response.content:
    if block.type == "text":
        print(block.text)

The tool calls and results come back in the response as mcp_tool_use and mcp_tool_result blocks. Anthropic's servers make the connection, which is why the MCP server must be publicly reachable over HTTP. The connector supports only MCP tools (not prompts or resources), is not eligible for zero data retention, and is not available on Amazon Bedrock or Google Cloud. To expose only some tools, set "default_config": {"enabled": false} on the toolset and enable the ones you want under configs. Reference: MCP connector docs. For LangChain and the OpenAI Agents SDK, see Adding probabilistic forecasting to LangChain, the OpenAI Agents SDK and the Claude API.

OpenClaw

OpenClaw can use Ephemeris through a skill or through MCP.

The skill, ephemeris-forecasting on ClawHub, teaches the agent when to forecast and calls the REST API with curl. Install it with openclaw skills install @tensorlink-dev/ephemeris-forecasting (or find it with openclaw skills search "ephemeris"). The skill declares EPHEMERIS_API_KEY as its primary environment variable, so you can hand it the key in ~/.openclaw/openclaw.json with a SecretRef that reads your environment instead of a pasted value:

json
{
  "skills": {
    "entries": {
      "ephemeris-forecasting": {
        "apiKey": { "source": "env", "provider": "default", "id": "EPHEMERIS_API_KEY" }
      }
    }
  }
}

The MCP route uses OpenClaw's mcp.servers config. OpenClaw substitutes ${VAR_NAME} from the environment in any config string. Set transport explicitly: when it is omitted, OpenClaw uses the older SSE transport.

json
{
  "mcp": {
    "servers": {
      "ephemeris": {
        "url": "https://ephemeris.cascade.industries/api/mcp",
        "transport": "streamable-http",
        "headers": { "Authorization": "Bearer ${EPHEMERIS_API_KEY}" }
      }
    }
  }
}

Check the connection with openclaw mcp doctor ephemeris --probe. In our experience the skill is the simpler path when you only need forecasts, because it carries the usage rules (send only real history, report the band) along with the call. References: OpenClaw MCP, MCP transports, MCP, skills and plugins config, env var substitution, Skills CLI.

Hermes Agent

Hermes Agent reads MCP servers from ~/.hermes/config.yaml under mcp_servers. A remote server takes a url and a headers map, and Hermes resolves ${VAR} placeholders in headers from the environment, including ~/.hermes/.env:

yaml
mcp_servers:
  ephemeris:
    url: "https://ephemeris.cascade.industries/api/mcp"
    headers:
      Authorization: "Bearer ${EPHEMERIS_API_KEY}"

Put EPHEMERIS_API_KEY=pc_live_... in ~/.hermes/.env, then start a new session or run /reload-mcp. hermes mcp test ephemeris reports what the server answered if the connection fails. Hermes registers the tools with a server prefix, so forecast appears as mcp_ephemeris_forecast. Reference: Hermes Agent MCP docs.

Anything else: plain REST

Any agent that can make an HTTP request can forecast without MCP. Clients that only run local (stdio) servers can also use the bridge: npx mcp-remote <url> --header "Authorization: Bearer <key>". Otherwise call the REST endpoint directly:

bash
curl --request POST \
  --url https://ephemeris.cascade.industries/api/v1/forecast \
  --header "Authorization: Bearer $EPHEMERIS_API_KEY" \
  --header "Content-Type: application/json" \
  --header "Idempotency-Key: forecast-2026-10-07-001" \
  --data '{
    "mode": "route",
    "series": [{ "values": [100.2, 101.4, 99.8, 103.1, 104.6, 102.2], "freq": "H" }],
    "horizon": 24,
    "quantiles": [0.1, 0.5, 0.9]
  }'

The full request and response format is in /llms-full.txt, and Adding probabilistic forecasting to LangChain, the OpenAI Agents SDK and the Claude API wraps this call as a function tool.

What should my first prompt be?

Start with a request that names the data, its frequency, the horizon and the band you want. For example: "Here are 48 hourly request counts, oldest first: [paste]. Forecast the next 24 hours with the 10th, 50th and 90th percentiles, and tell me which hours could exceed 1,200."

Give the agent real history. The models need at least a few dozen points, and more is better. Ask it to report the band, not only the median: the distance between the 0.1 and 0.9 lines is the uncertainty, and it is usually the useful part of the answer.

How do I check the connection works?

Ask the agent to "list the available forecasting models". That makes it call list_models, which returns the live panel without running a forecast (credits are only spent by forecasts that run), so a correct answer proves the URL, header and key all work. A 401 means the key is missing, wrong or revoked; check that the header value starts with Bearer followed by the key.

Each client also shows server status: claude mcp list or /mcp in Claude Code, MCP: List Servers in VS Code, openclaw mcp doctor in OpenClaw and hermes mcp test in Hermes Agent.

How do I keep the API key out of my repository?

Keep the key in the environment or a secret store and reference it from config. Every client above supports one way to do this: Claude Code's ${VAR} in .mcp.json, Cursor's ${env:NAME}, VS Code's password inputs, OpenClaw's SecretRef and ${VAR_NAME}, and Hermes Agent's ~/.hermes/.env.

Use a separate key for each agent or integration. Anyone holding a key can spend its account's credits, and per-agent keys let you revoke one without breaking the others.

How do I control what the agent spends?

Every successful forecast spends credits, and the response says how many: meta.billing.settled_mc is the cost of the call and balance_mc is what remains. One series with up to 1,024 points of context and a horizon up to 64 steps costs exactly one model's rate; route mode usually runs one model, and ensemble mode runs every compatible model, so it costs more. Per-model rates are on /pricing.

Four habits keep spending predictable:

  • Tell the agent in its instructions not to forecast speculatively or in a loop you did not ask for.
  • Have it pass an idempotency_key (8 to 128 letters, digits, dot, underscore, colon or hyphen) so a retried call is replayed without a second charge.
  • Have it call get_balance before a large batch. A 402 means the account is out of credits; the agent should tell you rather than retry.
  • Batch related series into one call (up to 64) instead of one call per series.

Do I need a hosted API, or can I run the model myself?

You can run the models yourself. Chronos-2, TimesFM 2.5, Toto 2, TiRex-2 and FlowState are published on Hugging Face under Apache-2.0, and PatchTST-FM r2 under OpenMDW-1.0, so you can load one in Python and expose it to your agent as your own MCP server or function tool. That is the better choice when data must not leave your network, when you forecast at a volume where a GPU you already own is cheaper than per-call credits, or when you want to fine-tune.

A hosted endpoint saves you serving the models, keeping several of them on compatible dependencies, and building routing or ensembling. Comparisons of the models themselves are in Chronos-2 vs TimesFM 2.5 vs Toto 2 vs TiRex-2: one harness, same data, and other hosted options in TimeGPT alternatives in 2026: hosted and open-source forecasting models.

Every setup above needs one Ephemeris API key. Create an account, make a key in the dashboard, and new accounts get a small free credit to try the first forecasts.

FAQ

Do I need to install anything to use a remote MCP server?

No, for clients that speak Streamable HTTP, such as Claude Code, Cursor, VS Code, OpenClaw and Hermes Agent: you only add the URL and header to their config. Clients that only run local servers need a bridge such as mcp-remote, which runs through npx and needs Node.js.

Which Ephemeris mode should an agent use?

Use route when unsure; it usually runs one model and is the cheapest good default. Use ensemble when calibrated uncertainty matters more than cost, and explicit only when the user names a model, after checking it in list_models.

Why does my Claude Desktop config fail on Windows?

Claude Desktop on Windows does not escape spaces inside args when it runs npx, which breaks a header like Authorization: Bearer .... Write Authorization:${EPHEMERIS_AUTH} with no space and put Bearer pc_live_... in the env block.

Can the agent make up data to forecast?

It should not, and the tool cannot tell. Instruct the agent to send only the history the user provided, oldest first, without padding or interpolating missing points.

Is calling list_models charged?

No. Credits are only spent by forecasts that run. Calling list_models once per session is the recommended way to learn current model names, limits and prices.