Blog · Updated · 8 min read
How to use TimesFM or Chronos-2 without hosting it
Short answer
You can self-host TimesFM 2.5 or Chronos-2 from their open-source Python packages, run TimesFM in BigQuery with a SQL function, deploy Chronos-2 into your own AWS account, or call a hosted forecasting API. Self-hosting gives full control; BigQuery suits data already stored there; an API needs no infrastructure at all.
On this page
- Key facts
- What are TimesFM and Chronos-2?
- What are my options for running them?
- How do I self-host TimesFM or Chronos-2?
- Do I need a GPU to run TimesFM?
- What does self-hosting cost me in practice?
- How do I use TimesFM on Google Cloud?
- Is there a managed option for Chronos-2?
- How do I call TimesFM or Chronos-2 through a hosted API?
- Which route should I choose?
- FAQ
- Related
Key facts
- TimesFM 2.5 is a 200M-parameter model from Google Research with up to 16k points of context; its weights are Apache-2.0, per the TimesFM README.
- TimesFM 3.0 weights are non-commercial; Google's README says commercial and production use of 3.0 is permitted through Google Cloud services such as BigQuery ML.
- BigQuery's
AI.FORECASTfunction runs TimesFM 2.5 by default, with horizons from 1 to 10,000 points and at most 15,360 points of context. - Chronos-2 is a 120M-parameter Apache-2.0 model from Amazon; its README recommends AutoGluon-Cloud or SageMaker JumpStart for production on AWS.
- TimesFM's own agent skill lists about 1.5 GB of RAM on CPU or about 1 GB of GPU memory for TimesFM 2.5.
- Ephemeris serves TimesFM 2.5 (
timesfm25) and Chronos-2 (chronos2) by name over REST and MCP, with context caps of 16,384 and 8,192 points.
What are TimesFM and Chronos-2?
They are time-series foundation models (TSFMs): neural networks pretrained on large collections of time series, so they can forecast a new series zero-shot, with no training on your data. You pass in history and get a forecast back.
Both return quantiles, not just one line. A quantile is a level the future value should fall below with a given probability; the 0.1 and 0.9 quantiles together give an 80% band.
TimesFM comes from Google Research. The README lists TimesFM 2.5 at 200M parameters with up to 16k points of context (the number of past points the model reads). Chronos-2 comes from Amazon and handles univariate, multivariate and covariate-informed forecasting, per the Chronos repository. A covariate is an extra input series that helps explain the target, such as a price or a holiday flag.
What are my options for running them?
There are three realistic routes: run the open weights yourself, use the cloud service that the model's publisher points to, or call a hosted API that runs the model for you.
| Route | Models | You manage | Licence notes | Interface |
|---|---|---|---|---|
| Self-host from the official package | TimesFM 2.5, TimesFM 3.0, Chronos-2 | Python environment, hardware, batching, upgrades | TimesFM 2.5 and Chronos-2 Apache-2.0; TimesFM 3.0 weights non-commercial | Python |
BigQuery AI.FORECAST | TimesFM 2.5 (GA), TimesFM 3.0 (Preview) | A Google Cloud project and your data in BigQuery | Commercial use of 3.0 permitted through Google Cloud, per the README | SQL |
| AWS (AutoGluon-Cloud or SageMaker JumpStart) | Chronos-2 | An endpoint in your own AWS account | Apache-2.0 | Python, endpoint |
| Hosted API such as Ephemeris | TimesFM 2.5, Chronos-2 and four other models | An API key and prepaid credits | Serves Apache-2.0 or OpenMDW-1.0 models; not TimesFM 3.0 | REST, MCP |
How do I self-host TimesFM or Chronos-2?
Install the official Python package, download the weights from Hugging Face on first load, and call the forecast function. Both packages are open source under Apache-2.0.
For TimesFM, the README installs it with pip install timesfm[torch], or pip install timesfm[mlx] for Apple silicon. This minimal TimesFM 2.5 example is from the official TimesFM agent skill in the same repository:
import numpy as np
import timesfm
model = timesfm.TimesFM_2p5_200M_torch.from_pretrained("google/timesfm-2.5-200m-pytorch")
model.compile(timesfm.ForecastConfig(
max_context=1024, max_horizon=256, normalize_inputs=True,
use_continuous_quantile_head=True, force_flip_invariance=True,
infer_is_positive=True, fix_quantile_crossing=True,
))
point, quantiles = model.forecast(horizon=24, inputs=[np.sin(np.linspace(0, 20, 200))])
print(point.shape) # (1, 24): the median forecastFor Chronos-2, the Chronos README installs it with pip install chronos-forecasting and works on pandas DataFrames with an ID column, a timestamp column and a target column:
import pandas as pd
from chronos import Chronos2Pipeline
pipeline = Chronos2Pipeline.from_pretrained("amazon/chronos-2", device_map="cuda")
context_df = pd.DataFrame({
"id": "store_1",
"timestamp": pd.date_range("2026-01-01", periods=120, freq="D"),
"target": [float(100 + i % 7) for i in range(120)],
})
pred_df = pipeline.predict_df(
context_df,
prediction_length=14,
quantile_levels=[0.1, 0.5, 0.9],
id_column="id",
timestamp_column="timestamp",
target="target",
)The README example passes future_df as well when you have future covariates, and uses device_map="cuda" for a GPU.
Do I need a GPU to run TimesFM?
Not for small jobs. The official TimesFM agent skill lists TimesFM 2.5 at about 800 MB on disk, about 1.5 GB of RAM on CPU or about 1 GB of GPU memory, and marks CPU mode as "slower but works". It asks for at least 4 GB of system RAM.
Memory grows with the number of series and the context length, so batch large jobs. The same skill suggests a per_core_batch_size of 8 on a CPU with 8 GB of RAM and 64 on a GPU with 8 GB of memory.
We did not find a comparable hardware table in the Chronos README. Its production options, SageMaker JumpStart endpoints, run on CPU or GPU.
What does self-hosting cost me in practice?
Mostly maintenance. You pin library versions, keep the weights cached, size the hardware, and track new releases. The TimesFM README itself says the open version "is not an officially supported Google product".
In our experience, self-hosting is the better choice when your data cannot leave your network, when you already have spare GPU capacity, or when you want to fine-tune. The TimesFM README links a fine-tuning example with LoRA.
How do I use TimesFM on Google Cloud?
Use BigQuery's AI.FORECAST SQL function, which runs a built-in TimesFM model with no model to create or train. The AI.FORECAST reference lists TimesFM 2.5 as the default and TimesFM 3.0 as Preview.
Key limits from that page:
- Horizon: 1 to 10,000 points for TimesFM 2.5, 1 to 1,024 for TimesFM 3.0.
- Context: at most 15,360 points for TimesFM 2.5 and 2,048 for TimesFM 3.0; older points are ignored.
- Output: a median (
forecast_value) and a prediction interval at aconfidence_levelyou set, default 0.95. - Covariates (
past_covariate_cols,future_covariate_cols) need TimesFM 3.0, which is under Pre-GA terms. - A series needs at least 3 data points.
Billing is at the BigQuery ML prediction rate, and the page says TimesFM 3.0 moves to token-based pricing from 1 December 2026. The TimesFM README also lists Google Sheets and a Vertex AI Model Garden endpoint.
From Python, send the SQL with the BigQuery client library:
from google.cloud import bigquery
client = bigquery.Client()
sql = """
SELECT *
FROM AI.FORECAST(
TABLE `mydataset.daily_sales`,
data_col => 'units_sold',
timestamp_col => 'sales_date',
horizon => 30)
"""
for row in client.query(sql).result():
print(row.forecast_timestamp, row.forecast_value,
row.prediction_interval_lower_bound, row.prediction_interval_upper_bound)This is the natural choice if your data already lives in BigQuery, and the route Google's README documents for commercial use of TimesFM 3.0.
Is there a managed option for Chronos-2?
On AWS, yes, though it runs in your own account rather than as a per-call API. The Chronos README recommends AutoGluon-Cloud for real-time, serverless or batch inference, or SageMaker JumpStart for real-time endpoints on CPU or GPU.
Choose this if you are on AWS and want the endpoint and data inside your account. You still pay for and manage the endpoint while it runs.
How do I call TimesFM or Chronos-2 through a hosted API?
Send your history to a forecasting API that already runs the model. Ephemeris is one example: it serves TimesFM 2.5 and Chronos-2 alongside Toto 2, TiRex-2, PatchTST-FM r2 and FlowState r1, over REST and as a remote MCP server for AI agents.
It has three modes:
explicit: run one model you name, such astimesfm25orchronos2.route: Ephemeris picks a model for the data.ensemble: run every compatible model and combine them, weighting the more accurate ones more.
A minimal explicit-mode call to TimesFM 2.5:
import os
import requests
resp = requests.post(
"https://ephemeris.cascade.industries/api/v1/forecast",
headers={"Authorization": f"Bearer {os.environ['EPHEMERIS_API_KEY']}"},
json={
"mode": "explicit",
"model": "timesfm25",
"series": [{"values": [100.2, 101.4, 99.8, 103.1, 104.6, 102.2], "freq": "H"}],
"horizon": 24,
"quantiles": [0.1, 0.5, 0.9],
},
timeout=60,
)
resp.raise_for_status()
print(resp.json()["forecasts"][0]["quantiles"]["0.5"])Change "model" to "chronos2" for Chronos-2, which also accepts past and future covariates. Model names, horizon limits and prices come from GET /api/v1/models, and each response reports the credits it cost.
Limits to know: up to 64 series per request, horizons up to 512 steps, and context truncated to each model's cap (16,384 points for TimesFM 2.5, 8,192 for Chronos-2). Ephemeris does not serve TimesFM 3.0, does not fine-tune, and does not run on-premises.
Which route should I choose?
Pick by where your data lives and how much infrastructure you want to own.
- Your data is in BigQuery, or you need TimesFM 3.0 commercially: use
AI.FORECAST. - You are on AWS and want Chronos-2 inside your account: use AutoGluon-Cloud or SageMaker JumpStart.
- Your data cannot leave your network, or you want to fine-tune: self-host the open weights.
- You want to call the models from any language or an AI agent, compare several models, or ensemble them, without running anything: use a hosted API.
FAQ
Is there an official TimesFM API?
Google offers TimesFM through BigQuery's AI.FORECAST SQL function, Google Sheets and a Vertex AI Model Garden endpoint, per the TimesFM README. Third-party APIs such as Ephemeris also serve TimesFM 2.5 per call.
Can I use TimesFM commercially?
TimesFM 2.5 weights are Apache-2.0, so yes. TimesFM 3.0 weights are non-commercial when self-hosted; the README says commercial use of 3.0 is permitted through Google Cloud services such as BigQuery ML.
Can I run TimesFM without a GPU?
Yes. The official TimesFM agent skill lists TimesFM 2.5 at about 1.5 GB of RAM on CPU and says CPU mode is slower but works.
Is there a hosted Chronos-2 API?
Amazon's recommended routes, AutoGluon-Cloud and SageMaker JumpStart, deploy Chronos-2 into your own AWS account. Third-party APIs such as Ephemeris serve Chronos-2 per call by the name chronos2.
Does Chronos-2 support covariates?
Yes. Chronos-2 handles past and future-known covariates in its own package, and Ephemeris passes them through when you call chronos2.
Related
- TimeGPT alternatives in 2026: hosted and open-source forecasting models
- Chronos-2 vs TimesFM 2.5 vs Toto 2 vs TiRex-2: one harness, same data
- Time-series foundation models: a practical guide (2026)
- Which time-series foundation model should I use? A decision guide
- How to give Claude, Cursor or any AI agent a forecasting tool (MCP)
- Model pages, pricing and the full API reference