What an AI weather model gives you, and what it does not
The named models are public, cheap and openly licensed. What they output is a gridded value at fixed hours, which is not what a contract measures.
A gridded forecast: one value per variable per 0.25-degree cell at six- or twelve-hour steps, from weights you can download and run. What it does not give you is the settlement variable — one station's maximum across a day that starts at midnight local standard time — or a calibrated probability at that station. Closing those two gaps is the work, and the model is the cheap part.
The question people arrive with is which model to run. It is the wrong end of the problem, and unusually easy to show why.
The public field is genuinely good, genuinely cheap and mostly openly licensed. You can download weights, run a global forecast on one machine in minutes, and get output competitive with a national meteorological centre. What you cannot do is read a temperature contract's settlement value off that output, because the model does not produce that quantity — not approximately, not at the wrong precision, but not at all. Two conversions stand between them, and both are station work rather than model work.
This page is about those conversions, and about what the published evidence says happens when people skip them. It is not about what any of this is worth; for what a price already contains, see how accurate the price is.
How it works
A global machine-learned forecast model is trained on reanalysis — usually ERA5, a model run backwards across roughly eight decades with observations assimilated into it, at 0.25 degrees and hourly resolution. At run time it takes an analysed state of the atmosphere and steps forward, emitting a full grid of variables at each step.
The named systems and what they actually emit:
- GraphCast, deterministic, and GenCast, its diffusion-based probabilistic successor, both from Google DeepMind and both in the WeatherNext repository alongside WeatherNext 2. GenCast's paper describes it as generating "an ensemble of stochastic 15-day global forecasts, at 12-hour steps and 0.25 degree latitude-longitude resolution" for "over 80 surface and atmospheric variables", in eight minutes, with "greater skill than ENS on 97.4% of 1320 targets we evaluated".
- AIFS, ECMWF's machine-learned system, which went operational on 25 February 2025 beside the physics-based IFS — the first of these to run in a national centre's operational suite. It runs at a grid spacing of about 28 km, updates every six hours, and ECMWF reports gains of up to 20% on some measures including tropical cyclone tracks. Its output is published in ECMWF's open data at 0.25 degrees under CC-BY-4.0, alongside IFS.
- Pangu-Weather, FourCastNet and Aurora, the rest of the public field, at the same order of resolution.
Every one of those outputs the same shape of thing: a value per variable, per grid cell, per step.
Gap one: the variable is not the variable
A contract settles on the maximum temperature recorded by one instrument at one named station. A model outputs 2-metre temperature for a 0.25-degree cell, roughly 28 km across, over the model's own idealised surface.
Those are different quantities, and the difference is structured rather than random. A cell containing a coastline averages water and land. A cell containing terrain averages elevations a station does not sit at. An airport sensor sits on a flat, open, often paved site whose behaviour at the daily maximum differs from the area around it in a way that is consistent from day to day. This is why the classical output of an operational forecast system is not the raw grid but statistically post-processed guidance: in the United States the National Blend of Models applies decaying-average and quantile-mapping bias correction and publishes at 2.5 km over the contiguous states, and version 5.0 has been operational since 5 May 2026.
Published verification usually scores models against reanalysis on their own grid, which measures whether the model reproduces the analysis — not whether it predicts an instrument. Those are not the same test and the second is harder.
Gap two: the step is not the day
GenCast steps at 12 hours. AIFS steps at 6. A daily maximum is the highest value across a continuous window, and on a US temperature contract that window runs midnight to midnight local standard time, which for most of the year is 1 AM to 1 AM on the clock.
Two instants twelve hours apart do not contain a maximum. Recovering one means interpolating to sub-daily resolution, aligning to the right window, and knowing which side of the boundary a warm late evening falls on. None of that is modelling; all of it decides contracts. What a temperature market settles on works through the window and the precision in detail.
Gap three: a run is not a probability
A deterministic model gives one number. A contract is priced in probability. An ensemble gives a distribution over grid cells, which is a description of the model's own uncertainty — a different object from the probability that a specific instrument reports a specific integer under a specific rounding rule.
Turning one into the other is calibration, it requires that station's history, and it is the step most likely to be skipped because the ensemble spread looks like a probability already. What a proper scoring rule measures, and why a well-calibrated forecast and an accurate one are not the same claim, is on what a forecasting score measures.
The published evidence that this is not hypothetical
Ennis and co-authors evaluated GraphCast, Pangu-Weather and NOAA's UFS GEFS against 60 heat waves across four boreal seasons and four regions of the contiguous United States, out to 20 days of lead time. They report consistent regional cold biases in the five to ten days of lead time before heat wave onset, in the machine-learned models and in the physics-based ensemble alike, persisting before and during the events, with Pangu-Weather in winter the exception — warm-biased before onset. GraphCast was the most skilful of the three.
Read that as a statement about regimes. The days on which a temperature contract's outcome is least obvious are the days on which the published systems are documented to share a bias, and a shared bias is invisible to anyone comparing two models against each other.
What it costs
The model is the cheap part, and that is the surprise.
Inference is minutes on one accelerator; GenCast's paper quotes eight for a 15-day ensemble. Weights are free to download. ECMWF's open real-time output is free. The costs that are real sit either side of the run.
Initial conditions. A model needs an analysed atmospheric state. ERA5 is the training set, not a real-time input: it updates daily with a latency of about five days, and its early release can differ from the final version published two to three months later. Running in real time means taking an operational analysis, which is a different product with different availability and different terms.
Forecast history. Scoring yourself requires past runs, and the open real-time archives are rolling — ECMWF's holds roughly the most recent twelve runs, two to three days. A forecast archive is something you accumulate from the day you start storing it, not something you download later. Where weather data comes from covers which archives are free, which are not, and what each licence permits.
Attribution and training-data terms. CC BY 4.0 on weights permits commercial use and requires credit and an indication of modifications. The WeatherNext repository adds a second question most readers skip: the ERA5 and HRES training datasets "may be governed by separate terms and conditions", and the repository tells you to check that you can comply before use. It also describes itself as research code with no guarantees of API stability.
The baseline you are implicitly buying out of. NBM is free, post-processed, station-relevant and published hourly. Any private pipeline's real cost is the difference between what it produces and what that already gives away.
What you can do about it
Close the window gap before the model gap. Aligning to the climatological day and recovering a maximum from sub-daily values changes a forecast's relationship to a contract more than swapping one model for another does, and it is deterministic work with a right answer.
Calibrate against the settlement product, not against reanalysis. The station archives are free. Score your output against the same text product the contract reads, at the same precision, for the station the contract names — not against ERA5, and not against the nearest grid cell.
Measure against the free baseline explicitly. The comparison that answers whether a pipeline is worth running is against post-processed public guidance, not against raw model output or against nothing. If it cannot beat what is already published for free, that is the finding.
Score before you price, and keep the two numbers apart. How well calibrated you are is one question; what a contract is worth is another, and running them together is how a forecast that is merely different from the price gets mistaken for one that is better. A market price is itself a forecast with a track record that can be scored the same way.
Treat a disagreement between a model and a price as a question, not a signal. The published bias result above is the concrete reason: when the systems agree with each other and are wrong together, the disagreement with a price is largest exactly when the model is least reliable. What a bot cannot fix is the general version of this — automating a pipeline does not add information to it.
State which model, which version and which initial condition you used, when you record a result. These systems are versioned and replaced: AIFS reached operations in February 2025 and has been updated since; NBM has moved through five major versions. A score recorded without a version is not reproducible, and the honest form of a calibration record is one someone else could repeat.
For the platforms where forecasts are submitted, scored and ranked in public — and for the scoring conventions they use — the forecasting platforms section carries the per-product detail.
Tools this bears on
Cards in the catalogue where what is above changes the decision.
OrcaLayer
Polymarket whale analytics indexed from Polygon, with a published farmer filter.
$9.99/moFree tier
Brier.fyi
Brier scores and letter grades for matched questions across four platforms.
FreeFree tierOpen source
Kalshi API
REST, WebSocket and FIX access to a CFTC-regulated event exchange.
Free tier onlyFree tier
FAQ
Which AI weather model should I run?
The question decides less than it appears to. The public field — GraphCast, GenCast, WeatherNext, Pangu-Weather, FourCastNet, Aurora, ECMWF's AIFS — all produce gridded output at about 0.25 degrees, and published comparisons separate them by margins smaller than the gap between any of them and a station's recorded daily maximum. Choose one, then spend the effort on the station and the window.
Can I use these models commercially?
Read two licences, not one. In Google DeepMind's WeatherNext repository the notebooks and code are Apache 2.0 and all other materials are CC BY 4.0, which permits commercial use with attribution — but the repository also warns that the ERA5 and HRES training datasets "may be governed by separate terms and conditions" and tells you to check before use. ECMWF's open real-time output is CC-BY-4.0 and may be redistributed commercially with attribution.
Is a model ensemble the same thing as a probability?
No. An ensemble is a set of runs whose spread describes the model's own uncertainty about a grid cell. A contract price is a probability about one station's reported value under a specific rounding and a specific window. Converting the first into the second requires calibration against that station's history, and an uncalibrated ensemble spread is usually too narrow at the extremes.
Why does a model with excellent published scores miss a station?
Because the published score is usually against reanalysis on the model's own grid, not against an instrument. A 0.25-degree cell is roughly 28 km across, and its 2-metre temperature is an area value over a model surface — not a shielded sensor on a patch of grass at one airport. Coastlines, elevation and urban surfaces make that difference systematic rather than random.
Do AI models handle heat waves as well as they handle ordinary days?
Published evidence says not, and the bias has a direction. Ennis and co-authors evaluated GraphCast, Pangu-Weather and NOAA's UFS GEFS across 60 heat waves, four seasons and four regions of the contiguous United States, and found consistent regional cold biases in the five to ten days of lead time before heat wave onset in all three systems, the exception being Pangu-Weather in winter.
Sources
- GenCast: Diffusion-based ensemble forecasting for medium-range weather (arXiv:2312.15796) — Price, Sanchez-Gonzalez, Alet, Andersson, El-Kadi, Masters, Ewalds, Stott, Mohamed, Battaglia, Lam and Willson, Google DeepMind,
- WeatherNext repository — GraphCast, GenCast and WeatherNext 2 — Google DeepMind, read
- ECMWF's AI forecasts become operational — European Centre for Medium-Range Weather Forecasts, read
- Turning Up the Heat: Assessing 2-m Temperature Forecast Errors in AI Weather Prediction Models During Heat Waves (arXiv:2504.21195) — Ennis, Barnes, Arcodia, Fernandez and Maloney, Colorado State University,
- ECMWF open data — real-time forecasts — European Centre for Medium-Range Weather Forecasts, read
- National Blend of Models versions — NOAA National Weather Service Meteorological Development Laboratory, read
- ERA5 hourly data on single levels from 1940 to present — Copernicus Climate Change Service, ECMWF, read
The catalogue next door
This page is background, not a listing. The products it bears on are in Forecasting Platforms, each filled in against the same schema, with the fields to narrow it yourself.
Last updated . Corrected in place: this is a reference page, not a dated post.