# Calibration

Also written reliability diagram, calibration curve.

*https://predictionmarkets.tools/glossary/calibration · next to Prediction Market Analytics & Dashboards*

**Definition:** Whether the things forecast at a given probability happen at that rate: of everything priced or forecast at 70%, about 70% should come true. It is a property of a set of forecasts, never of a single one, and it says nothing about whether the forecasts were useful. In this catalogue the word also names a job, a capability flag and a correction applied to a model's outputs, and the four uses are not interchangeable.

Calibration is the most-used word in this catalogue's scoring vocabulary and the least precise. A
venue's price, a forecaster's record and a training exercise are all described as calibrated, and
a card carries `capabilities.calibration_scoring` whether it draws a calibration curve or only
prints a Brier score. This page pins the strict meaning down and lists the looser ones next to it,
so that a claim can be read for what it actually asserts.

## How it works

Collect a set of resolved forecasts. Group them by the probability that was stated — everything
said at 5 to 15 percent in one bin, 15 to 25 in the next, and so on — and in each bin count how
often the event happened. Plot stated probability across and observed frequency up. The
forecast-verification group of the World Meteorological Organization calls this a reliability
diagram, and the field's own name for the property is **reliability**: the closer the points sit
to the diagonal, the better calibrated the forecasts. Points below the line mean the forecasts
were too high for that bin; points above it, too low.

Three things follow from the construction, and each is routinely forgotten.

**It needs a set.** One forecast of 70% on something that did not happen is not miscalibrated; it
is one draw. A curve with a dozen forecasts per bin is mostly noise, which is why a well-drawn
chart prints the count beside each point.

**The bins are a choice.** Where the edges fall, and which forecast from each question enters the
sample, change the picture. Manifold's calibration page makes its choice explicit: it samples
trades rather than moments, weights by trade, restricts itself to resolved binary questions with
15 or more traders, and says its markets may be more accurate than the chart shows, because large
miscalibrated trades are usually corrected immediately. On 27 September 2026 the same page
reported a platform-wide Brier score of 0.17475.

**It is one part of accuracy, not all of it.** Allan Murphy showed in 1973 that the Brier score
splits exactly into three terms: a reliability term that adds to the score, a resolution term that
subtracts from it, and an uncertainty term fixed by the questions themselves. Resolution is the
ability to sort events into groups that turn out differently. A forecaster who says the base rate
every time is perfectly calibrated and has no resolution at all; the WMO page says of exactly that
forecast that it does not discriminate between events and non-events.

### The other three things the word is used for

- **A job.** `job: calibration` on a forecasting card means the product is hired to *train* the
  property — trivia with confidence levels, questions that already resolved, a personal tracker.
- **A flag.** `capabilities.calibration_scoring` means the product computes some number about
  whether a forecast was good. Most of those numbers are Brier or log scores, which contain
  calibration but are not it.
- **A verb.** In forecast verification, to calibrate a forecast is to adjust
  its outputs after the fact so the curve moves to the diagonal. A model described as "calibrated"
  in that sense has been corrected, not measured.

## Why it matters here

**A calibrated venue tells you nothing about the contract in front of you.** Calibration is a
statement about thousands of prices taken together. A venue can be close to the diagonal overall
while the thin market you are looking at is badly off it, and nothing on the aggregate chart will
say which. [How accurate the price is](https://predictionmarkets.tools/guides/how-accurate-the-price-is) covers what the research
finds about where market prices drift from the line and when.

**"Well calibrated" is not "good".** Because the base-rate forecaster is perfectly calibrated, a
dashboard or a platform that reports calibration alone can be flattering a forecaster who never
said anything. Ask for the resolution half as well, or for a proper score that contains both —
the [Brier score](https://predictionmarkets.tools/glossary/brier-score) does, and so does the log score.

**Check what a dashboard sampled before you compare two of them.** A chart built from the last
trade before close, one built from the price a week out and one built trade-weighted across the
market's life are three different claims, and they will disagree even on the same markets.
[Brier.fyi](https://predictionmarkets.tools/tools/brier-fyi) lets you pick the criterion point — the market midpoint, thirty days
before close — and publishes how it matched the questions; Manifold's page says it is
trade-weighted. A curve that says neither is a picture, not a measurement.

**A calibration-training score and a track record are different artefacts.** A
[Quantified Intuitions](https://predictionmarkets.tools/tools/quantified-intuitions) exercise tells you how your confidence
behaves on trivia and questions that have already resolved. A record in
[Fatebook](https://predictionmarkets.tools/tools/fatebook) or [Confido](https://predictionmarkets.tools/tools/confido) is your own claim about your own
questions. Neither is the thing a stranger can check the way they can check a public tournament
record, and the [calibration-scoring collection](https://predictionmarkets.tools/collections/calibration-scoring) sorts its
eleven cards by who is being scored and whether the record travels anywhere.

## Where you will meet this

- [Adjacent](https://predictionmarkets.tools/tools/adjacent.md)
- [Airavat](https://predictionmarkets.tools/tools/airavat.md)
- [Apify prediction-market scrapers](https://predictionmarkets.tools/tools/apify-prediction-market-scrapers.md)
- [Artemis prediction-market metrics](https://predictionmarkets.tools/tools/artemis-prediction-markets.md)
- [Brier.fyi](https://predictionmarkets.tools/tools/brier-fyi.md)
- [CCXT](https://predictionmarkets.tools/tools/ccxt.md)
- [Confido](https://predictionmarkets.tools/tools/confido.md)
- [Convexly](https://predictionmarkets.tools/tools/convexly.md)
- [Crypto.com Prediction](https://predictionmarkets.tools/tools/crypto-com-prediction-markets.md)
- [DepthFeed](https://predictionmarkets.tools/tools/depthfeed.md)
- [Polymarket dashboards on Dune](https://predictionmarkets.tools/tools/dune-polymarket-dashboards.md)
- [Fatebook](https://predictionmarkets.tools/tools/fatebook.md)

## Sources

1. [A New Vector Partition of the Probability Score](https://doi.org/10.1175/1520-0450(1973)012<0595:anvpot>2.0.co;2) — Journal of Applied Meteorology, volume 12, number 4, 1973-06-01. The paper that splits the Brier score into reliability, resolution and uncertainty; later work extends the partition rather than replacing it.
2. [Forecast verification - methods, issues and FAQ](https://jwgfvr.github.io/forecastverification) — WMO WWRP Joint Working Group on Forecast Verification Research, read 2026-09-27
3. [Calibration](https://manifold.markets/calibration) — Manifold, read 2026-09-27

*Last updated 2026-09-27. A reference page, corrected in place — not a dated post.*
