How accurate a prediction market's price actually is

What research establishes about reading 73 cents as a 73% chance — where market prices are calibrated, the deviations that recur, and what the horizon does.

On a liquid market with weeks rather than years to run, the published research says a price is a decent estimate of a probability. Two deviations recur: prices compress towards fifty cents as the horizon lengthens, and contracts bought in single-digit cents have historically returned less than their price implies. What a price is evidence of is narrower than a poll: it is a clearing level between the people who showed up with money.

Seventy-three cents on a contract that pays a dollar. The number invites one reading — a 73% chance — and that reading is mostly defensible, which is the difficulty: mostly is not always, the exceptions are systematic rather than random, and they are largest in exactly the places a reader is most tempted to act on them.

This is the one page on this site whose central claim rests on peer-reviewed literature rather than on a venue's own document. There is no rulebook to cite here. A venue can tell you how a contract settles and who decides; whether its prices track the world is something somebody else has to measure, and every peer-reviewed measurement cited below was made on an exchange that predates the venue you are looking at. Where a finding comes from a particular era or a particular kind of market, this page says so, because that is usually the part that decides whether it applies to you.

Two neighbouring questions are handled elsewhere and are not repeated here. How a forecasting platform's leaderboard number is computed — Brier, log score, the crowd-relative variants — is what a forecasting platform's score measures. Whether there is a real price on the screen at all, and what a number with a thousand dollars behind it is worth, is where the liquidity comes from. This page assumes the liquid case: a two-sided book you could actually trade against. The question is how much to trust the number it prints.

How it works

Calibration is one property, and a narrow one. Take every contract a venue quoted at 20 cents, wait for all of them to resolve, and count. If 20% of them resolved yes, the venue's 20-cent prices were calibrated. Page and Clemen give the definition in one line: a market price is calibrated when the expected frequency of occurrence equals the price.

Notice what the definition does and does not promise. It is a statement about a population of contracts, never about the one in front of you — a perfectly calibrated venue still resolves one in five of its 20-cent contracts yes, and no individual price is ever shown to be right or wrong by its outcome. And it is silent on sharpness: a venue that priced every contract at the base rate of its category would be beautifully calibrated and useless.

A calibration chart plots price on the horizontal axis against the realised frequency on the vertical, with the diagonal as perfection. Points below the diagonal are prices that were too high; points above it, too low. Four questions decide what any such chart is actually showing, and they are rarely printed next to it.

Which price was sampled. A contract has a price at every moment of its life. The last trade, the midpoint of the book, a time-weighted average and a trade-weighted average are four different data sets and they do not produce the same curve. Manifold is the useful example because it publishes the choice: its calibration page reports a Brier score of 0.17375 over a stated 98k trades, sampled at 2% of past trades hourly on resolved binary questions with 15 or more traders, grouped into probability bands — and states that the figure is trade-weighted rather than time-weighted, adding that market accuracy may be better than reflected because large miscalibrated trades are usually corrected immediately. That is a methodology note doing its job: it tells you the number is a claim about trades, not about screen-time.

Which contracts are in the sample. Resolved ones. Anything voided, annulled or quietly delisted is absent, and those are disproportionately the questions whose criteria did not survive contact with the world. No chart can fix this and every chart inherits it.

Whether the contracts are independent. They are not. Five contracts on five candidates in one race are one event, and treating them as five observations makes a confidence band narrower than the evidence supports. Page and Clemen handle it by grouping interdependent markets into what they call competitions — all markets on the 2004 Democratic primaries are one cluster — and bootstrapping over clusters rather than over transactions, drawing ten transactions per competition so that a handful of huge markets cannot carry the estimate. A calibration curve published without that step is reporting more precision than it has.

How the price axis is cut. Fixed bins blur the ends, which is the region under dispute. Page and Clemen use a local linear regression with a window of 0.10 instead, specifically to estimate what happens near 0 and near 1 with usable precision.

Where the price is close to the probability

The evidence for accuracy is real, it is oldest in election markets, and it is strongest close to resolution.

Wolfers and Zitzewitz, in the survey published in the Journal of Economic Perspectives in 2004, report that in the week before the election the Iowa Electronic Markets predicted the two-party vote shares of the previous four US presidential elections with an average absolute error of around 1.5 percentage points, against 2.1 percentage points for the final Gallup poll of the same elections, and that the market's error shrinks steadily as the election approaches. Two qualifications travel with that number and are usually dropped: those were index contracts on vote share rather than binary contracts on a winner, so it is a statement about forecast error and not about calibration in the strict sense above; and it is a comparison against a poll, which is a mechanical benchmark rather than a hard one.

The modern equivalent is not yet in a journal. A 2026 preprint by Nam Anh Le, Decomposing Crowd Wisdom, measures calibration across 353 million trades and 429,000 binary contracts on Kalshi and Polymarket, and decomposes it by domain, time to resolution and trade size. It is worth reading for two reasons and citing with care for one: it has not been through peer review. Its useful conclusion is in its own summary — that a price's meaning depends on what, when and how much is traded. Its honest one is that under conservative event-clustered standard errors roughly half of the raw variation between cells is estimation noise, which is the sort of thing a headline number about "market accuracy" never mentions.

The defensible summary of all of it: on liquid contracts, close to resolution, in a domain with a steady supply of questions, market prices have repeatedly been found to be a reasonable estimate of the frequency they are followed by. Every word in that sentence is doing work.

The deviations that are not noise

Cheap outcomes and near-certainties

The favourite-longshot bias is the oldest documented pattern in this family and the most often misapplied. Its strongest evidence comes from horse racing, not from event exchanges. Snowberg and Wolfers, in the Journal of Political Economy in 2010, use all 6.4 million US horse-race starts from 1992 to 2001. Across that sample the rate of return on a position on a horse at odds of 100/1 or longer is about −61%, while a position on the favourite in every race loses 5.5%, and picking at random loses 23%.

Read those three numbers carefully, because the interesting thing is the slope and not the level. The 23% is mostly the track's own cut, taken out of every pool to fund operations. What the bias describes is that the loss is over ten times worse at the long end than at the short end, on the same track, under the same deduction. The paper's contribution is to test why: by checking whether the preferences that explain single-outcome positions also explain the pricing of compound ones — an exacta, a quinella, a trifecta — it finds the evidence favours misperception of probabilities, in the direction prospect theory predicts, over a taste for risk.

That is a horse-racing result. It does not transfer to an event exchange by assertion, and the prediction-market evidence is both weaker and more interesting. Page and Clemen estimate it directly on 512,612 transactions from 1,787 Intrade markets, grouped into 597 competitions, all of them markets that opened between August 2002 and February 2007 and all of them lasting more than a month:

Price on the screenShare that resolved yesWhich markets
0.2015.3%all 1,787
0.8087.4%all 1,787
0.2010.9%political only
0.8092.8%political only

The shape is the classic one — cheap contracts too dear, likely contracts too cheap — and it is markedly stronger in political markets than in the rest, a difference the authors report without explaining.

The most recent measurement on a venue a reader can actually use is another 2026 preprint, The Favorite-Longshot Bias in Prediction Markets by Marcos Cardozo and José Ignacio Rivero-Wildemauwe, over 588 million Polymarket trades by 2.48 million accounts. In observed transaction flows, purchases below 10 cents lose 19.3 cents per dollar while purchases at or above 90 cents earn 0.83 cents per dollar. The finding that matters most is its fragility: longshots lose 6.3 cents per dollar when every contract is weighted equally, and gain 4.1 cents when related contracts are first grouped under their parent event — and the two-sided pattern shows up in crypto and politics markets but not in sports. Same venue, same data, opposite sign, depending on a bookkeeping decision. Anyone quoting a single number for the size of this bias has chosen one of those conventions without telling you.

What happens at 1 cent and at 99

Two things are worth separating here, because the received wisdom overstates one of them.

The received wisdom is that very unlikely outcomes are always overpriced. Page and Clemen found the opposite at the extremes of their own data: their curve is S-shaped through the middle, and at prices approaching 0 and 1 the prices looked well calibrated — a feature their model does not predict and which they discuss rather than explain. So "longshots are always overpriced" is not what the prediction-market evidence says; it is what the racing evidence says.

The mechanical point is sturdier. At a one-cent price the minimum tick is the entire price: there is no way to express a belief of 0.3% on a venue quoting whole cents, so every belief between zero and about one and a half per cent is rounded into the same number. Add resolution risk — the chance that the question settles on a technicality rather than on the world — and a floor sits under the cheap side of every contract that has nothing to do with the event. Wolfers and Zitzewitz found the same thing empirically on TradeSports in 2003, where the contracts on extreme year-end levels of the S&P 500 were priced above their equivalents built from Chicago Mercantile Exchange options, and the gap implied a small arbitrage that persisted for most of that summer and reappeared the following year. Their conclusion was that prediction markets are likely to perform poorly at predicting small-probability events.

The practical version: below about 5 cents, stop reading the price as a probability and start reading the resolution criteria. Who decides the outcome matters more down there than any estimate of the event.

The horizon, which is the largest effect on this page

The clearest result in the literature is also the easiest to act on: calibration decays with time to resolution, in a known direction.

Page and Clemen give the mechanism before the measurement. Money committed to a contract earns nothing while it waits, so a trader whose belief is close to the price has no reason to trade at all, and that no-trade band widens as the wait lengthens. It widens asymmetrically: buying the expensive side ties up more money per contract, so the expensive side loses its marginal traders first. The result is a price pushed towards 50 cents — which is the favourite-longshot shape, arrived at from time preference alone, with no behavioural bias required.

Their data agrees. Splitting the sample at 100 days to expiration, the bias is stronger on the longer horizon, for political and non-political markets alike, and in the parametric estimation the coefficient on time left to expiration is negative and significant. Their conclusion is blunt: markets on events a year or more away will either fail to generate a price, or generate one that is systematically biased.

The obvious remedy is to pay interest on committed balances, and their data already contains a partial test of it — Intrade paid interest on balances over 20,000 dollars, and the long-horizon bias persisted anyway. Among today's venues, ForecastEx pays interest on posted collateral by rulebook, which removes one term from this equation without touching the rest of it. What the collateral is doing over those months, and who earns on it at each venue, is the subject of what your money does between the trade and the resolution.

So the same number means different things at different distances. On a contract resolving this month, 73 cents is close to 73%. On a contract resolving in 2029, 73 cents is evidence of a higher number than 73, and the further out it sits the more you should widen that adjustment — in direction, since nobody can give you the magnitude for your particular market.

What a price is evidence of

A price is the level at which two sets of people, both able to reach this venue and both willing to lock up money until resolution, stopped disagreeing enough to trade. That is a narrower object than a poll, and it is worth naming the four narrowings.

It is an estimate of the average belief of traders, under conditions. Wolfers and Zitzewitz's Interpreting Prediction Market Prices as Probabilities sets out sufficient conditions under which market prices correspond with mean beliefs, and finds that across a broad class of models prices are usually close to the mean belief of traders — their own summary calls them useful, though sometimes biased, estimates of average beliefs. The parameters that drive the gap are the degree of risk aversion and the shape of the distribution of beliefs.

It is a percentile of that distribution, not its middle. The clearing condition places the price not at the median belief but at a percentile pulled towards one half — the result Page and Clemen build on, from Ali and Manski. Combine it with the no-trade band above and the picture inverts the intuitive one: the price is set by the people who disagree with it most, because the people who agree with it have no reason to act.

It is a risk-neutral price. Wolfers and Zitzewitz are explicit that the price of a winner-takes-all security is essentially a state price, equal to the event's probability under the assumption of risk neutrality, and that if the event is correlated with traders' marginal utility of wealth then the two come apart. For most contracts the sums involved are small enough that the assumption is reasonable. For a contract on a recession, or on a rate decision, held by people whose jobs and portfolios move with the answer, it is doing more work than it looks.

It is the belief of whoever was allowed in. The population is self-selected, funded, KYC'd where the venue requires it, and geographically filtered — why a venue is unavailable where you are is a description of who is missing from every price on that venue. A market that excludes an entire country is not sampling that country's view of its own election.

What volume tells you, and what it does not

The instinct is that a heavily traded market is a better-calibrated one. The evidence does not support the strong form of it.

Page and Clemen tested exactly this and report that neither volume per transaction nor a market's total volume had a significant effect on calibration — with the 10th and 25th percentiles of market volume in their sample at 269 and 858 contracts, so the test had genuinely thin markets in it. Their own reading is that low volume does not appear to introduce additional systematic bias.

Trade size does appear in the modern data, and not in the flattering direction: Le's preprint reports that large political trades on Kalshi are associated with further compression towards 50% — a calibration-slope gap of roughly one half that survives clustered bootstraps on Kalshi, though not robustly on Polymarket. Whatever a whale's order is evidence of, it is not evidence that the resulting price is better calibrated.

What volume and open interest are genuinely evidence about is whether you can transact: whether the displayed number is reachable, how much of it you would move, and whether anyone is left to close against. That is a different question with its own measurements — depth at two cents from the midpoint, open interest rather than turnover — and it is set out in where the liquidity comes from.

What it costs

The price of reading a price wrongly is easiest to see through the opposite exercise: what a documented bias is actually worth to somebody trying to trade it out of existence.

Page and Clemen estimated their calibration curve on the first half of their sample and traded it on the second — markets that opened between January 2005 and February 2007 — with two simple rules. Buying when the price was between 0.08 and 0.10 or above 0.61 and selling at every other price returned an average of 5.25% per position; the cruder rule of buying everything above 50 cents and selling everything below it returned 9.55%. Both are rates of return on capital committed until resolution on markets lasting months, not annualised figures and not per-trade edges on a fast book.

Then the arithmetic that keeps the bias alive. At a discount rate of 10% the first strategy returns 2.47%. At 15% it returns nothing at all. At 25% the second strategy returns nothing either. And none of that includes what the venue charges: entry and exit are separate taxable events on most schedules here, which what a trade actually costs converts into comparable units.

The same shape appears in racing, where the average return at every odds level is negative because the track's cut is deducted first — the bias is the slope across the levels, not a positive number at any of them.

That is the cost, and it runs both ways. For a reader, a few cents of systematic bias on money locked up for a year is not an opportunity. For the market, the fact that it is not an opportunity is precisely why nobody removes it, and why the same pattern shows up in Intrade data from the 2000s and in a Polymarket study published in 2026.

What you can do about it

Ask the horizon before you ask anything else. It is the one adjustment with a direction the literature agrees on. Inside a month or two, take the price roughly at face value. Past a year, assume compression towards the middle and read a price below 50 as a ceiling and a price above 50 as a floor. Nobody can give you the size of the correction for your market; you can still get the sign right for free.

Convert the number into the sentence it actually supports. Not "there is a 73% chance", but "the people who could reach this venue, and were willing to tie up money until resolution, cleared at 73 cents". Then ask who is missing from that sentence — a jurisdiction, a professional class, anyone who found the capital cost too high — and whether their absence points in a particular direction on this question.

Below 5 cents and above 95, read the rules instead of the price. The tick is the price down there, and resolution risk is a larger term than anything about the event itself. This is where the cost of being wrong about who decides is highest, and where the price has least room to tell you anything.

Check calibration yourself — the ingredients are public. A venue's resolved-market list, its price history and its outcomes are usually all available through its API, which is enough to bin prices and count. Brier.fyi already does the cross-venue version on hand-matched questions across Kalshi, Polymarket, Manifold and Metaculus, publishes its pipeline as open source, and its card records that the pipeline is currently paused — which is itself a warning about how much maintenance this kind of scoring takes.

Demand a denominator from any accuracy claim, including a venue's own. How many resolved contracts, which price was sampled, over what date range, and what happened to voided markets. Manifold's calibration page answers all four and is the standard to hold others to. A claim of "X% accurate" with none of them attached is marketing, and the same is true of a chart with no sample size beside its points.

Do not read a confidence band that ignores clustering. If a chart's error bars were computed as though every contract on the same event were an independent observation, they are too narrow, and small apparent deviations from the diagonal are not evidence of anything.

And keep prices and scores in separate boxes. A forecasting platform's leaderboard number and a venue's price are measured differently, incentivised differently and not comparable as published — which is the whole of what a forecasting platform's score measures, and the reason a cross-venue accuracy claim needs hand-matched questions before it means anything at all.

Tools this bears on

Cards in the catalogue where what is above changes the decision.

  • Brier.fyi

    Brier scores and letter grades for matched questions across four platforms.

    FreeFree tierOpen source

  • Kalshi

    A CFTC-designated exchange for event contracts, settled in dollars against named sources.

  • Polymarket

    Self-custody event contracts on an on-chain order book, resolved by the UMA oracle.

  • Manifold

    Anyone can open a question, anyone can take a side, and the currency buys nothing.

    $5/moFree tierOpen source

FAQ

Does a price of 73 cents mean a 73% chance?

On a liquid contract resolving within weeks, that is a defensible reading and the research broadly supports it. On a contract resolving in a year or more, the documented pattern is compression towards the middle — low prices sit above the frequency they are followed by, high prices below it — so 73 cents on a long-dated question is evidence of a number higher than 73, in a direction you can name but a magnitude nothing in the literature pins down for your particular market.

What is the favourite-longshot bias, and does it apply to event contracts?

It is the long-documented pattern that cheap outcomes are priced above the rate at which they happen and likely outcomes below it. The strongest evidence is from pari-mutuel horse-race markets, not from event exchanges. On prediction markets the pattern has been measured too, but its size depends on how contracts are grouped, and one 2026 preprint finds it in crypto and politics markets on Polymarket and not in sports.

Does high volume mean a price is more trustworthy?

Not as evidence about calibration. In a study of 512,612 transactions that tested exactly this, neither volume per transaction nor a market's total volume had a significant effect on how well prices tracked outcomes. Volume tells you whether you can get in and out at the displayed price, which is a different and still useful thing.

How do I read a calibration chart?

The horizontal axis is the price, the vertical axis is the share of contracts at that price that resolved yes, and the diagonal is perfect calibration. Points below the line mean the price was too high. Before reading anything into it, ask which price was sampled, how many resolved contracts are behind each point, what happened to markets that were voided, and whether the confidence band accounts for contracts on the same event not being independent.

If the price is systematically wrong, why does nobody correct it?

Because the correction is small and slow. The one published attempt to trade a documented bias out of sample earned a single-digit to low-double-digit rate of return per position held to resolution, and a discount rate of 15 to 25 per cent on the capital cancelled it entirely — before any fee. A bias worth a few cents on money locked up for a year is not money on the floor, which is exactly why it survives.

Sources

  1. Prediction Markets National Bureau of Economic Research, working paper 10504; published in the Journal of Economic Perspectives, volume 18, number 2, pages 107-126, . The accuracy comparisons it reports are still the baseline later calibration work argues with rather than replaces.
  2. Interpreting Prediction Market Prices as Probabilities National Bureau of Economic Research, working paper 12200, . A derivation of when a price equals a mean belief, not a description of a venue; its conditions do not expire with the exchanges it predates.
  3. Explaining the Favorite-Long Shot Bias: Is it Risk-Love or Misperceptions? Journal of Political Economy, volume 118, number 4, pages 723-746, . Its horse-race data ends in 2001, and the pattern it explains is reported again on a current venue by the 2026 preprint cited here.
  4. Do Prediction Markets Produce Well-Calibrated Probability Forecasts? The Economic Journal, volume 123, number 568, pages 491-513, . Its Intrade data ends in 2007, which dates the evidence; no later work cited on this page contradicts the horizon result.
  5. Decomposing Crowd Wisdom: Domain-Specific Calibration Dynamics in Prediction Markets arXiv preprint 2602.19520, version 2 (not peer reviewed),
  6. The Favorite-Longshot Bias in Prediction Markets: Evidence from Polymarket arXiv preprint 2609.12878 (not peer reviewed),
  7. Calibration Manifold, read

The catalogue next door

This page is background, not a listing. The products it bears on are in Forecasting Platforms, each filled in against the same schema, with the fields to narrow it yourself.

Last updated . Corrected in place: this is a reference page, not a dated post.