Prediction Market Tools That Score a Forecast

Cards that compute a number for whether a forecast was any good. They do not compute the same number, and two point in opposite directions.

Last updated

What the flag claims

capabilities.calibration_scoring says the product computes a number for whether a forecast was any good — a Brier score, a log score, a calibration curve, a tracked record. Eleven of the sixty cards here carry it, and it is the one capability in this catalogue that is the product on eight of them rather than a feature on the side.

What the flag does not say is which number, what it is relative to, or who is being measured. Those three questions have at least seven answers between these eleven cards, and the answers are not interchangeable.

Who is being scored

You, on your own forecasts. Metaculus, Good Judgment Open, Manifold, Fatebook, Confido, Prediki, Quantified Intuitions and the Social Science Prediction Platform. Eight of the eleven, and all but Manifold settle in play money or in nothing at all — see play-money platforms for what that removes.

A platform, by somebody outside it. Brier.fyi matches questions across Kalshi, Polymarket, Manifold and Metaculus by hand and scores each platform's market on the outcome. Its matching rule is published, which is the part most cross-platform accuracy work skips: two markets may be grouped only if their differences would not have more than a one-percent chance of changing the resolution.

A stranger's wallet. Convexly's Edge Score composites three pillars — baseline-adjusted Brier, concentration of profit and loss, and count of resolved positions — against a frozen 8,656-wallet comparison set, and says in as many words that historical diagnostics do not establish future profit. It is the only card here scoring people who never opted in.

Seven numbers, and two of them point opposite ways

This is the trap, and it is arithmetic rather than judgement.

Confido's score runs backwards from everyone else's. The code computes (0.25 − (outcome − p)²) × 4: the range is −3 to 1, a 50% forecast scores 0, and higher is better. Every other Brier score in this catalogue is lower-is-better. The sign convention alone will catch anybody who skims.

Good Judgment Open reports three numbers, and the third is the one on the leaderboard. Your Brier score, the median of everyone's daily Brier scores on that question, and the Relative Brier Score — your average daily Brier minus the crowd's median, scaled by the share of days you had a forecast standing. Negative beats the crowd, and a question you skipped scores zero rather than penalising you.

Metaculus is built on the log score, with a Baseline score against a fixed benchmark and Peer scores that sum to zero across everyone on the same question by construction.

Quantified Intuitions uses a relative log score against the crowd's forecast at the vantage date — match the crowd exactly and you score zero — plus Greenberg interval scoring on its calibration and estimation tools. The word "Brier" does not appear in its codebase.

Hypermind has no scoring rule at all. You are ranked by the sum of the square roots of your profits minus the sum of the square roots of your losses, so the same total profit scores better extracted from many questions than from one. That measures skill at this market, not calibration in the abstract.

The Social Science Prediction Platform scores error, not probability. Mean absolute error against the realised estimate, standardised by the baseline standard deviation, with a relative measure against the mean forecast on the same question. What is elicited is an effect size, so there is no Brier score and no calibration chart anywhere on the site.

Prediki publishes none of it. Expertise points are described and not defined, and there is no formula, no calibration chart and no Brier or log score on the site.

What a forecasting platform's score actually measures is the page for the general version of this, and Metaculus vs Good Judgment Open vs Manifold for the three that a reader most often puts side by side.

What none of them scores

  • Numeric questions, on Confido. Automatic scoring is binary only; the scoring function returns nothing for a numeric answer space, and the institute's own guidance is to export CSV and compute your own.
  • Continuous and conditional markets, on Brier.fyi, which does not evaluate them yet — and a market with no trades is ignored rather than assumed to be at 50%.
  • Anything, until it resolves. The Social Science Prediction Platform scores against a study's reported estimate, which can take years, and some studies never report.

What to check before quoting a number

Which direction is better. What it is relative to — a crowd median, the median of the other markets on the same question, or a frozen cohort. Whether the record travels off the site. And who paid for the scoring: Brier.fyi states on its own about page that it was funded partly by the Manifold Community Fund, which is one of the four platforms it scores.

Nothing in this collection has been forecast on by this site with a scored account. Formulas, score ranges, cohort sizes and funding disclosures above are read from each platform's own documentation, code and reports, and dated on the cards.

All 11 of them

Showing 11 of 11

FAQ

Can I compare a score from one of these with a score from another?

Almost never without converting it first, and on two of these cards not at all. Confido computes (0.25 − (outcome − p)²) × 4, which runs from −3 to 1 and is higher-is-better; every other Brier score in this catalogue is lower-is-better. Hypermind ranks by the sum of the square roots of your profits minus the sum of the square roots of your losses, which is not a scoring rule at all. Read what the number is before reading the number.

Which of these score me, and which score somebody else?

Eight score you — Metaculus, Good Judgment Open, Manifold, Fatebook, Confido, Prediki, Quantified Intuitions and the Social Science Prediction Platform. Brier.fyi scores four platforms against each other on hand-matched questions. Convexly scores Polymarket wallets. Manifold is in both halves, because it publishes a platform-wide calibration curve for itself as well as scoring its own users.

Does a good score here travel anywhere?

Some do and some are local by construction. Good Judgment recruits Superforecasters each autumn from forecasters who have answered at least 100 GJ Open questions, so a record there is the one with a published destination. A Quantified Intuitions rank is a rank at a training site on questions that already resolved. Prediki publishes no scoring formula, no calibration chart and no Brier or log score at all, so there is nothing to lift out of it.

Is a platform's own accuracy number worth anything?

It is worth what its method and its data are worth. Manifold's platform-wide Brier score is published rather than claimed, and it is a statement about thousands of resolved questions in aggregate rather than about the one in front of you. Hypermind's July 2026 accuracy report is its own analysis of its own data, and it is checkable because it also published a 44 MB CSV of every trade since May 2014.