Forecasting Platforms

Scored questions and tracked records. What Brier, log and peer scores each measure, why a profit leaderboard is not an accuracy leaderboard, and what to check.

Last updated

The job is to be right, and to be scored for it

A prediction market venue and a forecasting platform look like the same object from outside: a question, a probability, a crowd, a chart moving. They are bought for two different jobs.

You go to a venue to take a position. The question is an instrument, the number on it is a price, and the thing you get at the end is a payout or a loss. You go to a forecasting platform to be right, on the record, in a way somebody can check afterwards. The question is a test item, the number on it is your stated belief, and the thing you get at the end is a score — usually permanent, usually public, usually attached to a username you have been using for years.

That difference sets everything else on these cards. There is usually no money, so usually no taker fee, no withdrawal and no regulator. Resolution is done by staff, by a question's author or by a vote of the people who traded it, rather than by a venue with a rulebook and an obligation, which makes it both more careful and less accountable. And the number the platform publishes is usually not a price at all — it is an aggregate of forecasts, typically a time-weighted median, which behaves quite differently from a price when new information arrives.

The exception is worth naming, because two platforms in this listing do have an order book or a market maker and do rank you by trading profit. What they do not have is a deposit: the currency is free, it is handed out once, and the only way money moves is a sponsor's prize pool divided among the forecasters who finished ahead. Read those two as contests with a market inside them rather than as venues you cannot withdraw from.

This is also the category where the audience is not the audience for the rest of the catalogue. The people who get value out of these platforms are researchers, analysts inside think tanks and government bodies, and people who want a defensible forecasting track record before they publish one. Several of the tournaments listed on these platforms exist because an agency or a funder commissioned them.

Scoring rules, in the plainest words they can be put

Three families of number, and the vocabulary is used inconsistently enough that it is worth pinning down here rather than on every card.

A proper scoring rule is one where the way to maximise your expected score is to state the probability you actually believe. This is not automatic. Score people on "the probability you gave the outcome that happened" and you will pay them to be more confident than they are, every time. The platforms here that rank you on a stated probability all use a proper rule, and the reason is exactly that.

Not every platform in this listing ranks you on a stated probability, though, and it is worth checking which kind you have landed on before reading a leaderboard. Two rank by trading profit in a free currency — one of those publishes the formula and the other awards undefined "expertise points" per topic on top of it. Another ranks by mean absolute error against a study's realised estimate, because what it elicits is an effect size rather than a probability. None of those is a proper scoring rule, and a rank under one of them does not mean what a rank under a proper rule means.

The Brier score is the squared error of a probabilistic forecast. Take your probability, take reality as 0 or 1, subtract, square, add up across the answer options. Forecast 70% on something that happens and your Brier score is (1 − 0.7)² + (0 − 0.3)² = 0.18; the best possible is 0 and the worst is 2. Halve that if a platform quotes 0.09 for the same forecast — some sum over both answer options of a yes/no question and some do not, and the two conventions differ by a factor of two on every number they produce. Lower is better, as the platforms that use it never tire of saying. Check which convention a published number uses before you compare it to anything: the form above sums both outcomes and tops out at 2, a common alternative scores one outcome and tops out at 1, and every figure differs by a factor of two between them. Platforms frequently publish a Brier score without saying which they mean.

Brier is intuitive, it is bounded, and its practical property is that it is forgiving of confident errors relative to the alternative.

The log score is the natural logarithm of the probability you gave the outcome that actually happened. It is always negative on a yes/no question — zero if you said 100% and were right, minus infinity if you said 0% and were wrong — which is why platforms that use it rescale it into something readable before showing it to anyone. Its practical property is the opposite of Brier's: it is brutal about confident errors. Moving from 99% to 99.9% gains you almost nothing when you are right and costs you a great deal when you are wrong. A platform's choice between Brier and log is a choice about how much it wants to punish overconfidence, and it changes which forecasters end up at the top.

A peer or relative score is not a third rule. It is one of the two above with the crowd subtracted from it: your score on a question, minus the average or median score of everyone else on the same question, usually weighted by how long your forecast was standing. Negative-is-better on a Brier-based relative score, positive-is-better on a log-based peer score, and the sign convention flipping between platforms is a real source of confusion.

Where the crowd term is the mean of everyone else, the result is zero-sum by construction: on a given question the participants' peer scores cancel out to zero. Where it is the median, or where the difference is scaled by how much each forecaster took part, they do not cancel and the leaderboard is relative without being zero-sum. Either way a peer score is a statement about the people you were standing next to and nothing else. Two forecasters with identical peer scores on two different platforms, or in two different tournaments on the same platform, have told you nothing about each other. Whenever you see a leaderboard in this category, the first question is which of these three things it is ranking.

Calibration, and why the accuracy board and the money board are different boards

Calibration is the property that your probabilities mean what they say. Of everything you called 70%, about 70% should happen. It is checked by bucketing a long run of forecasts and plotting the realised frequency against the stated probability — the closer the dots sit to the diagonal, the better calibrated the forecaster or the platform. Note what it does not require: a perfectly calibrated forecaster who says 50% to everything is useless and still perfectly calibrated. Calibration is a necessary property, not a sufficient one, and resolution in the technical sense — how far your forecasts move away from the base rate — is the other half.

This is where the two kinds of leaderboard part company, and it is the most common mistake made when comparing platforms across this category and the venues one.

An accuracy leaderboard ranks a proper score. It is invariant to how much you forecast in money terms, it is insensitive to liquidity, and on most platforms it is explicitly adjusted so that skipping a question costs you nothing and being early counts for more.

A profit leaderboard ranks money, or a play currency standing in for money. It rewards finding mispriced questions, trading size, being early in thin markets, and market-making. Those are real skills, and none of them is the same skill as stating well-calibrated probabilities. A trader can top a profit board by repeatedly taking 2% edges in questions whose outcomes they have no view on; a well-calibrated forecaster working on genuinely hard questions can sit in the middle of it. If a platform ranks by profit in a play currency, what you are reading is skill at that platform.

The practical test when you land on a platform is to find out which board is on its front page, and then whether it publishes anything about aggregate calibration at all — a calibration plot, an aggregate Brier score, a track-record page with the sample size and the date range on it. Some do and put a number on it. Some publish a record with no denominator.

"Real money" is mostly beside the point here, and it is a jurisdiction question anyway

On the venues side of this catalogue, whether a product settles in cash is the fact everything else hangs off, and whether a given reader can use it is a question about where they live. Neither carries over cleanly.

Most platforms in this category have no currency at all, and the ones that do have a currency you cannot get money back out of. Where cash appears it appears at the edges: a tournament prize pool divided among forecasters after the questions resolve, in proportion to their scores, and on one platform a periodic prize drawing with its own no-purchase rules and its own list of territories it is void in. Both are prizes, decided by rules the platform publishes, and neither has anything to do with the mechanics of holding a position.

So the usual argument — that money is what keeps a probability honest, and a number with nothing behind it is worth nothing — does not transfer as an argument about these platforms. What is behind the number here is a permanent public score under a name somebody has spent years building. Whether that is a weaker or a stronger incentive than money depends entirely on who is forecasting, and it is an empirical question the platforms themselves publish data about. Read the data rather than the slogan in either direction.

The one place "real money" does bite is history. At least one platform in this listing ran a real-money layer for a few months and closed it, and a great deal of third-party writing about it still describes that window as the present. Check the date on anything that tells you a forecasting platform pays out.

Resolution means four different things, and only the body of a card can hold them

Same word, four questions, and a card that answers only the first has not told you anything.

Who decides. Platform staff, the person who wrote the question, a panel, or — on one platform here — the traders themselves, by a vote weighted by the capital each of them put into the question, opened by any user once the owner lets a settlement go overdue. Staff resolution is slow and consistent; author resolution is fast and only as good as the author, and the usual recourse when it goes wrong is a report to a moderator rather than a rulebook.

Against what. A named source with a publication date, or credible open-source reporting assessed case by case. Both are defensible; they fail differently. A named source can fail to publish, publish late, or publish something obviously wrong, and the good platforms document what they do in each case.

On what timetable. Whether a question closes when the event happened or when it was reported — these differ, and one platform in this listing closes retroactively on the former, so only forecasts made before the real close are scored.

And what happens when the question turns out to be unanswerable. This is the one with no equivalent on a venue. A platform with no money at stake can annul a question and score nobody, and the better ones do. A venue cannot — money is sitting on the answer, so a resolution has to land somewhere. If you are choosing a platform to run a serious forecasting exercise on, its annulment policy tells you more about its editorial standards than its question count does.

What to check before you commit a season to one

Three things, in this order.

Can you get your data out. This varies more within this category than anything else, and it changes without an announcement. One platform's API now requires a token for every request and returns aggregate values on a few dozen questions out of thousands; another publishes an unauthenticated read API with a documented rate limit plus bulk dumps that have not been refreshed since 2024; a third publishes no API and withholds permission for automated extraction in its terms of service. If you are planning research on top of a platform, settle this before you start, and settle it against the platform's own API documentation and terms rather than against a blog post.

Whether the questions are the questions you care about. Question supply in this category is set by staff and by whoever commissioned a tournament, not by whatever is liquid. That is a feature if the commissioner cares about what you care about and a problem if not. The subject mix on these platforms skews to infectious disease, AI capability and policy, geopolitics, energy, macro indicators and elections, because that is what the funders of this work care about.

What the score buys you outside the platform. On one of these, an outstanding record is a route to paid professional forecasting work, and the selection criteria are published. On another it is a public track record and a share of a prize pool. On a third it is a number denominated in a currency that cannot leave the site. All three are legitimate; they are not interchangeable, and the difference is worth more to your decision than any feature comparison.

All 9 tools in Forecasting

Compiled from each vendor’s own documentation, pricing page and terms — no card here is marked hands-on yet.

Showing 9 of 9

Head to head

Background

How this part of the industry works, rather than which product to pick.

How to

One task per page, done with cards from this listing.

FAQ

What is the difference between a forecasting platform and a prediction market?

What you are there to do. On a venue you take a position and your reward is the payout. On a forecasting platform you state a probability and your reward is a score and a record. Almost nothing else on the two cards agrees — settlement, fees, who resolves, whether a regulator is involved.

Which scoring rule should I care about?

Whichever the platform ranks you by. Brier is the squared error of a forecast; the log score is the logarithm of the probability you gave the outcome that happened. Both are proper. A peer or relative score subtracts the crowd from whichever one is underneath it, which makes it a statement about one particular crowd and incomparable across platforms.

Does a platform being play-money make it less accurate?

Not by itself, and the question is empirical rather than obvious. The incentive that replaces money is a public, permanent score, which works on people who care about a record. What you should check instead is whether the platform publishes a calibration curve or an aggregate Brier score you can inspect, and how far back it goes.

Can I get the data out?

Ask before you invest a season in a platform, because the answers differ sharply and change without notice. Some publish an authenticated API with aggregate values restricted to a handful of questions, some publish an unauthenticated API plus bulk dumps, and at least one withholds permission for automated extraction in its terms of service.