Metaculus vs Good Judgment Open vs Manifold: which score you are chasing

Three free platforms, three formulas — a log-based peer score, a relative Brier against one crowd, and profit in a currency that cannot leave the site.

The choice is the scoring rule, and nothing else is close

All three are free, all three are open worldwide with no identity check, all three settle in something that is not money, and all three cards set calibration_scoring: true. None of that separates them.

What separates them is the formula each one ranks you by. On a platform whose entire product is a record, the formula is the product: it decides what gets rewarded, and after a season of chasing it, what you will have got good at. Read the three formulas first and the feature lists second.

Three formulas, and what each one pays for

Metaculus rests everything on the log score, which is proper — the way to maximise it on average is to state the probability you actually believe. Two rescalings sit on top because a raw log score is always negative and reads badly. The Baseline score compares you to a fixed chance benchmark, so it can be positive for everyone on an easy question. The Peer score compares you to everyone else on the same question, and the participants' Peer scores sum to zero by construction, which makes it the harsher and more informative of the two. Both are time-averaged over the life of the question, so a forecast left standing earns for every day it stood.

Good Judgment Open ranks you by the Relative Brier Score: your average daily Brier score, minus the crowd's average daily median on the same question, multiplied by your Participation Rate. Lower is better and negative beats the crowd. A question you skipped scores zero rather than penalising you. Two things follow from the formula that do not follow from Metaculus's: because it subtracts a median and scales by participation, it does not cancel out to zero across a question — the leaderboard is relative without being zero-sum. And because Brier is a squared error bounded at 2, it treats a confident miss more gently than a log score does.

Manifold ranks the headline number on a profile by mana profit. That is not a scoring rule; it is what is left after taking positions against an automated market maker and the limit orders resting on it. Manifold does publish a platform-wide calibration figure — a Brier score of 0.17375 across all resolved markets on its calibration page on 19 September 2026 — but that is a claim about the aggregate, not about the ranking a profile displays.

If you maximise this number for a season, what will you have learned

This is the question worth asking before signing up, and the three answers are genuinely different.

Metaculus. State the number you believe, state it early, and update it as things change. The time-averaging pays for standing correct through the boring middle of a question rather than arriving late with the answer, and the log score underneath means you cannot buy rank with confidence you do not have. What you come away with is calibration, measured under the rule that is least tolerant of bravado.

Good Judgment Open. Be better than the specific people in your Challenge, on the question flow Good Judgment and its Challenge sponsors chose — geopolitics, elections, macro, armed conflict. Participation Rate multiplies your score, so breadth and early entry matter as much as being right, and you cannot withdraw or delete a forecast, by policy, so the record cannot be tidied afterwards. What you come away with is skill at beating one particular crowd — a crowd that, as the card notes, includes a cohort trying to get hired.

Manifold. Find questions priced wrong, be early in thin ones, subsidise liquidity, size up. Those are real skills and none of them is the skill of stating a well-calibrated probability. The card is blunt about it: someone trading many small mispricings in thin markets and someone who is well calibrated on hard questions end up in very different places on a mana leaderboard, and the leaderboard cannot tell you which you are looking at. And the reward for correcting somebody else's price is mana, bought at a flat 100 per US dollar and, by Manifold's own terms, not cashable out, sellable, transferable or exchangeable for anything off the site. The familiar argument that a price disciplines itself because being right pays arrives here paid in a consumable that only buys more Manifold.

Who resolves the question you were scored on

A score is only as meaningful as the resolution behind it, and these three hand that job to three different parties.

Metaculus staff resolve against the criteria published on the question, and annul a question whose criterion turns out not to be evaluable — annulled questions score nobody. Good Judgment resolves against credible open-source reporting unless the question names a source, and reserves a review for plain error in a named source. Its timetable carries a trap worth knowing before you plan a week around a question: questions close on when the event occurred rather than when it was reported, so retroactive closing dates are normal and only forecasts made through the calendar day before the official close are scored.

On Manifold, the creator resolves their own question. That is a design choice rather than a defect, and it is what makes thirty-second question creation and tens of thousands of questions possible. But it changes what your number means: your mana profit is partly a verdict on how well a series of strangers read resolution criteria they wrote themselves. Moderators step in only in exceptional circumstances, resolving to N/A is restricted to moderators and staff, and the card's own advice is to check a creator's track record before taking a side. Neither of the other two ever asks that of you.

Two of these three do not make a price at all

Six of the nine cards in this category carry economics.liquidity_model: none — they take a stated probability, score it, and there is nothing to buy or sell at any point. Metaculus and Good Judgment Open are both in that six. Manifold is one of the three that do make a price, through an automated market maker with limit orders on top, and it is the only one of these three with an order book, live trading, or a second home in the venues category.

That split decides the choice more than any feature does.

If you want something to act on — take a side, watch the number move, run a bot against it — only Manifold gives you an object with a price, and only Manifold gives you the access to work against it: unauthenticated read endpoints, 500 requests per minute per IP, a websocket, bots explicitly permitted. Metaculus now requires an account token on every request and returns aggregate values on roughly 50 questions to an ordinary account, about 250 open and 250 resolved to an approved bot account. Good Judgment Open publishes no API at all and its terms withhold permission for collection, aggregation or automated extraction.

If you want to be measured, a price is in the way, because it pays for trading rather than for being right, and the two come apart exactly where the questions get hard.

The published numbers, and the line they do not cross

Two of the three publish an aggregate, and they should not be lined up against each other. Manifold read 0.17375 on 19 September 2026, across all resolved markets. The empirical Baseline and Peer figures Metaculus publishes — a median Baseline of about +17 on resolved binary questions — are marked as of November 2023, so the benchmark you would compare your own score against is three years old. Different rules, different corpora, different vintages. Neither card records which Brier convention its number uses, and the two conventions in circulation differ by a factor of two on every figure they produce. Good Judgment Open publishes no platform-wide number at all.

The limit that applies to all three is the same one. A platform aggregate describes what a crowd did across thousands of resolved questions. It promises nothing about the question open in front of you — which may be thin, on a subject this crowd is weak at, and resolved by a party you have not checked. Metaculus's card makes the same point about its own leaderboard: it is a record about forecasters, not a claim that the aggregate is right in any particular case.

Which one to take, and when to move

Take Metaculus if you want a record that reads as calibration to somebody outside the site. It is the only one of the three with a proper rule, staff resolution, an annulment policy, an open codebase you can run yourself, and any programmatic access at all under your own account. Move on when you need aggregate data in bulk without filing an application, or when the questions you care about are simply not listed — question supply here follows what research funders fund.

Take Good Judgment Open if the score is a job application. It is the only one of the three where an outstanding record has a published route out of the platform: recruitment each autumn from forecasters with at least 100 answered questions, weighing average accuracy per closed question most heavily, then comment quality and collegiality, then a three-month probation. Accept in exchange that it is the worst of the three on every other axis — no API, no automation, no export, no data permission — and that its number is meaningless outside the Challenge it was earned in. Move the moment your reason for being there becomes data rather than a record.

Take Manifold for volume, speed and your own question listed in thirty seconds. It is the only one of the three where you can ask something nobody has asked, get a price on it the same day, and build against it without asking permission. Treat the profile number as skill at Manifold, not as accuracy, and treat the perpetuals as the separate product they are: oracle-priced, leveraged, liquidatable, the one thing on the platform that charges a fee, and outside the calibration argument entirely. Move to Metaculus the moment you want your number read as accuracy by anyone who does not already use Manifold.

If you cannot decide, the honest split is Metaculus for the record and Manifold for repetition — but put the volume where the formula is the one you want to be judged by, because three short histories under three incomparable formulas add up to no track record at all.

FAQ

Which of these three scores hardest on an overconfident forecast?

Metaculus. Its Baseline and Peer scores are rescalings of the log score, which punishes a confident error far more than a squared-error score does. Good Judgment Open's Relative Brier Score is also a proper rule underneath, but Brier is bounded at 2 and is the more forgiving of the two on exactly that mistake. Manifold's headline profile number is mana profit, which is not a scoring rule at all.

Can I compare my Metaculus Peer score with my Relative Brier Score on GJ Open?

No, and the Good Judgment Open card says so directly. Both subtract a crowd, but one is built on the log score and the other on Brier, the sign conventions run in opposite directions, and Metaculus Peer scores cancel to zero across a question while a Relative Brier Score does not. Each number describes the people you stood next to, on that platform, on those questions.

Does Manifold's published Brier score mean I can trust the price on the question I am looking at?

It does not promise that. The 0.17375 on Manifold's calibration page on 19 September 2026 is an aggregate across all resolved markets. The question in front of you may be thinly traded, on a subject the crowd is bad at, and it will be resolved by whoever wrote it. The aggregate is evidence about the platform, not about your question.

Which of the three can I get data out of?

Manifold, comfortably — unauthenticated read endpoints, a documented 500 requests per minute per IP, a websocket, bots explicitly permitted, with bulk dumps that have not been refreshed since July 2024. Metaculus needs an account token on every request and returns aggregate values on roughly 50 questions to an ordinary account. Good Judgment Open publishes no API and its terms withhold permission for automated extraction.

Do I have to choose one?

Only if you want the record to mean something. Forecasts spread across three platforms produce three short track records under three incomparable formulas. Pick the formula you want to be measured by, put the volume there, and use the others for practice.