What a forecasting track record is worth off the platform

Who reads a forecasting record earned on one site, why it rarely converts into anything elsewhere, and how expert panels recruit from public leaderboards.

Mostly, it is worth something inside the house that scored it. Good Judgment recruits Superforecasters from Good Judgment Open and Metaculus picks Pro Forecasters from its own leaderboards, both reading data they computed themselves. Anywhere else a record travels as a claim, a rank or a percentile beside a handle, which the reader cannot recompute unless you kept the forecasts, the questions and the formula yourself.

A season of forecasting leaves you with a number on a profile page. The number is real, and the work behind it was real. What is less obvious is that almost everything that gives the number its meaning — the forecasts themselves, the questions they were made on, the dates, the crowd you were measured against, the formula — stays on the site that computed it. What leaves with you is a handle and a figure, and a figure with nothing behind it is something a stranger has to take on trust.

This page is about that gap: who actually reads a forecasting record, where they read it, and what it takes for a record to mean something to somebody other than the platform that holds it. What a score measures in the first place is its own page, and so is how accurate a market price is; neither is repeated here.

How it works

A track record has three parts, and the platform holds all three:

  1. The forecasts — every probability you entered, with the day you entered it and every update after. Where scoring is time-averaged, the updates are most of the record, not a footnote to it.
  2. The corpus — which questions were asked, who wrote them, and which of them resolved rather than being voided or left open. Two people with the same score on different corpora have done different things.
  3. The formula — the rule, the convention, the crowd term and the aggregation, each of which moves the number without moving your judgement.

A figure on a profile is the output of all three. Take it off the site and you have the output without the inputs, which is exactly the form in which a claim cannot be checked.

What each platform lets you take with you

Whether the inputs can leave is a product decision, and the platforms in this category made very different ones. From our own cards:

PlatformWhat leavesWhat that means for a record
FatebookOne-click CSV of every prediction you made, plus an APIThe whole record is yours, in a file
ManifoldA public API with keyless read endpoints; bulk JSON dumps that the data page says were last refreshed in July 2024Your trades are retrievable; the dumps are not current
MetaculusAn API that requires an account token; with an ordinary account it returns your own dataYour own forecasts, yes; the crowd you were scored against, mostly not
Good Judgment OpenNothing — no API, and the terms do not permit automated extractionThe record exists only as the profile page
Quantified IntuitionsNo API, no exportPractice that leaves no portable trace

Two closures in this catalogue show what a record without an export costs when a platform stops. INFER's archive kept final standings and crowd trajectories but not individual forecast histories, with no export button (the INFER page). PredictionBook now answers every path with the same retirement notice, and a record nobody exported before February 2025 survives only where the Internet Archive happened to crawl a public page (the PredictionBook page). A record that lives only on a profile page lasts exactly as long as the profile page.

Who reads a record, and where they read it

There are three kinds of reader who act on a forecasting record, and they differ most in one respect: whether they can see the inputs.

Good Judgment reads its own platform

Good Judgment's professional Superforecasters descend from a research programme. The published account of it describes the winning strategy in a US intelligence-community forecasting tournament as culling the top performers each year and putting them into elite teams. Good Judgment's own recruiting page adds the thresholds: the original superforecasters were taken from the top 2% of forecasters in each experimental condition, and only from those who had answered at least 50 questions in a tournament season.

After the tournament ended in mid-2015, the same page says, new professional Superforecasters have come from Good Judgment Open, the firm's free public platform. The stated route, as captured on 12 September 2026:

  • forecast on at least 100 GJ Open questions, cumulatively rather than in one year;
  • the metric looked at most closely is average accuracy score per closed question, among "many metrics";
  • comment quality and collegiality are assessed as well, because clients read the rationales and Superforecasters work in teams;
  • recruiting happens each autumn, and a candidate who completes a three-month probation becomes a full Superforecaster.

Read that list for what it depends on. Every input is something Good Judgment computed on its own servers: the per-question scores, the comment history, compliance with the site's terms. The GJ Open FAQ adds the detail that makes this work — a forecast cannot be withdrawn or deleted, so that nobody can tidy a record once a question turns against them. The record is trustworthy to the firm because the firm watched it being made. The title that results, which Good Judgment prints with a registered-trademark sign, is issued by one company and describes a standing with that company.

Metaculus reads its own leaderboards, and says it will look elsewhere

Metaculus hires Pro Forecasters for client projects, and its FAQ gives the selection method as a weighted combination of its Peer Accuracy, Baseline Accuracy and Comments leaderboards across several leaderboard periods, with Peer weighted highest and good scores needed in every category. The floor is at least 75 resolved questions across multiple subject areas and at least a year of forecasting. The same answer says Metaculus primarily recruits from its own community but will consider forecasters who have demonstrated excellent ability elsewhere.

That last sentence is the only published door we found in this category through which an outside record walks in, and it does not say how an outside record is weighed. The inside record is weighed by a formula Metaculus publishes; the outside one, by whatever the applicant can show.

A closed panel reads other people's platforms

The third reader recruits on records it did not compute. The Swift Centre, a London-based forecasting organisation with a not-for-profit research wing and an advisory company, describes its panel on its research page as drawn from the top 1% of forecasters worldwide, "people with a measured, public record of accuracy, not credentials alone". Neither that page, the advisory page nor the archive articles we read name the source of the percentile: which platform, which leaderboard, which period, out of how many.

That is not a peculiarity of one organisation. It is what a record looks like once it has left the venue that computed it. A percentile quoted off-venue has to be taken on trust by the client, because the client cannot see the corpus, the formula or the crowd, and a percentile with none of those attached could have been earned on almost any denominator. From the outside, a closed panel's composition is a description rather than a dataset.

Why a record needs a sample before anyone trusts it

Both readers that publish their method set a minimum count before looking at accuracy at all: 50 questions in a season for the original superforecasters, 100 on GJ Open for candidates since, 75 resolved questions and a year on Metaculus. The floors exist because a short record carries a lot of luck, and the research behind the superforecaster idea says so in its own terms.

The paper's authors describe the expectation they had to beat: that forecasters picked for a strong first year would regress toward the mean, because chance could be the dominant driver of any one year's results. For the group selected as superforecasters that did not happen — their accuracy held and improved in the second and third years. For the comparison group of strong forecasters who fell just short of the cut, and for everyone else, it did: both got worse. The authors call the superforecaster result extremely unlikely to be a lucky accident, and they reached that conclusion with three years of data, several accuracy measures and hundreds of questions — not with one leaderboard.

Two further findings from the same paper change what a record means once it leaves:

  • Forecasters chose their own questions. Because self-selection could reward somebody for picking easy questions, the researchers standardised Brier scores within each question before averaging. A public leaderboard where everyone answers a different subset is making the same comparison without that correction, unless its formula subtracts a crowd — which brings its own problem, covered on the scoring page.
  • Part of the accuracy came from the environment. Putting top performers together in elite teams improved their accuracy more than either training or ordinary teaming did, and the paper concludes that superforecasters are partly discovered and partly created. A record earned inside a team, with teammates' rationales in view, is not only a record of the individual.

So the floors in the recruiting rules are not paperwork. A figure from forty questions in one subject area is a number with a wide error bar attached, whether or not anybody prints the bar.

Why two records from two sites do not add up

The arithmetic of why a Metaculus Peer score and a GJ Open Relative Brier Score cannot be compared is on the scoring page — different rules, different crowd statistics, opposite signs. The consequence for a record is blunter: records do not combine. Sixty questions on one site and sixty on another are not a 120-question record. They are two 60-question records, each under its own formula, on a corpus somebody else chose, and each below at least one of the recruiting floors above.

Three things go wrong when people try to combine them anyway:

  • Averaging the numbers. The two numbers are in different units. An average of a log-based score and a Brier-based one is not a score in either.
  • Averaging the ranks. A top-5% finish in a 300-person tournament and a top-5% finish on a site-wide leaderboard of thousands are statements about different crowds, and the crowd is half of what a relative rank measures.
  • Quoting the best one. A forecaster active on three sites who quotes the site where they ranked highest has selected on the outcome — the same move as quoting only the questions that went well, one level up.

The one honest form of a multi-site record is a list, not a total: each site, each formula, each question count, each date range, side by side, with the reader left to weigh them.

What a panel's own record is evidence of

A closed panel is recruited on individual records earned elsewhere, and then sells a collective record of its own. That second record is subject to every rule on this page, and the Swift Centre's published write-ups are useful precisely because they show their working.

The year-end score. In its 2025 review, dated 18 December 2025, the Swift Centre reports an aggregate Brier score of 0.115 across the 13 public likelihood forecasts that resolved in 2025, and states in the same sentence that the score is not difficulty-adjusted. It also reports that two of the thirteen account for most of the error, and that the score would have been 0.040 had those two gone the other way. Both disclosures are to its credit, and both describe the size of the sample: in a set of thirteen, two questions move the headline figure by a factor of almost three.

The same page supplies a reference scale — 0.25 for guessing 50/50 on everything, 0.22 to 0.25 for untrained people, 0.17 to 0.20 for frontier AI models, below 0.15 for superforecasters described as the top 2% — without naming the study each band came from. Two things are worth checking before any figure is read against a scale like that:

  • The convention. A scale on which a coin flip scores 0.25 is the one-outcome Brier. Good Judgment Open publishes Brier on the convention whose worst case is 2, where the same coin flip scores 0.5. A benchmark carried from one convention to the other is wrong by a factor of two.
  • The corpus. A band describing forecasters on one tournament's questions says nothing about thirteen questions chosen by somebody else. The paper above had to standardise within each question to compare people on different questions inside a single tournament.

The comparison with a market. A May 2026 write-up sets the panel's forecasts against Polymarket prices on five markets chosen in late January 2026, when the panel had identified several that appeared, in its words, fundamentally mispriced. Four are reported as the return on a hypothetical position; the fifth is recorded as no position, because the panel judged that price efficient. The same page then picks five more markets it considers inaccurate. Read it for its structure: the questions were chosen by one side of the comparison, mostly because that side disagreed with the price; the unit is a simulated profit rather than a score; and the sample is five. A head-to-head built that way can show that a disagreement existed and which way it resolved. It cannot show how the panel would have done on the markets it agreed with or did not pick, which is where a comparison of accuracy lives. What a market price is evidence of in the first place is taken apart on the price page.

None of this is an argument that the panel is inaccurate. It is an argument that a panel's own record is a short record, on a hand-picked corpus, reported by the panel — the same three properties that make an individual's off-venue record hard to read.

What it costs

No platform in this category charges for a record. What a record costs is time, and the units are set by the reader you want to convince rather than by you.

  • Question count, before accuracy is looked at. 100 closed GJ Open questions for a Good Judgment candidate; 75 resolved questions, several subject areas and a year of forecasting for a Metaculus Pro. On questions that run for months, the count is the long part.
  • Written rationales. Both published selection methods score comments alongside accuracy. Forecasts entered without reasoning are a record that meets half of each stated test.
  • Staying on one site. Every threshold above counts questions on the reader's own platform. Effort spread across three sites builds three records, each short of the floor.
  • Recency. Metaculus's FAQ computes medal ranks from points — 10 for gold, 4 for silver, 1 for bronze — that decay by 2× per year. A strong season ages out of the current rank on a profile whether or not the forecaster has stopped; only the separate best-ever rank keeps it.
  • Permanence, in both directions. GJ Open does not allow a forecast to be withdrawn or deleted, so a bad stretch stays in the record. And when a platform closes, the record can go with it, as INFER's and PredictionBook's users found.

What you can do about it

Decide who the record is for, then build it where they read. If the reader is Good Judgment, the record that counts is on Good Judgment Open, and the stated test is 100 questions, per-question accuracy, and comments. If it is Metaculus, it is Metaculus's own Peer, Baseline and Comments leaderboards. A record built anywhere else is, to either of them, an outside record, and only Metaculus says it will consider one.

Keep your own copy of every forecast. Not a screenshot of the profile — the forecasts, with dates and the question text. Fatebook exports everything you entered there as CSV; Metaculus returns your own data through its API with an account token; Manifold exposes your trades through a public API. Where a platform offers no export, as GJ Open does not, write the forecasts down as you make them. A record you hold is one you can hand to somebody to recompute. A record you do not hold is one you can only describe.

Quote a record with its denominator. Site, formula, question count, date range and — for a relative score — the crowd. "Top 1%" or "Superforecaster-level" with none of those attached is the off-venue form of a record, and a careful reader discounts it for exactly that reason.

Use one handle everywhere. A stable username across sites is what lets a stranger find all of your records rather than the one you chose to mention, and it is the only thing that connects a record in a platform's archive to you after the platform has gone.

When you are the reader, ask for the inputs. For a panel selling forecasts, or an individual offering a record: which platform, which formula, how many questions, which were excluded, and whether the questions were chosen before or after the forecaster knew their view. A panel that publishes its misses and its sample size, as the Swift Centre's 2025 review does, has given you something to work with; a percentile with no platform attached has given you a description.

What you can do about it

Tools this bears on

Cards in the catalogue where what is above changes the decision.

  • Good Judgment Open

    Brier scores against the crowd, and the recruiting ground for Superforecasters.

    FreeFree tier

  • Metaculus

    Proper scoring and public track records on questions nobody can take a position in.

    FreeFree tierOpen source

  • Fatebook

    Write down what you think will happen, in Slack or a browser, and get scored on it.

    FreeFree tierOpen source

  • Manifold

    Anyone can open a question, anyone can take a side, and the currency buys nothing.

    $5/moFree tierOpen source

FAQ

Can I take my forecasting score from one platform to another?

The number, yes, as a claim you quote. The record behind it usually not. The forecasts, the questions, the crowd and the formula stay on the site that computed them, and some sites offer no export at all. A stranger reading the number has nothing to recompute it from unless you kept the forecasts yourself.

How does Good Judgment choose Superforecasters?

From its own public platform. Its recruiting page asks for at least 100 Good Judgment Open questions, looks most closely at average accuracy per closed question, weighs comment quality and collegiality, recruits each autumn and ends with a three-month probation. The original cohort came from the top 2% of a government-sponsored research tournament.

Does Metaculus accept a track record from another site?

Its FAQ says it primarily recruits Pro Forecasters from its own leaderboards, requiring at least 75 resolved questions across several subjects and a year of forecasting, and that forecasters with excellent records elsewhere may be considered as well. It does not say how an outside record is weighed.

Can I add up my records from several platforms?

No. The scores use different rules and conventions, the ranks are against different crowds, and the questions were chosen by different people. Sixty questions on each of two sites are two short records, not one long one. List them side by side with the formula and the question count for each.

What should I ask a forecasting panel that says it recruits the top 1%?

Top 1% of what. Which platform, which leaderboard, which period, out of how many forecasters, and how many questions each member has resolved. Then ask for the panel's own record with its sample size, its misses, and whether its questions were chosen before or after the panel knew its view.

Sources

  1. Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions — Perspectives on Psychological Science, volume 10, issue 3 (Mellers, Tetlock and others), . The published account of how the original superforecasters were selected and tracked; a study of a finished tournament, not superseded.
  2. How to Become a Superforecaster (Internet Archive capture of 12 September 2026) — Good Judgment, read
  3. FAQ — Good Judgment Open, read
  4. Metaculus FAQ, "What are Metaculus Pro Forecasters?" and "What are medal ranks?" (Internet Archive capture of 6 September 2026) — Metaculus, read
  5. Swift Centre Research — Swift Centre, read
  6. 2025 - A Year in Review — Swift Centre,
  7. Polymarket vs Swift Centre: Round 2 — Swift Centre,

The catalogue next door

This page is background, not a listing. The products it bears on are in Forecasting Platforms, each filled in against the same schema, with the fields to narrow it yourself.

Last updated . Corrected in place: this is a reference page, not a dated post.