Social Science Prediction Platform
Forecast what a study will find, before it finds it, and be scored against the estimate.
by BITSS
Last updated
What it is
A platform for forecasting what a study will find. A researcher with an experiment already designed but not yet run posts it here, forecasters predict the effect size, and when the results come in the forecasts are compared with the real estimate. The purpose is stated plainly on the homepage: to enable the systematic collection and assessment of expert forecasts of the effects of untested social programs.
It is run by the Berkeley Initiative for Transparency in the Social Sciences, part of the Center for Effective Global Action at UC Berkeley, and led by Stefano DellaVigna at Berkeley and Eva Vivalt at Toronto, with funding from the Alfred P. Sloan Foundation and an anonymous foundation. The advisory board is a list of names a reader of this category will recognise, Philip Tetlock among them.
What you submit is not a probability. Each study is a Qualtrics survey written by the researchers and uploaded to the platform, so the elicitation varies: typically a point estimate of an average treatment effect, given in standard deviations, percentage points or percent, on a slider or in a text box. Probabilities appear only where a particular study asks for one. Four questions are appended to the end of each survey by default, the first being how confident you are in your own predictions for that study.
Scoring is error, not a scoring rule. The public leaderboard ranks by mean absolute error — the average absolute distance between prediction and truth, each prediction standardised by the baseline standard deviation, lower being better — with a relative error column that compares you against the mean forecast on the same question, negative being better than average. To be ranked you need at least ten numeric forecasts with reported results across at least five studies. Names are anonymised unless you opt out, and half the top ten currently read "Anonymous". There is no Brier score, no log score and no calibration chart anywhere on the site, and given what is elicited there could not be.
Availability
Open to anyone worldwide who is 18 or over. You register with an email address, a Google account or an ORCiD; there is no institutional email requirement, no approval step for an account, and no identity check. The background questionnaire explicitly branches for people who are not graduate students, faculty, post-docs or non-academic researchers. Some surveys are posted as public links that need no account at all.
Two things are gated, and neither is the account. Surveys are reviewed by platform staff for completeness before they go live, and elicitors must affirm that they hold their own institution's IRB approval. And money is gated hard — see below.
The terms are governed by California state and US federal law, with UC Berkeley as the responsible institution and its privacy office named for GDPR requests. The terms page is dated March 2024.
Pricing
Free, in the FAQ's own words, for everyone — both taking surveys and posting them.
Money flows the other way for a minority of users, and the conditions are worth reading before counting on it. The Forecasting Panel pays USD 200 after a first six-month semester and a further USD 200 after the second, with a USD 300 bonus for panelists who complete twelve months and finish in the top half for accuracy. It asks for at least 80% of the platform's surveys each semester, roughly 15 to 25 projects, and the incentives are restricted to graduate students, faculty and researchers at other institutions — the general public may forecast but is not eligible. Payment runs by bank transfer through the Berkeley Existential Risk Initiative, and the first time you are due a payment the platform may verify that you are a researcher by looking for a .edu address, a RePEc, ORCID, Google Scholar or ResearchGate profile, a university page or a peer-reviewed paper. Larger amounts may require a W-9 or W-8BEN.
Individual researchers may attach their own incentives to a survey, but the platform says it does not provide a mechanism to distribute those payments — that is arranged by email between the researcher and the forecaster.
Markets & resolution
The question set is whatever researchers have posted: field experiments and surveys in development economics, behavioural science, labour, public policy and firm productivity, run in 54 countries. The public archive lists 158 studies, the homepage counts 170 studies, 82,754 predictions and 400 active forecasters, and the newest study visible in the archive on 19 September 2026 carried a July 2026 date.
Resolution has no adjudicator, because the answer is the study's own reported estimate. That gives the platform a property no venue and no other forecasting site has: the truth arrives on the researcher's timetable, not the question's. The platform's own analysis of its first 100 projects, posted between 2020 and 2024, had results for 66 of them. A forecast made here may sit unscored for years, and some studies never report at all.
That analysis is the reason to take the platform seriously and is worth reading before signing up. DellaVigna and Vivalt, NBER working paper 34493, dated November 2025 and revised April 2026, covers 53,298 forecasts across those 100 projects. Its findings: forecasters on average over-estimate treatment effects, but the average forecast is quite predictive of the actual effect; academics are more accurate than non-academics while subject-matter expertise is not associated with accuracy; a panel of motivated repeat forecasters does better, though repeat forecasting alone does not; and confidence in one's own forecasts is perversely associated with lower accuracy.
Integrations
None. There is no API, no export of forecasts, and the platform code is not public — the FAQ says the team is preparing both a clean downloadable dataset and the code for release, and puts the dataset two to three years out. What is downloadable today is the instrument rather than the data: each archived study offers its Qualtrics .qsf file, along with metadata, abstract and citation.
No licence is stated for any of it. The only reuse clause in the terms is an anti-scooping one: you agree to credit project authors for major ideas taken off the platform.
Limitations
The scoring is thin for a platform whose entire premise is measured accuracy. The leaderboard is marked BETA, shows only ten names, and as read on 19 September 2026 was stamped as of 25 February 2026 — seven months stale, against a stated intention to update it every two months. Whether a signed-in forecaster gets a personal accuracy page is not visible from outside, and there is no public per-user profile.
The numbers on the site do not reconcile with each other. The homepage says over 2,400 forecasters from 340-plus institutions; the platform's own August 2026 newsletter describes more than 4,700 forecasters on the first 100 projects alone. Nothing on the site explains what each figure counts.
Feedback is slow by construction, which is the opposite of what a forecaster training calibration needs: you will wait years to learn whether you were right, and on a third of projects you never will. And the open-survey list is behind the login, so from outside you cannot tell what is currently answerable.
Finally, an error score is not a calibration record. A good rank here says your effect-size estimates land close to realised estimates on social-science experiments. It does not transfer to probability forecasting, and nobody should read it as though it did.
Alternatives
Metaculus for probability forecasting on a proper scoring rule, with tighter feedback and a record a research audience already recognises. Good Judgment Open for a Brier score against a crowd on geopolitical questions. Fatebook if what you want is to record your own predictions about your own work and get a calibration chart out of it.
Specs
- Interfaces
- none
- Export
- None
- Available in
- Global
- KYC required
- No
- Market subjects
- Science, Macro, Business
- Resolved by
- The study's own reported estimate. A forecast is scored once the researchers run the experiment and report the result, which can take years, and some studies never report.
- Maker fee
- None
- Platforms
- Web
- AI features
- None
- Capabilities
- Calibration scoring
- Pricing verified
- Availability verified
- Markets verified
- Capabilities verified
Background
How this part of the sector works, rather than which product to pick.
- Scientific and government forecast hubs — Public-health agencies run open, scored forecast hubs anyone can join. The unit of submission is a model, not a person, and nothing there is a personal record.
- How accurate a prediction market's price actually is — What research establishes about reading 73 cents as a 73% chance — where market prices are calibrated, the deviations that recur, and what the horizon does.
- What a forecasting platform's score actually measures — Proper rules, the two Brier conventions that differ by a factor of two, crowd-relative scores that do not travel between sites, and why profit is not accuracy.
Also worth comparing
- Manifold — Anyone can open a question, anyone can take a side, and the currency buys nothing.
- Metaculus — Proper scoring and public track records on questions nobody can take a position in.
- Prediki — A wiki with a market attached — anyone writes the question, edits it, votes the result.
- Confido — Open-source forecasting workspace you host yourself, with a Brier score pointing upwards.
- Fatebook — Write down what you think will happen, in Slack or a browser, and get scored on it.
- Good Judgment Open — Brier scores against the crowd, and the recruiting ground for Superforecasters.
FAQ
Do you have to be an academic to use the platform?
No. Anyone 18 or over can create an account with an email address, a Google account or an ORCiD and forecast for free, and the background questionnaire has a branch for people who do not work in research. The paid forecasting panel is different — its incentives are restricted to graduate students, faculty and researchers.
What do you actually forecast here?
Not a probability. Each study is a Qualtrics survey written by the researchers, and you are usually asked for a point estimate of a treatment effect — in standard deviations, percentage points or percent — entered on a slider or in a text box, plus how confident you are.
How are forecasters scored?
By mean absolute error against the realised estimate, standardised by the baseline standard deviation, with a relative error measure alongside it that compares you with the mean forecast on the same question. There is no Brier score and no calibration chart, because what is elicited is an effect size rather than a probability.
Can you download the forecasts?
Not yet. The FAQ says the collected data is not currently public and that the team aims to release a clean downloadable dataset within two to three years. The survey instruments themselves are downloadable per study as Qualtrics .qsf files.