Peer score
Also written Metaculus Peer score
A score for a forecast measured against the other forecasts on the same question rather than against the outcome alone. On Metaculus, where the name comes from, it is the average gap between your log score and everyone else's, multiplied by 100 and averaged over the life of the question, and all participants' scores on a question add up to zero. Elsewhere the phrase is used loosely for any crowd-relative score, which is where comparisons go wrong.
The forecasting category page uses "peer score" as a family name for scores that subtract a crowd, and that is fair usage. This page is about the one number that carries the name as a product — Metaculus's Peer score — because it has properties the rest of the family does not, and a reader who assumes the family shares them will misread every leaderboard in the category.
How it works
Take one moment in a question's life. Every forecaster with a prediction standing has a log score against the eventual outcome: the natural log of the probability they gave to what happened. Your Peer score at that moment is the average of the differences between your log score and each other forecaster's, multiplied by 100. Metaculus's scoring code computes the same thing a shorter way — your log score minus the log of the geometric mean of all standing forecasts, yours included, scaled by N/(N − 1) — and the two are algebraically identical. On a continuous question the result is halved.
Then the moments are averaged. Scores are time-averaged from the question's open date to its scheduled close, and three rules in the scores FAQ decide what goes into that average:
- Before your first forecast you score 0. Not "not counted" — zero, averaged in with everything else. The fraction of the question's life you had a prediction standing is published separately as your coverage.
- A forecast stands until you replace it. A day you did not log in is scored on whatever you said last.
- Early resolution pads with zeros. If a question resolves before its scheduled close, the stretch in between counts at 0 for everyone.
Two more details come out of the code rather than the FAQ. When only one person has a forecast standing, the scaling factor is set to zero, so a lone forecaster scores nothing, right or wrong. And the spot Peer score skips the averaging entirely and reads one moment, by default the moment the Community Prediction is revealed.
Because each forecaster's score is an average of differences with everybody else, adding up every participant's Peer score on a question gives exactly zero. The FAQ lists the extremes Metaculus's probability limits allow — +691 and −691 on a binary question — and, as of November 2023, a median observed Peer score of +2. The Community Prediction's own Peer score is positive, and the FAQ explains why: an aggregate is less noisy than most of the people it is built from.
What else is called a peer score, and is not this
Good Judgment Open and Fatebook both publish a relative Brier score: yours minus a crowd median, computed on the Brier scale, where lower is better. A median is not an average, so those scores do not cancel to zero, and the arrow points the other way. Metaculus's own older Relative score was also median-based and on the log scale; the FAQ describes it as being replaced by the Peer score from late 2023, while still in use in many open tournaments.
Why it matters here
A positive Peer score says you beat the room, not that you were right. On a question where everyone was confidently wrong, the least wrong forecaster has a positive score. On an easy question where everyone was right, a forecaster who was merely right scores about zero. The number is information about the company you kept, which is why Metaculus publishes it beside a Baseline score that compares you to chance instead. Read the two together; either alone misleads.
Two Peer scores from different rooms do not compare. A +15 earned in a tournament of specialists and a +15 earned on a question most people forecast carelessly are the same number about different crowds. The same applies across platforms, with an extra problem: a Metaculus Peer score is log-based and higher-is-better, a Good Judgment Open relative Brier is Brier-based and lower-is-better, and there is no conversion between them because the two rules punish a confident miss differently. Which score you are chasing puts Metaculus, Good Judgment Open and Manifold side by side.
Late entry is not free. Because the stretch before your first forecast scores zero rather than being skipped, joining a question halfway caps what you can earn on it, and joining at the last minute to copy a settled crowd earns close to nothing. That is a design choice — it removes the reward for waiting — and it means a per-question Peer score and a per-question log score describe different things even for the same forecast.
A tournament prize is not a linear function of it. Metaculus ranks tournament entrants by the question-weighted sum of Peer scores and divides the pool in proportion to that sum squared, with nothing paid on a negative sum. The FAQ calls this essentially proper only for a sufficiently large number of independent questions; on a short tournament, what maximises the expected prize is not always your honest forecast. The proper scoring rule page works through why.
A zero-sum score cannot tell you the crowd improved. If every forecaster on the site got better, every Peer score would still sum to zero per question. Progress over time has to be read off an absolute score — log or Brier — on comparable questions, which is what tracking your own forecast record sets up.
Where you will meet this
Cards in the catalogue whose own text uses the term.
Sources
- Scores FAQ (snapshot of 17 July 2026) — Metaculus, read
- scoring/score_math.py — Metaculus, read
Updated