From books

A forecast becomes measurable when it is a number on a question with a deadline, and the Brier score shows who was calibrated, not who sounded convincing.

Barbara Mellers, Philip E. Tetlock et al. · Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions · 2015 · Mellers, Tetlock et al., «Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions», Perspectives on Psychological Science (2015), p. 273, secțiunea despre stilurile cognitive ale superprognozatorilor2 minutes read
This worldview predisposes superforecasters to treat their beliefs more as testable hypotheses and less as sacred possessionsBarbara Mellers, Philip E. Tetlock et al. · Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions · 2015 · Mellers, Tetlock et al., «Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic Predictions», Perspectives on Psychological Science (2015), p. 273, secțiunea despre stilurile cognitive ale superprognozatorilor

A forecast with no deadline and no resolution criterion cannot be lost, so it cannot teach you anything.

Between 2011 and 2015 IARPA, the research agency of the US intelligence community, ran geopolitical forecasting tournaments: hundreds of questions with a verifiable answer and a fixed deadline, on which participants gave probabilities they could revise at any time. The Good Judgment Project, led by Philip Tetlock and Barbara Mellers, won, and the book Superforecasting (Tetlock and Gardner, 2015) tells how. The yardstick was the Brier score: the sum of squared differences between the probability given and what happened (1 if the event occurred, 0 if not). Zero is perfect; always saying 50% earns 0.5 on a two-outcome question. The method has four steps. Phrase the question so that a stranger could decide at the deadline whether it happened; give a number, not a maybe; update often, in small steps, as evidence arrives; at resolution, record the score and check calibration — of the claims made at 70%, did roughly 70% come true? It pays off for any judgement that recurs. The trap is choosing vague questions that can never be lost: with no deadline and no criterion there is no score, and so no learning.

Why it mattersAn analyst who keeps no score cannot know whether their ‘likely’ means 60% or 90% — and their reader knows even less.

Questionwith aNumericalprobability,Resolutionand BrierRecalibrationfrom your
Without the last step the loop breaks: you forecast, but you do not learn.

Back to the feed