What each forecaster says, and what is true
The true probability of the event given a forecaster's signal, and what three forecasters report. The truth is Φ(-0.5 + 1.2w) and the honest forecaster reports it, so its curve and the truth are the same line. The loud forecaster pushes every probability towards the ends and the hedged one pulls every probability towards the middle; both have their mean report held at the base rate of 0.3744, so each crosses the truth exactly once and neither can be caught by checking its average. Their reliability terms are 0.008500 and 0.013025 against the honest forecaster's zero, and all three have the same area under the ROC curve, 0.868312.
A forecast that is a probabilitywide8 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
The Brier score of five forecasters, split into what each one is charged for saying the wrong probability, what it is credited for telling one case from another, and what the world's own uncertainty costs everybody. The three add to the score exactly — the largest gap across the table is 2.22e-16 — so they are a rewriting of the score rather than a model of it. The uncertainty term is the same 0.234237 for all five, because it is a property of the world. The forecaster that reports the base rate every time has nothing in the other two, and scores exactly the uncertainty; the honest forecaster has nothing in the first and 0.092758 in the second.
A reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.
The resolution of six forecasters that are all perfectly calibrated, each reporting the true probability of the event given a signal that carries more or less of the latent state. The largest reliability anywhere in the family is 2.0e-33, so a calibration check passes every one of them. They are not equally good: resolution runs from exactly 0.00 for the forecaster that issues the base rate every time to 0.092758 for the one that sees everything, and their Brier scores run from 0.234237 — which is the world's own uncertainty, and the score of a table of base rates — to 0.141479. Calibration is a necessary condition that a constant forecast satisfies exactly.
The expected score of reporting each probability on the axis when the event's true probability is 0.25, for four scoring rules. The Brier, logarithmic and spherical scores each bottom out at 0.25 — found by search rather than assumed, to 8 decimal places — which is what makes them proper: a forecaster with a genuine belief cannot improve its expected score by reporting anything else. The absolute-error score is a straight line in the reported value, p + r(1 − 2p), so it has no interior minimum at all; its optimum is 0.0, a distance of 0.250 from the truth, and taking it saves 0.125.
The calibration error shown by a forecaster with no miscalibration in it at all, at five record lengths, drawn against one over the square root of the length so that the closed form is a straight line through the origin. It is 0.1252 at fifty forecasts and 0.0090 at ten thousand, against a closed form of √(2K/πn) times the mean root bin variance which gives 0.1257 and 0.0089. The threshold drawn across it is 0.02, a figure routinely read as evidence that something is wrong; the mean falls under it at 1976 forecasts and the 95th percentile at about 4111. Below that, an honest forecaster and a miscalibrated one are being told apart by a statistic that is mostly the sample size.
The ROC curve of an honest forecaster whose signal has correlation 1 with the latent state — every threshold on its probability, from the bivariate normal — and two forecasters that only ever say 0 or 1. The one an absolute-error score pays for says 1 wherever the honest probability exceeds a half: it has a true-positive rate of 0.6742 and a false-positive rate of 0.1375, and its "curve" is the two straight segments through that point, with area (TPR + TNR)/2 = 0.7684. The honest curve's area is 0.8683. Thresholding at the base rate of 0.3744 instead of at a half gives the largest area any two-valued forecaster can have here, 0.7818, and it is still 0.0866 short.
The ROC curve of the honest probability at full signal, area 0.8683, beside the same forecaster rounded to the nearest whole number, area 0.7684 — the two-valued liar — to the nearest half, area 0.8205, and to the nearest tenth, area 0.8650. A vocabulary of v values is v points on the square joined by straight segments, and tied reports count half.
Where it is used
7 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 7 different questions.
- An identity in three terms A forecast that is a probability
- A curve that is a binning A forecast that is a probability
- Calibrated and useless A forecast that is a probability
- A score that rewards lying A forecast that is a probability
- The miscalibration a perfect forecaster shows A forecast that is a probability
- The liar with two answers A forecast that is a probability
- A forecaster that rounds A forecast that is a probability