Concept

ROC curve — where it appears

The trace of a forecaster's hit rate against its false-positive rate as the threshold for acting on it is swept from one end to the other. The area under it is the chance a randomly chosen event that happened was given a higher probability than one that did not, and it depends on the ordering of the forecasts rather than on their values.

Named by 5 essays across 2 fields — each of them below, with the objects they name alongside it.

The operating characteristic, and the point that minimises harm at 0.10% prevalence. The published pair — 90% sensitive, 95% specific — is the open mark. With a miss costing 100 times a false alarm and a prevalence of 0.10%, the threshold that minimises expected cost sits at 74.7% sensitivity and 98.81% specificity, with a predictive value of 5.9%.

The test is a point somebody chose

A test reported as 90% sensitive and 95% specific is not two properties of a test. It is one property read at a threshold, and the threshold that minimises harm runs from 3.05 standard deviations of the score at a prevalence of one in ten thousand to −0.12 at one in two — 45% of cases detected at one end and 99.9% at the other.

screening · Baserate
Six forecasters, all calibrated, not equally useful. The resolution of six forecasters that are all perfectly calibrated, each reporting the true probability of the event given a signal that carries more or less of the latent state. The largest reliability anywhere in the family is 2.0e-33, so a calibration check passes every one of them. They are not equally good: resolution runs from exactly 0.00 for the forecaster that issues the base rate every time to 0.092758 for the one that sees everything, and their Brier scores run from 0.234237 — which is the world's own uncertainty, and the score of a table of base rates — to 0.141479. Calibration is a necessary condition that a constant forecast satisfies exactly.

Calibrated and useless

Six forecasters that are calibrated to 2·10⁻³³ run from resolution exactly 0 to 0.092758, and three forecasters with reliabilities from 0 to 0.013025 have areas under the ROC curve identical to every bit a double carries. Each measure is exactly blind to what the other one sees.

calibrate · Calibration
Where each rule says to put the number. The expected score of reporting each probability on the axis when the event's true probability is 0.25, for four scoring rules. The Brier, logarithmic and spherical scores each bottom out at 0.25 — found by search rather than assumed, to 8 decimal places — which is what makes them proper: a forecaster with a genuine belief cannot improve its expected score by reporting anything else. The absolute-error score is a straight line in the reported value, p + r(1 − 2p), so it has no interior minimum at all; its optimum is 0.0, a distance of 0.250 from the truth, and taking it saves 0.125.

A score that rewards lying

An absolute-error score pays a forecaster exactly ⅛ of a point to replace a true quarter with a zero, and over two hundred records a liar beats a truthful forecaster on 200 of 200. A skill score against the forecaster's own average buys 0.012633 of reported skill for 0.002035 of real score.

calibrate · Calibration
One curve, and two forecasters that are each a single point. The ROC curve of an honest forecaster whose signal has correlation 1 with the latent state — every threshold on its probability, from the bivariate normal — and two forecasters that only ever say 0 or 1. The one an absolute-error score pays for says 1 wherever the honest probability exceeds a half: it has a true-positive rate of 0.6742 and a false-positive rate of 0.1375, and its "curve" is the two straight segments through that point, with area (TPR + TNR)/2 = 0.7684. The honest curve's area is 0.8683. Thresholding at the base rate of 0.3744 instead of at a half gives the largest area any two-valued forecaster can have here, 0.7818, and it is still 0.0866 short.

The liar with two answers

The forecaster an absolute-error score pays for says only 0 or 1, and on the ROC square it is a single point: its area is (TPR + TNR)/2 = 0.7684, against the honest forecaster's 0.8683, and it falls below the honest one on 200 of 200 counted records. No relabelling of its two answers returns what it threw away — the best recovers a Brier score short of the honest one by exactly the 0.022154 of resolution lost — and below a signal correlation of 0.7332 the same score prefers saying no every time to an honest forecast.

calibrate · Calibration
The honest curve, and the same forecaster in three coarse vocabularies. The ROC curve of the honest probability at full signal, area 0.8683, beside the same forecaster rounded to the nearest whole number, area 0.7684 — the two-valued liar — to the nearest half, area 0.8205, and to the nearest tenth, area 0.8650. A vocabulary of v values is v points on the square joined by straight segments, and tied reports count half.

A forecaster that rounds

An honest probability issued in tenths loses 0.0033 of ROC area and 0.000708 of resolution — the variance its bands average away, and 89.5% of the 0.000792 it adds to the Brier score. Two hundred records of two thousand forecasts show that loss on 189; it takes about 3,300 forecasts to put it two standard errors from zero. And 3.207 in every thousand forecasts in tenths are a 0% on an event that happened, which a logarithmic score charges without limit.

calibrate · Calibration

Named alongside it

The objects these essays reach for when they reach for this one.

Base rateBrier scoreDiscriminationProbability forecastProper scoring ruleReliability termResolution termClimatological forecastClosed formForecast calibrationImproper scoring ruleMurphy's decomposition

All concepts