Concept

Murphy's decomposition — where it appears

The rewriting of a probability score as reliability minus resolution plus uncertainty. It is an identity rather than a model when every forecast inside a group is the same number, and on any coarser grouping it leaves a further term whose size is set by how coarsely the grouping was made.

Named by 4 essays across one field — each of them below, with the objects they name alongside it.

What each forecaster says, and what is true. The true probability of the event given a forecaster's signal, and what three forecasters report. The truth is Φ(-0.5 + 1.2w) and the honest forecaster reports it, so its curve and the truth are the same line. The loud forecaster pushes every probability towards the ends and the hedged one pulls every probability towards the middle; both have their mean report held at the base rate of 0.3744, so each crosses the truth exactly once and neither can be caught by checking its average. Their reliability terms are 0.008500 and 0.013025 against the honest forecaster's zero, and all three have the same area under the ROC curve, 0.868312.

An identity in three terms

Reliability minus resolution plus uncertainty is quoted as a rewriting of a probability score. It is an identity to 2.6·10⁻¹⁵ on the one grouping where reliability is the whole score and resolution exactly cancels uncertainty, and it is out by 0.004125 on the coarsest grouping anybody would actually draw.

calibrate · Calibration
One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.

A curve that is a binning

A forecaster with no miscalibration in it at all reads 0.001429 at five bins and 0.014100 at fifty, on the same five hundred forecasts. The closed form is K/n times the forecaster's own irreducible score, and subtracting it returns zero.

calibrate · Calibration
One curve, and two forecasters that are each a single point. The ROC curve of an honest forecaster whose signal has correlation 1 with the latent state — every threshold on its probability, from the bivariate normal — and two forecasters that only ever say 0 or 1. The one an absolute-error score pays for says 1 wherever the honest probability exceeds a half: it has a true-positive rate of 0.6742 and a false-positive rate of 0.1375, and its "curve" is the two straight segments through that point, with area (TPR + TNR)/2 = 0.7684. The honest curve's area is 0.8683. Thresholding at the base rate of 0.3744 instead of at a half gives the largest area any two-valued forecaster can have here, 0.7818, and it is still 0.0866 short.

The liar with two answers

The forecaster an absolute-error score pays for says only 0 or 1, and on the ROC square it is a single point: its area is (TPR + TNR)/2 = 0.7684, against the honest forecaster's 0.8683, and it falls below the honest one on 200 of 200 counted records. No relabelling of its two answers returns what it threw away — the best recovers a Brier score short of the honest one by exactly the 0.022154 of resolution lost — and below a signal correlation of 0.7332 the same score prefers saying no every time to an honest forecast.

calibrate · Calibration
The honest curve, and the same forecaster in three coarse vocabularies. The ROC curve of the honest probability at full signal, area 0.8683, beside the same forecaster rounded to the nearest whole number, area 0.7684 — the two-valued liar — to the nearest half, area 0.8205, and to the nearest tenth, area 0.8650. A vocabulary of v values is v points on the square joined by straight segments, and tied reports count half.

A forecaster that rounds

An honest probability issued in tenths loses 0.0033 of ROC area and 0.000708 of resolution — the variance its bands average away, and 89.5% of the 0.000792 it adds to the Brier score. Two hundred records of two thousand forecasts show that loss on 189; it takes about 3,300 forecasts to put it two standard errors from zero. And 3.207 in every thousand forecasts in tenths are a 0% on an event that happened, which a logarithmic score charges without limit.

calibrate · Calibration

Named alongside it

The objects these essays reach for when they reach for this one.

Brier scoreProbability forecastReliability termBase rateProper scoring ruleResolution termBinningClosed formDiscriminationForecast calibrationQuadratureReliability diagram

All concepts