A forecast that is a probability

Calibrated and useless

Six forecasters that are calibrated to 2·10⁻³³ run from resolution exactly 0 to 0.092758, and three forecasters with reliabilities from 0 to 0.013025 have areas under the ROC curve identical to every bit a double carries. Each measure is exactly blind to what the other one sees.

Worth reading first: An identity in three terms · What the model says next.

A forecaster that issues 0.3744 every single day, having read nothing and looked at nothing, is perfectly calibrated. Of the occasions it said 0.3744, the event happened at exactly that rate, because those occasions are every occasion and 0.3744 is the base rate. There is no approximation in that sentence and no small print. The check that a probability forecast can be put to without knowing anything about how it was made is passed, exactly, by a table of base rates.

That is usually conceded as a caveat and then dropped. Taken seriously it is a measurement, and the measurement has an exact zero in it. Six forecasters that each report the true probability of the event given their own signal are calibrated to within 2·10⁻³³ — the largest reliability anywhere in the family — and they run from resolution exactly 0 to 0.092758, with Brier scores from 0.234237 to 0.141479. Every one of them passes. One of them is the best forecaster the world admits and one of them reads nothing, and no calibration check distinguishes them.

The converse has a sharper zero still. Three forecasters whose reliabilities are 0, 0.008500 and 0.013025 — the honest posterior, the same posterior pushed towards the ends, the same posterior hedged towards the middle — have areas under the ROC curve of 0.868311689903, 0.868311689903 and 0.868311689903. Not close. Identical to every digit a double carries, with a spread across the whole family of exactly zero, and the reason is arithmetic rather than approximate.

Six forecasters, all calibrated, and not equally good

The family is one dial. A forecaster observes a signal correlated λ with the latent state and reports the true probability of the event given what it saw. At λ = 0 the signal carries nothing and the honest report is the base rate; at λ = 1 the signal is the state and the honest report is everything the world will admit. Every member is calibrated by construction, because a conditional expectation given a forecaster’s own information is calibrated — that is what a conditional expectation is.

Six forecasters, all calibrated, not equally useful. The resolution of six forecasters that are all perfectly calibrated, each reporting the true probability of the event given a signal that carries more or less of the latent state. The largest reliability anywhere in the family is 2.0e-33, so a calibration check passes every one of them. They are not equally good: resolution runs from exactly 0.00 for the forecaster that issues the base rate every time to 0.092758 for the one that sees everything, and their Brier scores run from 0.234237 — which is the world's own uncertainty, and the score of a table of base rates — to 0.141479. Calibration is a necessary condition that a constant forecast satisfies exactly.
Fig. 1 The resolution of six forecasters that are all perfectly calibrated, each reporting the true probability given a signal carrying more or less of the latent state. Resolution runs from exactly zero to 0.092758 and their Brier scores from 0.234237 down to 0.141479, while the largest reliability anywhere in the family is 2·10⁻³³.

Resolution reads 0.000000, 0.005310, 0.021420, 0.042677, 0.064391 and 0.092758 at signal correlations 0, 0.25, 0.5, 0.7, 0.85 and 1. Their scores are 0.234237, 0.228927, 0.212817, 0.191560, 0.169846 and 0.141479, which are the uncertainty term minus the resolution term exactly, because reliability is zero throughout. The whole difference between the worst and the best forecaster in the family lives in a term a calibration check does not touch.

The first row is the one to keep. Its score, 0.234237, is the variance of the events themselves — the number every forecaster in this world is charged and none can reduce — so a table of base rates scores exactly the uncertainty and exactly nothing else. Its area under the ROC curve is 0.50, which is what a coin gets. It is, on the one check people describe as the fundamental requirement of a probability forecast, flawless.

The comparison is also the reason a benchmark has to be in the table rather than in the discussion. A set of forecasters ranked against each other says nothing about whether any of them beats reading nothing, and the forecaster that reads nothing is the one every skill score divides by — which is why putting a benchmark into the field of candidates rather than beside it changes what a comparison can conclude. Here the benchmark is not merely a low bar; it is a member of the calibrated family in good standing, and any diagnostic that ranks by calibration alone has it tied for first.

So the standard prescription — maximise sharpness subject to calibration — is not a refinement of the calibration requirement. It is a second requirement doing all the work, with the first one contributing a constraint that a constant already satisfies. Stating it as a single slogan hides which half is binding, in the same way that quoting one number for a forecast comparison hides whether the difference between two forecasters would survive being asked for a standard error.

What a calibration check cannot see is how far the forecasts go

Sharpness is the spread of the probabilities a forecaster is willing to issue, and for a calibrated forecaster it equals the resolution term exactly: both are 0.092758 for the best-informed one and both are 0.005310 for the worst. That equality is not a coincidence and it is the cleanest statement of what a calibration check misses. A calibrated forecaster’s reports have the same variance as the truths behind them, so the only way to be more useful is to say more extreme things and be right about them.

How far a forecaster spreads its probabilities. The mean probability issued in each tenth of a forecaster's own record, for three honest forecasters in a world whose base rate is 0.3744. Every one of them is perfectly calibrated — each reports the true probability of the event given the signal it saw — and they differ only in how much of the latent state their signal carries. The best informed runs from 0.0079 to 0.9286 across its deciles and has a resolution of 0.092758; the worst runs from 0.2523 to 0.5068 and has 0.005310. A forecaster that issued the base rate every time would be the flat line at 0.3744, and it would be calibrated too.
Fig. 2 The mean probability issued in each tenth of a forecaster’s own record, for three honest forecasters differing only in how much of the latent state their signal carries. The best informed runs from 0.0079 to 0.9286 across its deciles; the worst runs from 0.2523 to 0.5068. All three are perfectly calibrated.

The best-informed forecaster’s lowest decile averages 0.0079 and its highest 0.9286. The worst-informed one runs from 0.2523 to 0.5068 — it never says anything below a quarter or above a half, and it is right about all of it. A forecaster that issued the base rate every time would be the flat line at 0.3744, and it too would be right about all of it.

That picture is what a calibration diagnostic is blind to, and the blindness is total rather than partial. Reliability is a statement about the gap between a forecast and the truth behind it; a forecaster whose gap is zero has said nothing about how far from the base rate it was willing to go. Nothing that reads only that gap can rank the three curves above, and the three curves are the entire difference between a useful forecaster and a useless one.

There is a practical version of this that shows up whenever forecasts are compared without a benchmark. A forecaster asked to be well calibrated, and evaluated on whether it is, has an available strategy: hedge towards the base rate until the check is satisfied. It costs resolution and the check does not measure resolution. That is the same structure as a rate held at 5% while a different rate runs free — a promise kept exactly, and a different promise that nobody asked about broken as far as necessary to keep it.

Resolution is the term that separates them, and it is in the decomposition already

Nothing new had to be invented for this. The three-term rewriting of a probability score already carries the distinction: reliability charges a forecaster for saying the wrong number and resolution credits it for telling one case from another, and the family above holds the first at zero while sweeping the second.

Three terms that add to the Brier score. The Brier score of five forecasters, split into what each one is charged for saying the wrong probability, what it is credited for telling one case from another, and what the world's own uncertainty costs everybody. The three add to the score exactly — the largest gap across the table is 2.22e-16 — so they are a rewriting of the score rather than a model of it. The uncertainty term is the same 0.234237 for all five, because it is a property of the world. The forecaster that reports the base rate every time has nothing in the other two, and scores exactly the uncertainty; the honest forecaster has nothing in the first and 0.092758 in the second.
Fig. 3 The Brier score of five forecasters split into what each is charged for saying the wrong probability, what it is credited for telling one case from another, and what the world’s own uncertainty costs everybody. The forecaster that reports the base rate every time has nothing in the first two terms and scores exactly the uncertainty, 0.234237.

What the family adds is that the two terms are independently controllable, which the decomposition alone does not say. A rewriting into named parts is compatible with the parts being tied together — it would be a much weaker finding if every forecaster with low reliability also happened to have high resolution. Here they are orthogonal by construction: λ moves resolution with reliability pinned at zero, and the confidence multiplier moves reliability with resolution pinned at 0.092758. Both dials exist and neither disturbs the other.

That is the reason the identity checked in three terms is worth having despite being exact only where it is empty. It does not need to be an exact identity on a readable binning to be the right vocabulary; what it names are two genuinely separate properties of a forecaster, and the whole of this essay is a measurement that they are separate.

The ranking measure is blind in the other direction, and exactly so

Turn the family the other way. Fix the information at its maximum and move the confidence multiplier, so that every forecaster reports a strictly increasing function of the same signal and differs only in how loudly it says what it knows. Reliabilities then run from 0 to 0.017887 at the most distorted setting, and the areas under the ROC curve do not move at all.

What a ranking measure cannot see. The area under the ROC curve of five forecasters, computed from the bivariate normal rather than counted. Three of them — the honest posterior, the same posterior said too loudly, and the same posterior hedged — have areas that agree to nine decimal places at 0.868311690, because each reports a strictly increasing function of the same signal and a relabelling does not reorder anything. Their reliabilities are 0.000000, 0.008500 and 0.013025 and their Brier scores 0.141479, 0.149979 and 0.154504. What the measure does see is information: a forecaster with a weaker signal reads 0.751105, and one that issues the base rate every time reads exactly 0.50.
Fig. 4 The area under the ROC curve of five forecasters, computed from a bivariate normal rather than counted. The honest posterior, the same posterior said too loudly and the same posterior hedged agree to nine decimal places at 0.868312, while their reliabilities are 0, 0.008500 and 0.013025.

The measured spread across six confidence multipliers is 0.00e+0. Not small — zero, in double arithmetic, and the reason is that a ranking measure reads only the ordering of the forecasts. A strictly increasing relabelling does not reorder anything, so the set of events ranked above any given event is the same set before and after, and the area under the curve is the same number by construction rather than by luck. It is checked against a Mann–Whitney count over a sample that was never shown the formula, which is the second route the closed form needs.

What the same measure does see is information: 0.587909, 0.677325, 0.751105, 0.808483 and 0.868312 across the signal correlations, and exactly 0.50 for the forecaster that reads nothing. So the two instruments partition the properties of a forecaster cleanly and neither is a weak version of the other. Calibration sees the numbers and not the ordering; discrimination sees the ordering and not the numbers. A forecaster can be perfect on either and worthless on the other, and the two failures look nothing alike.

There is a second reason the ranking measure is the wrong instrument for this question, and it is independent of the blindness. An area under a ROC curve is an average over every threshold at once, and a probability forecast is normally read at one threshold that a decision fixes — which is where a test’s sensitivity and specificity stop being the numbers that matter and the base rate takes over. A forecaster whose probabilities are wrong but whose ordering is right supports a decision only once somebody has recovered the right cut-off, and recovering it is exactly the recalibration the last two sections price.

This is worth stating against the way the two are usually contrasted, which is that discrimination is a coarse summary that misses subtle miscalibration. It does not miss it subtly. It cannot see it at all, and no amount of data changes that: the blindness is a property of the functional, not of the estimate.

A weaker signal and a louder voice are different objects

The two dials produce forecasters that a single summary can confuse and the pair of summaries cannot.

What each forecaster says, and what is true. The true probability of the event given a forecaster's signal, and what three forecasters report. The truth is Φ(-0.5 + 1.2w) and the honest forecaster reports it, so its curve and the truth are the same line. The loud forecaster pushes every probability towards the ends and the hedged one pulls every probability towards the middle; both have their mean report held at the base rate of 0.3744, so each crosses the truth exactly once and neither can be caught by checking its average. Their reliability terms are 0.008500 and 0.013025 against the honest forecaster's zero, and all three have the same area under the ROC curve, 0.868312.
Fig. 5 What three forecasters report against the signal, with an honest forecaster at signal correlation 0.5 drawn quietly behind them. The loud and hedged forecasters cross the truth exactly once each, because their mean report is held at the base rate; the less informed honest forecaster does not cross it at all, because it is the truth given less.

An honest forecaster at signal correlation 0.5 and a hedged forecaster at correlation 1 both issue probabilities that stay nearer the base rate than the truth does. Their Brier scores are 0.212817 and 0.154504, so on the score alone one is much better than the other, but the shape of what they issue is similar enough that a reader shown two records could easily read the second as the first. The measures separate them immediately and in opposite ways: the areas under the ROC curve are 0.677325 and 0.868312, and the reliabilities are 0 and 0.013025.

So the diagnosis is available and it takes both instruments. A forecaster that is calibrated with low resolution is uninformed, and there is nothing to fix — it needs a better signal. A forecaster with high discrimination and non-zero reliability is informed and mislabelled, and there is something to fix, because the information is already in the ordering and only the numbers attached to it are wrong. That distinction is what the next two sections are about, and it is the reason the exact zero above is useful rather than merely tidy.

What a fitted repair recovers, and what it costs

If a distortion is a monotone relabelling of the truth, then composing the forecaster’s reports with the true calibration curve returns the honest posterior itself, and the whole deficit goes. In population that is exact and it is worth writing down as an identity rather than as a result: the score is reliability minus resolution plus uncertainty, a repair sets reliability to zero and leaves resolution alone, so the population deficit is exactly the reliability term — 0.008500 for the loud forecaster and 0.017887 for the most distorted one, to the digit.

On a record it is not exact, and the way the measurement is set up is the whole of what it says. The calibration curve is fitted by isotonic regression on one half of a record of two thousand forecasts and applied to the other half, over sixty pairs of records.

What a monotone repair gets back. The Brier score of a forecaster before and after a recalibration fitted on one half of its record and scored on the other, over 60 pairs of 2000 forecasts. Every forecaster on this axis reports a strictly increasing function of the same signal, so in population the repair is exact and returns the honest posterior itself; on a sample it recovers 85.5% of the loud forecaster's deficit and 93.9% of the most distorted one's, and never all of it. The line it is trying to reach is the true posterior's own score. The honest forecaster is the case worth reading twice: recalibrating it moves its score from 0.140781 to 0.141991, because a curve with a step per observation costs something even when there is nothing to correct.
Fig. 6 The Brier score of a forecaster before and after a recalibration fitted on one half of its record and scored on the other, across six confidence multipliers. It recovers 85.5% of the loud forecaster’s deficit and 93.9% of the most distorted one’s, and never all of it. Recalibrating the honest forecaster moves its score from 0.140781 to 0.141991.

The loud forecaster goes from 0.150623 to 0.143000 against an oracle of 0.141706, recovering 85.5% of its deficit. The most distorted one recovers 93.9% and the hedged one 91.4%. Never all of it, and the shortfall is the fitting: an isotonic fit on a thousand points is estimating a monotone curve from a thousand noisy Bernoulli outcomes, and what it gets wrong it carries into the other half.

The row that matters most is the honest forecaster’s. It has no deficit to recover — its reliability is zero — and recalibrating it moves its out-of-sample score from 0.140781 to 0.141991. It gets worse. A fitted correction is not free where there is nothing to correct, and the cost here is about a sixth of the 0.007623 the same repair recovers on the loud forecaster. That is the cleanest available statement of what fitting a curve costs, and it is the same shape as the repair that was exactly right and still made things worse: a correction derived correctly, applied where its premise does not hold, and paid for anyway.

The reading anybody computes first is impossible

The sweep was written the other way round at first, which is the ordinary way: fit the recalibration on a record and score it on the same record. Both readings are kept, because the second is the one a first attempt produces and the gap between them is the finding.

Scored in sample, the repair recovers 140% of the loud forecaster’s deficit, and across the family the figures run from 125% to 322%. A share above one is not a large effect; it is an impossible one — it says the repaired forecaster is better than the forecaster it was repaired towards. The arithmetic confirms it directly: the in-sample repaired score is 0.137394 against the true posterior’s 0.140718 on the same events. The fitted curve beats the truth.

The failure is a selection rather than a bug, in the same family as the effects that look large because significance was the filter: nothing was computed incorrectly, and the quantity reported is not the quantity anybody wanted. What it is, of course, is a curve with a step available per observation. Isotonic regression pools adjacent points only where they violate monotonicity, so on a record where they mostly do not it returns something close to the outcomes themselves, and scoring that against the outcomes it was fitted to measures how well a thing reproduces itself. Fitted on the honest forecaster, which has nothing wrong with it, it still improves the in-sample score by 0.003288 — a deficit of exactly zero, apparently repaired.

That is the reason the whole sweep is split, and it is the same reason a criterion has to be read as a prediction of a hold-out rather than as a statement about the sample it was computed on. It is also the reason the honest forecaster is in the sweep at all: a family in which every member has something to repair could not have shown this, because there would have been no row whose true answer was known to be zero. An impossible number is the most useful kind of failure, because it cannot be argued into a small effect.

Where this stops holding

Every forecaster measured here reports a strictly increasing function of one honest posterior. That is not a mild regularity condition; it is what produces all three results. The exact blindness of the ranking measure is a statement about monotone relabellings. The exactness of the population repair is a statement about monotone relabellings. Even the equality of resolution across the confidence family follows from it.

A forecaster whose calibration curve is not monotone in its own report — well calibrated in the middle of its range and confused at both ends, say, which is a shape real forecasting records show — would break all three. Its ordering would carry a genuine defect, so discrimination would move; a monotone repair could not return it to the honest posterior, so the population recovery would be under one; and its resolution would differ from the honest forecaster’s. Nothing here measures such a forecaster, and the results above should not be read as though it did.

The second limit is the record length. Everything about the repair is measured at two thousand forecasts per half, and the shortfall from a total recovery is a fitting cost that shrinks with the record. Whether it shrinks fast enough to matter at the two hundred forecasts a real evaluation usually has is not measured here, and the direction is not encouraging — the same record length at which a reliability diagram’s own reported number depends on how many bins it was drawn in by a factor of ten: the cost of fitting a curve with a free step per point does not fall quickly. What is claimed is the exact statement in population, the exact zero in the ranking measure, and the counted out-of-sample shares at the record length stated beside them.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Base rateBrier scoreClimatological forecastDiscriminationForecast calibrationIsotonic regressionMonotone transformationOut of sampleOverconfidenceProbability forecastRecalibrationReliability termResolution termROC curveSharpness