An identity in three terms
Worth reading first: What the model says next.
A forecast of 30% cannot be wrong about the day it was issued for. It can only be wrong about a collection of days, which is why every account of what a probability forecast is worth reaches for the same rewriting: the score a forecaster earns is what it is charged for saying the wrong number, minus what it is credited for telling one case from another, plus what the world’s own uncertainty costs everybody. Three terms, quoted as though they were a fact of arithmetic.
They are a fact of arithmetic, and the interesting part is where. In a world stated closely enough that every expected score is an integral rather than a count, the three terms reproduce the score to 2.2·10⁻¹⁶ — the last bit a double carries, across five forecasters that differ in what they know and in how loudly they say it. On a finite record the same three terms, computed on a reliability diagram of five bins, are out by 0.004125. That is not sampling noise. It is a fourth term, its size set by how coarsely the diagram was drawn, and it disappears exactly when the diagram becomes too fine to read.
So the decomposition is an identity, and it is an identity on a grouping nobody uses. On the one partition where the three terms are exact — a bin for each distinct forecast value, which for continuous forecasts is a bin per observation — reliability becomes the whole score at 0.14138 and resolution exactly cancels uncertainty at 0.23347 apiece. Everything the decomposition was supposed to separate has collapsed into one term. That is the shape of the finding, and the rest of this is what produces it.
A world where the truth is available
Nothing below is read off a record. An event happens with probability Φ(−0.5 + 1.2Z) for a latent standard normal Z, a forecaster sees a signal W correlated λ with Z, and it reports Φ(A + cBw). Two dials: λ is how much it knows, c is how loudly it says it. The base rate is 0.3744 in closed form and the variance of the events themselves is 0.234237, which is the score of a forecaster that issues the base rate every time and reads nothing.
The intercept is not free. It is fixed so that the mean forecast equals the base rate at every c, which removes the one kind of miscalibration nobody needs a diagram to find. A forecaster whose average probability is wrong is caught in a week by comparing two numbers. What is left in the family is the shape of the calibration curve rather than its level, and that is the kind a reliability diagram exists to detect.
The reason for stating a world rather than reading one is that both routes to every number are then available. A forecaster’s true reliability and its true information content are known in closed form, so a quantity computed from a sample has something to be checked against that is not a second sample of itself. That is the discipline the essay on computing every number twice argues for, applied to a subject where the second route is usually missing: a real forecasting record has no oracle in it, and every diagnostic drawn on one is therefore being trusted rather than checked.
Five forecasters are read throughout. The honest posterior; the same posterior from a weaker signal; the base rate every time; the honest posterior said too loudly, at c = 1.8; and the honest posterior hedged, at c = 0.5. Three of them are perfectly calibrated and two are not, and the two that are not are wrong in opposite directions by construction.
The third of those is worth naming now because it does the most work later. A forecaster that issues 0.3744 on every occasion is perfectly calibrated: of the occasions it said 0.3744, exactly the base rate happened, which is the whole of what calibration asks. It is also useless. That a check can be passed by a table of base rates is not a defect of this particular world — it is what the definition says, and the arithmetic here only makes it exact. The same reading of a probability as a claim about a reference class is what makes a positive test result on a rare condition mostly wrong, and what naming the screening arithmetic as a posterior update turns into an instance of a rule rather than a puzzle.
Each term is a definition, not a residual
The rewriting is worth writing down in the form it is actually computed in, because every one of the three terms is a definition and none of them is what is left over. Group the forecasts somehow, let bin k hold n_k of them with mean forecast f̄_k and observed event rate ō_k, and let ō be the overall rate:
Reliability is the mean squared gap between what was said and what happened, group by group; it is zero for a forecaster reporting what it actually believes. Resolution is how far its groups’ event rates spread away from the base rate; it is zero for a forecaster that says the same thing every time. Uncertainty is a property of the events and of nothing else. The first is a cost and the second is a credit, which is why the sign in front of resolution is the one that surprises people.
The population table is where the terms stop being an accounting device. The honest posterior’s reliability is exactly zero and its resolution is 0.092758, so its score is 0.234237 − 0.092758 = 0.141479 — computed a second time, directly, as an expected squared error, with the two routes agreeing to 2.2·10⁻¹⁶. The forecaster issuing the base rate scores 0.234237, which is the uncertainty term to the last digit. Neither of those is a fit. They are two ways of evaluating the same integral that share the quadrature grid and nothing else.
It is worth being clear about what kind of object this is, because a three-way split of a squared error into named parts invites the reading that it is a fit. It is not. Nothing was estimated to produce those terms, no parameter was chosen to make them add up, and the arithmetic would go through on a forecaster produced by a coin. That is the difference between this and the variance decomposition behind a share of variance explained, which is not a measure of fit either and rises when noise is added to the predictors. Resolution cannot rise when a forecaster is given a signal that carries nothing, because it is the spread of observed rates across the groups a forecaster’s own reports define, and a report that carries nothing defines groups whose rates are the base rate.
The loud forecaster’s reliability is 0.008500 and the hedged one’s 0.013025, against resolutions identical to the honest forecaster’s. That last equality is the first thing in this field that is stranger than it looks: distorting a forecast monotonically does not change how well it tells one case from another, because the true probability given the report is unchanged when the report is merely relabelled. The whole of the distortion lands in reliability and none of it anywhere else.
The residual is not small, and it is not noise
A decomposition checked only where it holds is a decomposition nobody has checked. So the residual — the score minus the three terms — is computed on every binning as well as on the one the algebra was written for.
On the partition that gives every distinct forecast value its own bin, the residual is 2.6·10⁻¹⁵. On five equal-count bins it is 0.004125, on ten 0.001276, on twenty 0.000547 and on fifty 0.000211. Those are means over a thousand records, and the standard error of each is far below the gaps between them: the sequence is monotone in the bin count, which rules out sampling as the explanation before any arithmetic is done. A quantity that fell as 1/√n would not fall as the bins were multiplied while n stayed at five hundred.
The measurement is a count over a thousand records rather than a demonstration on one, for the reason the essay on seeding every figure gives: a residual computed on a single record could be right by luck, and a sequence of four residuals falling monotonically in the bin count could be four lucky draws. Over a thousand records each, they are not.
Put the magnitude beside something. The honest forecaster’s whole score is 0.141479, and the largest of these residuals is 0.006805 across the forecasters and binnings the assertions sweep — nearly 5% of a score, discarded silently by an identity that was quoted as exact. The residual at five bins alone is half the reliability that a genuinely miscalibrated forecaster, the loud one, actually carries in population.
The fourth term is the binning’s, and it has a formula
Where the residual comes from is available without any measurement, and it is the reason the exact partition is the exact one. Writing for how far a forecast sits from its own bin’s mean, the algebra leaves
and adding it to the three makes the statement an identity again. The largest four-term residual across every forecaster and every binning tried is 6.7·10⁻¹⁵, which is the same order as the three-term residual on the exact partition — because on that partition every d_i is zero and the fourth term vanishes term by term, not on average.
That is the whole of it. The within-bin term is exactly the spread of forecasts inside a bin that the diagram threw away when it replaced them with their mean. A reliability diagram is a picture that summarises each bin by two numbers, and the arithmetic of the decomposition absorbs the discarded third one. Coarser bins throw away more: at five bins the term is −0.004052 and at fifty it is −0.000046, a factor of nearly ninety for a factor of ten in bins.
Which leaves the finding stated plainly. The three-term identity holds exactly when every forecast inside a bin is the same number. For continuous forecasts that means one bin per observation, and there the decomposition is not merely exact but empty: reliability rises to 0.14138, the entire score, because every bin holds one forecast and one outcome and the gap between them is the whole error; resolution rises to 0.23347, which is the uncertainty term, because each bin’s rate is 0 or 1. A decomposition whose two informative terms have collapsed onto the score and the uncertainty is telling nobody anything. The identity is exact where it is useless and approximate where it is read.
Two proper scores, two orders
The same three objects can be built for the logarithmic score, with a divergence in place of a squared distance and an entropy in place of a variance. Both scores are proper — both are minimised in expectation by reporting what the forecaster actually believes — so both put the honest forecaster first, and the question is whether they agree about anything else.
They do not. By the Brier score the loud forecaster beats the hedged one, 0.149979 against 0.154504, a margin of 0.004525. By the logarithmic score the order reverses: 0.501048 against 0.476965, a margin of 0.024083 the other way. Nothing about either forecaster changed. Both scores are proper, both were computed by quadrature on the same grid, and they disagree about which of two miscalibrated forecasters is the better one.
The pair is not a curiosity found by searching. Fixing the hedged forecaster and solving for the confidence multiplier at which the two have equal expected Brier score gives c = 1.8985; a step below that value the Brier score strictly prefers the loud forecaster and the logarithmic score strictly prefers the hedged one, so any pair straddling the tie is a disagreement. The reversal is a region rather than an example, in the same sense that Simpson’s reversal is a region rather than a famous table.
What a confident mistake costs is a choice
The reason the two scores disagree is not subtle and is worth stating as arithmetic rather than as temperament. The Brier score charges at most 1 for a forecast that is completely wrong. The logarithmic score charges −log(ε), which is unbounded. A forecaster that pushes its probabilities towards the ends is buying small charges on the many occasions it is right at the price of large ones on the few occasions it is wrong, and whether that trade is good depends entirely on how the large ones are priced.
The logarithmic terms are the same three objects. Uncertainty is the base rate’s binary entropy, 0.661281 nats; resolution is how much of that entropy the forecaster’s signal removes, 0.229005 for the honest one; reliability is the mean divergence between what is true given the report and the report itself, which is zero for a forecaster reporting its own belief, 0.068772 for the loud one and 0.044690 for the hedged one. Their sum reproduces the honest forecaster’s logarithmic score of 0.432276, and the largest gap between the three terms and the score across all five forecasters is 4.4·10⁻¹⁶ — so both decompositions are identities in population and neither is a fit to the other.
The loud forecaster’s reliability under the logarithmic score is more than eight times its reliability under the Brier score, while the hedged one’s is only three and a half times. That ratio is the disagreement, written in the terms rather than in the totals: the two rules do not weight miscalibration differently in general, they weight this kind of miscalibration differently, and overconfidence is the kind an unbounded rule punishes hardest.
So choosing between two proper scores is not choosing between two measurements of one thing. It is choosing what a confident mistake costs, and the choice changes the answer. That is a different complaint from the one usually made about scoring rules, which is that improper ones exist; two proper rules ranking the same pair in opposite orders is a much narrower and more awkward fact, and it is the reason the field that compares point forecasts by squared error cannot simply be lifted one level up. There, the loss function is a modelling choice made once. Here it decides an ordering, and the ordering is the output. The same asymmetry shows up in the second question that field asks about a pair of forecasters — whether either of them is redundant given the other — where the answer turns out not to move with the series at all. Redundancy is a statement about information and information is what a monotone relabelling leaves alone; accuracy is a statement about numbers, and numbers are exactly what a relabelling changes.
One of the two numbers depends on where certainty is cut off
An unbounded score cannot be computed at all without a rule for what happens at zero, and the rule is usually buried. Every report here is clamped into [ε, 1 − ε] before a logarithm is taken of it, and what that clamp is worth was measured rather than assumed.
Moving it from 10⁻¹⁰ to 10⁻¹⁴ — four orders of magnitude — leaves the honest forecaster’s expected logarithmic score unchanged to the last bit, a difference of exactly zero. It moves the loud forecaster’s by 6.3·10⁻⁷. The Brier score of the same loud forecaster does not move at all, because it is bounded by one and there is nothing at the ends for a clamp to reach.
Six decimal places is not where the disagreement above lives, so nothing in the ranking is at risk. The point is the direction of the dependence: an unbounded score’s value for a very confident forecaster is set by how confident it is allowed to be, and that is a parameter of the evaluation rather than of the forecaster. Tightening it further would not help, because the arithmetic runs out first — the normal distribution function used throughout returns exactly zero below about −8.5, so a clamp below 10⁻¹⁴ would be a clamp against a quantity that has already collapsed. That is the same class of defect as an interval whose endpoint is a property of its estimator’s own breakdown: a number that looks like a measurement and is a property of the machinery.
What is claimed, and what is left
Three things are claimed. The three-term rewriting of a probability score is an identity, in population to 2.2·10⁻¹⁶ and on a sample only on the partition by distinct forecast value. On any readable binning it is out by the within-bin term, whose size the binning sets and which is 0.004125 at five bins. And two proper scores built from the same three objects rank the same pair of forecasters in opposite orders, by margins of 0.004525 and 0.024083.
What is not claimed is anything about which score to use. The disagreement is stated and its cause is named — a bounded charge against an unbounded one — and choosing between them is a decision about what a confident error should cost, which is not a question a measurement answers. Nor is anything claimed here about estimating these terms: every number in this essay is either an integral or a mean over a thousand records, and the behaviour of the reliability term on one record of five hundred forecasts is a different and worse story. The essay that reads a diagram as a binning is where that goes, and it is where the residual measured here stops being an accounting curiosity and starts charging blameless forecasters.
One further limit belongs here rather than there. Everything above is a statement about a family of forecasters whose reports are strictly increasing functions of one signal. That is what makes the resolutions of the honest, loud and hedged forecasters identical, and it is a real restriction: a forecaster that is well calibrated in the middle of its range and confused at both ends is not in this family, none of the three terms behaves the same way for it, and nothing here measures one.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The miscalibration a perfect forecaster shows — both name binning, climatological forecast, forecast calibration, probability forecast, reliability term
Named objects
A flat tag is an object no other essay names yet.
Base rateBinningBrier scoreClimatological forecastForecast calibrationLogarithmic scoreMurphy's decompositionProbability forecastProper scoring ruleQuadratureReliability termResolution termUncertainty termVariance decompositionWithin-bin variance