A forecast that is a probability

A curve that is a binning

A forecaster with no miscalibration in it at all reads 0.001429 at five bins and 0.014100 at fifty, on the same five hundred forecasts. The closed form is K/n times the forecaster's own irreducible score, and subtracting it returns zero.

Worth reading first: An identity in three terms · What the model says next.

A reliability diagram is the picture that answers whether a probability forecast means what it says. Group the forecasts, count how often the event happened in each group, plot the one against the other, and read the departure from the diagonal. It is the only check a probability forecast can be put to without knowing anything about how it was made, and it is drawn in every account of forecast evaluation there is.

Take a forecaster with no miscalibration in it whatsoever — one reporting the true probability of the event given its own signal, whose population reliability is exactly zero by construction — and draw its diagram on five hundred forecasts. At five equal-count bins the reliability term reads 0.001429. At fifty it reads 0.014100. Same forecaster, same five hundred forecasts, same outcomes, a factor of ten between the two numbers, and the true value is zero.

That is not a sampling accident to be averaged away. Both figures are means over a thousand independent records, their standard errors are 3.5·10⁻⁵ and 1.0·10⁻⁴, and the sequence across bin counts is monotone. It is a bias rather than sampling variation, it has a closed form that can be written down before any outcome is looked at, and the factor in that closed form turns out to be the forecaster’s own irreducible score. A reliability diagram is a picture of a forecaster and a binning, and the number read off it belongs to both.

One record, at three resolutions

The construction below is the ordinary one. Five hundred forecasts from the honest posterior in a world whose base rate is 0.3744, sorted, cut into K groups of equal count, and each group summarised by its mean forecast and its observed event rate.

One forecaster, one sample, and the bins it was read inA reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.00.2500.5000.750100.2000.4000.6000.8001the mean probability the bin was issuedthe share of that bin's events that happenedone sample, binnedthe same forecaster, in populationperfect calibration500 forecasts, 10 equal-count bins, dot area is the bin's countreliability 0.002262
Fig. 1 One record of five hundred forecasts in equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The slider moves the number of bins.

The slider is the whole argument in one control. Nothing about the forecaster changes as it moves; nothing about the record changes; the five hundred pairs of numbers underneath are the same five hundred pairs at every setting. What changes is how many of them are averaged together before the average is plotted, and the picture goes from a handful of points sitting almost exactly on the diagonal to a scatter that looks like a forecaster in trouble.

The direction is worth being explicit about, because the intuition points the wrong way. More bins is more resolution, and more resolution usually means seeing more of what is there. Here it means seeing more of what is not: a bin’s observed rate is an average of n/K Bernoulli outcomes, its standard error grows as √(K/n), and the diagram plots that standard error as though it were a departure. The reliability term squares those departures and sums them with the bin counts as weights, so the noise enters at first order in K and the signal does not enter at all.

One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 5 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.001392 and its expected calibration error 0.0329, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.
Fig. 2 The same record at five bins. A hundred forecasts in each group, so each observed rate is an average of a hundred outcomes and the points sit close to the diagonal; the reliability term reads 0.001429 on average across a thousand such records.

At five bins each group holds a hundred forecasts and each plotted rate is an average of a hundred Bernoulli draws, so the vertical scatter is roughly 0.04 and the points look convincing. At fifty bins each group holds ten, the scatter is roughly 0.12, and the same forecaster’s diagram is a cloud. A reader shown only the second picture and told it was a real forecaster’s record would have something to say about it, and would be wrong.

One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 50 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.011677 and its expected calibration error 0.0768, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.
Fig. 3 The same record again at fifty bins. Ten forecasts to a group, so each observed rate is one of eleven possible values, and the diagram of a forecaster whose true reliability is zero scatters across the whole square.

This is the same defect the essay that read one bootstrap three ways found in a different instrument: the quantity a picture reports is not the quantity the picture is about, and the gap between them is a property of how the picture was constructed. There the choice was a block length; here it is a bin count, and unlike a block length it is almost never reported at all.

The bias is the forecaster’s own irreducible score, divided by n and multiplied by K

The closed form follows from the construction and needs no simulation. Within bin k the observed rate ō_k is a mean of n_k independent Bernoulli draws whose probabilities are the forecasts themselves, so under the null that the forecaster is calibrated,

E[(fˉkoˉk)2]  =  1nk2ikfi(1fi)\mathbb{E}\big[(\bar f_k - \bar o_k)^2\big] \;=\; \frac{1}{n_k^2}\sum_{i \in k} f_i(1 - f_i)

and summing over bins with weights n_k/n gives, for equal-count bins,

E[REL^]  =  KnE[f(1f)]\mathbb{E}\big[\widehat{\mathrm{REL}}\big] \;=\; \frac{K}{n}\,\mathbb{E}\big[f(1 - f)\big]

The factor is E[f(1 − f)], and that quantity is 0.141479 — which is exactly the honest forecaster’s own expected Brier score, since a calibrated forecaster’s score is the mean variance of the outcomes it cannot predict. So the sentence the formula says is that a reliability diagram charges a blameless forecaster K/n of the score it can never avoid, and there is no bin count at which that charge is zero.

A blameless forecaster's diagram is not blank. What the reliability term of a reliability diagram reports for a forecaster whose true reliability is exactly zero, over 1000 samples of 500 forecasts. It is not zero at any binning and it grows in proportion to the number of bins: 0.001429, 0.002857, 0.005617, 0.014100 at 5, 10, 20, 50 equal-count bins, against a closed form of K·E[f(1 − f)]/n which gives 0.001415, 0.002830, 0.005659, 0.014148. The factor E[f(1 − f)] is 0.141479 — the honest forecaster's own irreducible score — so a diagram charges a blameless forecaster K/n of what it can never avoid. Equal-width bins charge more again, by 17.8% at fifty bins, because the bias is a sum of reciprocal bin counts and equal-width bins are unequal in count.
Fig. 4 The reliability term reported for a forecaster whose true reliability is exactly zero, at four bin counts, counted over a thousand records of five hundred forecasts and against the closed form K·E[f(1 − f)]/n. It is not zero at any binning and it grows in proportion to the number of bins.

Counted against computed: 0.001429 against 0.001415 at five bins, 0.002857 against 0.002830 at ten, 0.005617 against 0.005659 at twenty, 0.014100 against 0.014148 at fifty. Every gap is inside the standard error of the count, and the two routes share nothing — one is a mean over a thousand simulated records and the other is an integral over the forecaster’s population law.

The per-bin figures are the cleaner statement of the same thing. Dividing each count by its own bin count gives 2.86·10⁻⁴, 2.86·10⁻⁴, 2.81·10⁻⁴ and 2.82·10⁻⁴, against a closed form of E[f(1 − f)]/n = 2.83·10⁻⁴. One constant, four binnings, no trend: the bias per bin is a property of the forecaster and the record length, and the bin count only says how many of them to add up.

Every cell of that sweep is read on the same thousand records. That is not only cheaper than drawing four sets; it is the comparison the argument actually needs. Five bins against fifty on one forecaster’s own record is a paired comparison, and the alternative — four independent sweeps — would put a difference between binnings and a difference between draws into the same number.

Equal width is not equal count, and it costs more

The other way of drawing a reliability diagram cuts the probability scale into equal widths rather than the record into equal counts, and it is at least as common. It costs more, at every bin count.

Equal-width bins report 0.001591, 0.003268, 0.006544 and 0.016604 at five, ten, twenty and fifty bins, against the equal-count figures above: dearer by 11.3%, 14.4%, 16.5% and 17.8%. The reason is in the closed form and is arithmetic rather than empirical. The bias is a sum over bins of terms in 1/n_k, that sum is minimised for a fixed total when the n_k are all equal, and a forecaster’s reports are not uniformly distributed — they pile up where its signal puts them, so an equal-width rule produces some crowded bins and some nearly empty ones. The nearly empty ones dominate the sum.

Equal-width bins are drawn anyway, and for a reason that is not carelessness: the horizontal axis of an equal-width diagram is the probability scale itself, so the picture can be read against a fixed grid and two forecasters’ diagrams can be laid over each other. An equal-count diagram’s bins move when the forecaster does, which makes it the better estimator and the worse illustration. Both properties are real and they point in opposite directions, which is the ordinary situation whenever a statistic and a picture are asked to be the same object.

That is a second dial, and it points the same way as the first. Two analysts drawing a reliability diagram of the same record, one with ten equal-count bins and one with fifty equal-width bins, will report reliability terms differing by a factor of 11.6, with no disagreement about a single forecast or a single outcome between them. The rule for constructing the picture is part of the result in exactly the sense a stopping rule is part of a p-value: nothing in the data changed, and the number did.

The correction is computable before an outcome is looked at

The useful part of the closed form is what goes into it. The bias depends on Σ f_i(1 − f_i) inside each bin — on the forecasts alone. Nothing about the outcomes enters it. So the correction can be computed the moment the forecasts are issued, before anybody knows what happened, and subtracted from whatever the diagram later reports.

The estimator is the plain one: subtract from the reported reliability the quantity

REL^0  =  1nk1nkikfi(1fi)\widehat{\mathrm{REL}}_0 \;=\; \frac{1}{n}\sum_k \frac{1}{n_k}\sum_{i \in k} f_i(1 - f_i)

which is the sampling term evaluated on the record in hand rather than in population, so it adapts to whatever bins the forecasts actually fell into.

What is left once the sampling is taken out. The reliability an honest forecaster's diagram reports, and the same figure with the sampling term subtracted, over 1000 samples of 500 forecasts. The correction is not a fudge: within a bin the observed rate is an average of Bernoullis whose variance is Σ f(1 − f)/n_k², and that quantity is computable from the forecasts alone before any outcome is looked at. Subtracting it leaves 1.6e-5, 3.1e-5, -3.7e-5, -3.4e-5 at 5, 10, 20, 50 bins, each inside its own standard error of zero — which is what a forecaster with no miscalibration in it should read. The raw figures it corrects are 0.001429, 0.002857, 0.005617, 0.014100.
Fig. 5 The reliability an honest forecaster’s diagram reports, and the same figure with the sampling term subtracted, at four bin counts over a thousand records. What is left is inside its own standard error of zero at every binning, which is what a forecaster with no miscalibration in it should read.

What is left reads 1.6·10⁻⁵ ± 3.5·10⁻⁵ at five bins, 3.1·10⁻⁵ ± 4.8·10⁻⁵ at ten, −3.7·10⁻⁵ ± 6.6·10⁻⁵ at twenty and −3.4·10⁻⁵ ± 1.0·10⁻⁴ at fifty. Four estimates, two positive and two negative, every one inside its own standard error of zero, across a range of bin counts over which the uncorrected figure moved by a factor of ten. The correction is not a fudge factor tuned until the answer came out right; it was written down from the construction and it happens to work.

The shape of the repair is worth recognising, because it recurs. An estimator built by substituting sample quantities into a population formula inherits whatever the substitution costs, and the cost is usually available in closed form and usually ignored — it is the same reason a plug-in estimate of a variance is not the variance of the thing it is plugged into, and the same reason an estimate at the maximum of a fitted surface is not an estimate of the maximum. Here the substitution is replacing a bin’s true event rate with the rate observed in it, the cost is that rate’s own sampling variance, and the unusual part is only that the cost depends on nothing the outcomes tell.

It does not make the diagram bin-count-free, and the standard errors say why. The corrected estimate at fifty bins has three times the standard error of the one at five, because subtracting a bias does nothing about the variance that produced it. What the correction buys is that the number stops depending systematically on a choice nobody reports; what it does not buy is precision.

A forecaster that is genuinely wrong is also read through the same lens

Everything so far is about a forecaster with nothing wrong with it. The same arithmetic applies with a real distortion underneath, and it applies in a direction worth stating because it is the opposite one.

One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 20 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster pushes its probabilities towards the ends by a factor of 1.8, so its population curve is genuinely off the diagonal and the points scatter round it. The sample's reliability term reads 0.015853 and its expected calibration error 0.0954, against a true reliability of 0.008406. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.
Fig. 6 A record from a forecaster that pushes its probabilities towards the ends by a factor of 1.8, in twenty equal-count bins, over the curve it genuinely has in population. The population curve is off the diagonal in a systematic S: below it at the low end, above it at the high end.

Its population reliability is 0.008500, and the diagram of it at ten bins does not report that even with the sampling term removed. The population diagram — the curve with no sampling error at all, computed by integrating between the bins’ own boundaries — has a reliability of 0.008124. The missing 4% is the within-bin variation in the true gap, averaged away when each bin was summarised by two numbers.

So the binning does two things at once and they push in opposite directions. It adds a sampling bias proportional to K, which makes an honest forecaster look miscalibrated. And it removes part of a real distortion, which makes a miscalibrated forecaster look better than it is — by more at coarse binnings, since a coarse bin averages over more of the curve. A diagram at ten bins is overstating a blameless forecaster’s error by 0.002857 and understating a distorted one’s by 0.000376 in the same picture.

The largest bin-level gap in that population curve is 0.153535, and the curve is a clean S crossing the diagonal once — the signature of a monotone distortion with its mean held at the base rate. It is worth knowing what that looks like, because it is what a diagram is for, and because the scatter in the fifty-bin picture of an honest forecaster above is not it.

The expected calibration error does not escape the dial

Reliability is one summary of a diagram and the expected calibration error is the other: the count-weighted mean absolute distance between a bin’s forecast and its rate, rather than the mean squared one. It is the number reported when a single figure is wanted, and it inherits the problem in a slightly worse form.

For the same blameless forecaster it reads 0.02849, 0.03989, 0.05574 and 0.08839 at five, ten, twenty and fifty bins. That is a factor of 3.10 across a factor of ten in bins, which is √10 to two figures — as it must be, since an absolute deviation scales as a standard deviation and the standard deviation of a bin’s rate goes as √(K/n). The squared summary grows as K and the absolute one as √K, so the second looks tamer and is not: it is the same dial with a square root over it, and it has no correction as clean as the reliability term’s because the expectation of an absolute value does not decompose the way the expectation of a square does.

One consequence follows immediately and is worth separating from the arithmetic. A summary that grows with a construction parameter cannot be compared across records drawn differently, so two published calibration errors are comparable only if both papers said how many bins they used and used the same rule to make them. That is a stronger requirement than it sounds, and it is not usually met. The comparison people actually want — is this forecaster better calibrated than that one — is being made between two numbers that each contain a term nobody reported, in the same way that twenty intervals producing one expected miss is a statement about a procedure rather than about any interval in the set.

The number 0.02 circulates as a threshold for the expected calibration error — a figure below which a forecaster is taken to be well calibrated. Every one of the four figures above is over it, from a forecaster with nothing wrong with it, on five hundred forecasts. What that costs an honest forecaster over a range of record lengths is a measurement in its own right, and it is worse than five hundred forecasts makes it look.

And the decomposition’s residual moves with the same dial

The bin count is also what sets the fourth term in the score’s decomposition, which is where this essay meets the identity checked in three terms. The three-term statement — score equals reliability minus resolution plus uncertainty — is out by 0.004125 at five bins and 0.000211 at fifty, and exact to 2.6·10⁻¹⁵ on the partition that gives every distinct forecast value its own bin.

The identity holds where the diagram is unreadable. How far the Brier score sits from reliability minus resolution plus uncertainty, averaged over 1000 samples of 500 forecasts from an honest forecaster. On the partition that gives every distinct forecast value its own bin the three terms add up exactly — 2.61e-15, which is the last bit a double carries — and the decomposition is also empty there, because reliability becomes the whole score and resolution cancels uncertainty. On 5 equal-count bins the same three terms are out by 0.00412, and the gap closes as the bins get finer: 0.00412, 0.00128, 0.00055, 0.00021 at 5, 10, 20, 50 bins. What is missing is the within-bin term, and adding it makes the statement an identity again at every binning, to 3.61e-16.
Fig. 7 The gap between the Brier score and its three named terms, at four bin counts and on the partition by distinct forecast value. It is the within-bin term, it falls as the bins get finer, and it reaches machine precision exactly where the diagram has one forecast per bin.

Reading the two dials together is the point. The bias in the reliability term rises with K and the residual in the decomposition falls with K, so there is no bin count at which a diagram is both an honest estimate and an exact accounting. Fine bins buy an identity and charge a blameless forecaster 2.8·10⁻⁴ per bin for it; coarse bins buy a low charge and lose 0.004125 of the score into a term the picture does not show. The two failures are not two aspects of one mistake; they are a trade, and the bin count is what trades them.

What is not answered here, deliberately

The obvious next question is which bin count to use, and it is not answered. The machinery to answer it is present — the bias is known in closed form at every K, the variance is counted at every K, and a distorted forecaster’s detectability could be swept the same way — so leaving it out is a decision rather than an omission.

The reason is that a recommended bin count would undo the finding. The claim of this essay is that the number read off a reliability diagram is not a property of the forecaster; a rule saying “use ten bins” would restore exactly the impression that it is, by making one reading canonical and the others mistakes. What a record supports instead is reporting the binning alongside the number, or reporting the debiased figure, or both. The habit that would fix it is the one the essay on what a p-value does not say asks for in a different setting: a statistic quoted without the construction that produced it is not a smaller claim than the full one, it is a different and unfalsifiable claim. A bias–variance optimum over K exists and would be worth measuring; it would answer a different question, which is how best to detect a distortion of a stated size, and that question has the distortion’s size in it as an input.

What is claimed is narrower and is checked in both directions. A forecaster with no miscalibration in it reads positive at every binning, in proportion to the bin count, at a rate the closed form predicts to within the standard error of a thousand records; equal-width bins cost 11 to 18% more than equal-count ones; and the correction, computed from the forecasts before any outcome is seen, returns an estimate inside its own standard error of zero at every binning tried. The assertion that would catch this essay being wrong is the last one, and it is the one that has to hold at all four bin counts rather than at a convenient one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Base rateBinningBrier scoreDebiased estimatorEstimation errorExpected calibration errorForecast calibrationMurphy's decompositionProbability forecastQuadratureReliability diagramReliability termSampling variationUncertainty termWithin-bin variance