A forecast that is a probability

The miscalibration a perfect forecaster shows

A forecaster whose true reliability is exactly zero shows a calibration error of 0.1252 on fifty forecasts and 0.0090 on ten thousand. Every one of 1,200 blameless hundred-forecast records exceeds the 0.02 routinely read as evidence of a problem, and the mean does not fall under it until 1,976 forecasts.

Worth reading first: An identity in three terms · What the model says next.

A forecaster reporting the true probability of an event given everything it knows has a reliability of exactly zero. Nothing about it is miscalibrated, there is nothing to fix, and any diagnostic run on it should read nothing.

On a record of fifty forecasts its expected calibration error reads 0.1252. On a hundred, 0.0884. On five hundred, 0.0403; on two thousand, 0.0203; on ten thousand, 0.0090. Those are means over 1,200, 1,200, 700, 250 and 100 independent records, and the whole of every one of them is the arithmetic of averaging a bin’s Bernoulli outcomes. The forecaster is not wrong at any length. The instrument is.

The figure that circulates as a threshold — a calibration error under 0.02, below which a forecaster is taken to be well calibrated — is therefore not cleared by an honest record until about 1,976 forecasts, and not cleared with any reliability until about 4,111. At a hundred forecasts, 100% of 1,200 blameless records fail it. Not most of them. All of them, at both of the two shortest lengths measured.

What a hundred honest forecasts look like

A calibration error is a single number summarising a reliability diagram: the count-weighted mean distance between each bin’s mean forecast and the share of that bin’s events that happened. Ten equal-count bins throughout what follows, which is the ordinary choice, and the bin count is a dial the number moves with rather than a detail.

What 100 honest forecasts look like. The calibration error of 1200 separate records of 100 forecasts, every one of them from a forecaster whose true reliability is exactly zero. Not one record reads zero. The distribution is centred on 0.0884, its median is 0.0870 and its 95th percentile is 0.1269, so 100.0% of these blameless records show more than the 0.02 that is routinely read as evidence of a problem. The worst bin of the average record is off by 0.2329. The closed form for the mean is √(2K/πn) times the mean root bin variance, which gives 0.0889.
Fig. 1 The calibration error of 1,200 separate records of a hundred forecasts, every one of them from a forecaster whose true reliability is exactly zero. Not one record reads zero. The distribution is centred on 0.0884 and its 95th percentile is 0.1269.

The distribution has no mass anywhere near zero, and that is the first thing to notice about it. A calibration error is a mean of absolute values, so it cannot be negative and cannot cancel: every bin contributes its own sampling departure with a positive sign, and ten of them are added up. A quantity built that way has a positive expectation under a null that is exactly true, and no amount of averaging over records moves it, because averaging is what produced the number in the first place.

It is also a distribution rather than a number, which is why 1,200 records were drawn at this length rather than a summary computed once. A single record’s calibration error is one draw from the picture above, and a claim about what an honest forecaster’s diagram strays to cannot be made from a mean: the mean at a hundred forecasts is 0.0884 and the record at the 95th percentile reads 0.1269, and it is the second number that decides whether a threshold is cleared by anybody rather than on average.

Its 95th percentile at this length is 0.1269 and the worst single bin of the average record is off by 0.233. A reader handed one such diagram — a bin where the forecaster said 0.7 and the events happened about 47% of the time — would have a story about overconfidence, and the story would be about ten Bernoulli draws.

This is the same shape as twenty intervals producing one expected miss: a procedure behaving exactly as advertised, generating an observation that reads as a failure of it. The difference is that a miss is a discrete event that somebody has counted the expected number of, and a calibration error is a continuous number with no advertised value at all — so there is nothing to compare the observation against, and the comparison people actually make is against zero.

The closed form, and the constant it hides

The expectation is not something to be simulated and then trusted. A bin’s observed rate is a mean of n/K Bernoulli outcomes, its standard deviation is therefore √(v_k K/n) with v_k the mean of f(1 − f) inside the bin, and the mean absolute deviation of something nearly normal is √(2/π) of its standard deviation. Adding those over K bins with weights 1/K,

E[ECE]    2Kπn1Kkvk\mathbb{E}[\mathrm{ECE}] \;\approx\; \sqrt{\frac{2K}{\pi n}} \cdot \frac{1}{K}\sum_k \sqrt{v_k}

which falls as 1/√n, rises as √K, and is never zero.

How long an honest forecaster looks broken for. The calibration error shown by a forecaster with no miscalibration in it at all, at five record lengths, drawn against one over the square root of the length so that the closed form is a straight line through the origin. It is 0.1252 at fifty forecasts and 0.0090 at ten thousand, against a closed form of √(2K/πn) times the mean root bin variance which gives 0.1257 and 0.0089. The threshold drawn across it is 0.02, a figure routinely read as evidence that something is wrong; the mean falls under it at 1976 forecasts and the 95th percentile at about 4111. Below that, an honest forecaster and a miscalibrated one are being told apart by a statistic that is mostly the sample size.
Fig. 2 The calibration error shown by a forecaster with no miscalibration in it at all, at five record lengths, drawn against one over the square root of the length so that the closed form is a straight line through the origin. The threshold drawn across it is 0.02.

The mean root bin variance is 0.352391, integrated over the ten bins’ own boundaries, so the closed form gives 0.12574, 0.08891, 0.03976, 0.01988 and 0.00889 at the five lengths. Counted against computed, the worst gap anywhere is 2.3%, at n = 2,000 where the count is over 250 records and its own sampling noise is largest. The two routes share nothing: one is a mean over simulated records and the other is an integral over the forecaster’s population law with a normal approximation to the mean absolute deviation on top.

The cleaner way to read the same agreement is as one constant. Multiplying each counted mean by √n gives 0.8853, 0.8841, 0.9006, 0.9096 and 0.8955 across a factor of two hundred in record length, against a closed form of 0.8891. One number, five lengths, no trend — which is the 1/√n law stated as a measurement rather than fitted to one.

That is the sense in which a calibration error is not interpretable on its own. It is approximately 0.889/√n times a quantity that depends on the forecaster’s own sharpness and on the bin count, and none of those three things is usually reported beside it. A calibration error quoted without its record length is a number whose largest component has been omitted — the same structural complaint a p-value quoted without its sample size attracts, arriving in a setting where nobody has thought to make it.

The threshold, and the record it needs

Setting the closed form equal to 0.02 and solving gives the record length at which an honest forecaster’s mean calibration error finally falls under the threshold. It is 1,976 forecasts. Solved, not searched: the expression is monotone in n and inverts directly.

One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.
Fig. 3 One record of five hundred forecasts from the honest forecaster, in ten equal-count bins, drawn over its population curve. The population curve is the diagonal exactly, so every departure the points show is sampling, and the record’s calibration error is 0.0403 on average across a thousand such records.

The mean is the generous reading. Half of all honest records at 1,976 forecasts are above the threshold by construction, so a forecaster wanting to be confident of clearing it needs the 95th percentile under 0.02 instead, which the same law puts at about 4,111 forecasts. The share of blameless records failing the threshold runs 100%, 100%, 98.9%, 50.0% and 0.0% across the five lengths — that fourth figure being the definitional half at a length close to 1,976, and the fifth being what a threshold looks like once the record is long enough for it to mean something.

The gap between those two record lengths is itself worth reading. A factor of 2.08 in n separates the length at which half of honest records clear the threshold from the length at which nineteen in twenty do, and under a 1/√n law that factor has to be the square of the ratio between the 95th percentile and the mean. That ratio is 0.1789 against 0.1252 at fifty forecasts and 0.0576 against 0.0403 at five hundred — about 1.43 either way, and 1.43² is 2.04 against the 2.08 the two crossings give. So the distribution’s shape is close to constant across lengths and only its scale moves, which is what makes a single crossing point meaningful at all. A quantity whose shape changed with n would not have one.

Two thousand daily forecasts is five and a half years. Ten thousand is twenty-seven. Very few forecasting records anybody actually evaluates are in the range where a fixed threshold on a calibration error separates an honest forecaster from a distorted one, and every record shorter than that produces a number above the threshold whatever the forecaster does. The threshold is not too strict or too lenient; it is a threshold on a quantity whose scale is set by something the threshold does not mention.

What 2000 honest forecasts look like. The calibration error of 250 separate records of 2000 forecasts, every one of them from a forecaster whose true reliability is exactly zero. Not one record reads zero. The distribution is centred on 0.0203, its median is 0.0199 and its 95th percentile is 0.0298, so 50.0% of these blameless records show more than the 0.02 that is routinely read as evidence of a problem. The worst bin of the average record is off by 0.0544. The closed form for the mean is √(2K/πn) times the mean root bin variance, which gives 0.0199.
Fig. 4 The same measurement at two thousand forecasts, over 250 records. The distribution has moved down by a factor of more than four and has not moved to zero: its centre is 0.0203 and its 95th percentile 0.0298.

At two thousand forecasts the worst bin of the average record is still off by 0.054, which is a bin where the forecaster said 0.35 and the events happened about 30% of the time. The distribution has narrowed and it has not moved anywhere near zero, and it never will: there is no record length at which an honest forecaster’s diagram is flat, only lengths at which the departure is smaller than whatever else is being argued about.

The squared summary is the same story, and it has a correction

The reliability term — the mean squared bin gap rather than the mean absolute one — behaves the same way and is worth reading beside it, because unlike the calibration error it has a correction that works.

For the same honest forecaster it reads 0.027742, 0.013795, 0.002894, 0.000727 and 0.000141 at the five lengths, against a closed form of K·E[f(1 − f)]/n which gives 0.028296, 0.014148, 0.002830, 0.000707 and 0.000141. Falling as 1/n rather than as 1/√n, because a square of a sampling departure is what it is, and never zero for the same reason.

What is left once the sampling is taken out. The reliability an honest forecaster's diagram reports, and the same figure with the sampling term subtracted, over 1000 samples of 500 forecasts. The correction is not a fudge: within a bin the observed rate is an average of Bernoullis whose variance is Σ f(1 − f)/n_k², and that quantity is computable from the forecasts alone before any outcome is looked at. Subtracting it leaves 1.6e-5, 3.1e-5, -3.7e-5, -3.4e-5 at 5, 10, 20, 50 bins, each inside its own standard error of zero — which is what a forecaster with no miscalibration in it should read. The raw figures it corrects are 0.001429, 0.002857, 0.005617, 0.014100.
Fig. 5 The reliability an honest forecaster’s diagram reports at four bin counts, and the same figure with the sampling term subtracted. What is left is inside its own standard error of zero at every binning.

The correction is computable from the forecasts alone, before any outcome is looked at, and subtracting it leaves an estimate inside its own standard error of zero at every bin count. Nothing equivalent exists for the calibration error, and the reason is exactly the absolute value: the expectation of a square decomposes into a bias plus a true value and the expectation of an absolute value does not, so the sampling term cannot be subtracted off, only approximated and then subtracted approximately. The summary that is easier to read is the one that cannot be repaired.

There is a second reason to prefer the squared summary, and it is about what happens when a real distortion is present. The correction subtracts a quantity that depends only on the forecasts, so it is the same subtraction whether the forecaster is honest or not: a distorted forecaster’s corrected reliability is an estimate of its true reliability rather than of zero. The calibration error has no such split available. Its expectation under a distortion is not the true error plus a sampling term — absolute values of a signal and a noise do not add — so even an approximate correction would be approximating a different quantity at every distortion size. That is the difference between a bias and an offset, and only one of the two can be taken off.

That is an argument for reporting the reliability term with its correction rather than the calibration error with a threshold, and it is worth being clear that it is an argument about instruments and not about forecasters. Both numbers are describing the same diagram. One of them has a known bias with a closed form and the other has a known bias without an estimator.

A test with the right reference is wrong in both directions at once

The disciplined answer to a biased summary is to stop reading it against zero and start reading it against its own null distribution. That is available here and it is honest arithmetic: under the null that each f_i is the true probability, each y_i is Bernoulli(f_i) independently, so within a bin the sum of y_i − f_i has mean zero and variance Σ f_i(1 − f_i), and

X2  =  k(ik(yifi))2ikfi(1fi)X^2 \;=\; \sum_k \frac{\big(\sum_{i \in k}(y_i - f_i)\big)^2}{\sum_{i \in k} f_i(1 - f_i)}

is a sum of K nearly-standard-normal squares to be read against a χ² on K degrees of freedom. No parameter has been fitted, which is what separates this from the same statistic computed after a logistic regression, and is why the degrees of freedom are K rather than K − 2.

What a calibration test can see, and when. The rate at which a calibration test rejects a forecaster that is calibrated by construction, and one that pushes its probabilities towards the ends by a factor of 1.25, at five sample sizes. The statistic is a sum of 10 standardised bin residuals against a χ² reference with 10 degrees of freedom, and no parameter has been fitted, so the reference is the right one asymptotically. It is not right at fifty forecasts: the test rejects an honest forecaster 7.7% of the time against a nominal 5.0%, and catches the real distortion only 27.7% of the time. Both failures go the same way as the sample grows — 4.0% and 100.0% at ten thousand — so a calibration check on a short record is a check that is wrong in both directions at once.
Fig. 6 The rate at which the calibration test rejects a forecaster that is calibrated by construction, and one that pushes its probabilities towards the ends by a factor of 1.25, at five record lengths. Both failures shrink with the record and neither is gone at five hundred forecasts.

The reference is right asymptotically and it is not right at fifty forecasts. Against a critical value of 18.3070 its error rate under a true null — the rate at which it rejects an honest forecaster — is 7.7% at a nominal 5%, then 7.2%, 6.4%, 4.8% and 4.0% as the record grows. The mean statistic under the null is 9.9, 9.9, 10.1, 10.4 and 9.7 against ten degrees of freedom, so the reference distribution’s centre is right at every length — what is wrong at short lengths is its upper tail, which is where the test is read. That is the ordinary situation and not a surprise: an approximation is worst exactly where it is used, and a mean that lands on its target says nothing about a 95th percentile.

The other half is worse. Against a real distortion — the same honest posterior pushed towards the ends by a factor of 1.25, with its mean report held at the base rate so that only the shape of its calibration curve is wrong — the test’s power is 27.7% at fifty forecasts and 29.3% at a hundred. It reaches 62.9% at five hundred and 100.0% at two thousand. The mean statistic under that alternative runs 17.3, 19.1, 25.6, 56.0 and 206.7, so the signal is there and growing at every length; what is missing at fifty forecasts is enough of it to clear a critical value set for ten degrees of freedom.

Reading the mean statistic beside the rejection rate is what makes that diagnosis available, and it is the same move as reading a whole distribution of p-values rather than counting how many fell under 0.05 — a rejection rate cannot tell a reference distribution that is wrong in the tail from one that is wrong everywhere, and a mean of 9.9 against ten degrees of freedom rules out the second immediately. What is left is a tail that is too heavy at short lengths, which is what a sum of ten squared quantities each built from five Bernoulli draws should be.

So at fifty forecasts the check rejects half again as often as it should when there is nothing there, and misses seven distortions in ten when there is. A short record is a record the check is wrong about in both directions at once, which is a specific and unusual failure — a test with the right size and no power has learned to say nothing, and a test with power and the wrong size has learned to say everything, and this one has neither excuse. Both halves are asserted together for that reason: measuring only the size would have made the test look mildly conservative to fix, and measuring only the power would have made it look merely underpowered.

What this is not: a forecaster that reads nothing

One reading has to be blocked, because everything above pushes towards it. The finding is not that calibration checks are useless and long records are the answer, and the evidence against that reading is a forecaster this field has already measured.

Six forecasters, all calibrated, not equally useful. The resolution of six forecasters that are all perfectly calibrated, each reporting the true probability of the event given a signal that carries more or less of the latent state. The largest reliability anywhere in the family is 2.0e-33, so a calibration check passes every one of them. They are not equally good: resolution runs from exactly 0.00 for the forecaster that issues the base rate every time to 0.092758 for the one that sees everything, and their Brier scores run from 0.234237 — which is the world's own uncertainty, and the score of a table of base rates — to 0.141479. Calibration is a necessary condition that a constant forecast satisfies exactly.
Fig. 7 The resolution of six forecasters that are all perfectly calibrated, running from exactly zero for the one that issues the base rate every time to 0.092758 for the one that sees everything. The largest reliability anywhere in the family is 2·10⁻³³.

A forecaster that issues the base rate every time has a reliability of zero in population and, at every record length measured here, the smallest apparent calibration error of anything in this world — because its forecasts have no spread, its bins collapse, and there is nothing for sampling to scatter. It is also worth nothing: its resolution is exactly 0 and its score is 0.234237, which is the uncertainty term. So the instrument’s bias is not merely noise added to a good measurement; it rewards hedging, since a forecaster that says less has less to be caught being wrong about.

That is the sting in this essay and it is why it sits beside the measurement that a calibration check passes a table of base rates. The two failures compound: the check cannot distinguish a useful forecaster from a useless one in population, and on a finite record it actively prefers the useless one. A forecaster optimising for a calibration diagnostic on a short record has an available strategy, and it is the strategy that destroys the thing the forecast was for.

What is not measured, and it would change the numbers

Every outcome above is independent given its forecast. That is what makes the χ² reference the right one and what makes the closed form for the calibration error a sum of independent bin variances, and it is exactly the assumption a real forecasting record breaks. Neighbouring days are correlated; whatever made yesterday’s event likely is still there today.

The direction is not in doubt and the size is not measured here. Dependence inflates the variance of every bin’s observed rate, so the apparent calibration error grows and the test’s size grows with it — which is the same arithmetic that makes fifty correlated observations worth about six, applied to a bin rather than to a series. Turning that into numbers needs a dependence structure stated and defended rather than assumed, since the answer is a function of it, and nothing here does that. It is named as the next measurement rather than gestured at, because a correction of unstated size is worse than an honest independent-case figure with a warning attached.

Two smaller limits belong beside it. Everything is at ten equal-count bins, and the calibration error rises as √K, so every figure in this essay is a statement about a diagram drawn one particular way — which is the point the field’s second essay makes and is not repaired by any of this. And the power figures are against one alternative, a confidence multiplier of 1.25 with the mean report held fixed; a distortion of a different shape, or one that moves the level as well, would be found at a different rate, and the 27.7% at fifty forecasts should be read as one alternative’s number rather than as the test’s general ability to see anything.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BinningCalibration testChi-squareClimatological forecastCritical valueDebiased estimatorError rateExpected calibration errorForecast calibrationNull hypothesisProbability forecastReliability diagramReliability termSampling variationStatistical power