The miscalibration a perfect forecaster shows
Worth reading first: An identity in three terms · What the model says next.
A forecaster reporting the true probability of an event given everything it knows has a reliability of exactly zero. Nothing about it is miscalibrated, there is nothing to fix, and any diagnostic run on it should read nothing.
On a record of fifty forecasts its expected calibration error reads 0.1252. On a hundred, 0.0884. On five hundred, 0.0403; on two thousand, 0.0203; on ten thousand, 0.0090. Those are means over 1,200, 1,200, 700, 250 and 100 independent records, and the whole of every one of them is the arithmetic of averaging a bin’s Bernoulli outcomes. The forecaster is not wrong at any length. The instrument is.
The figure that circulates as a threshold — a calibration error under 0.02, below which a forecaster is taken to be well calibrated — is therefore not cleared by an honest record until about 1,976 forecasts, and not cleared with any reliability until about 4,111. At a hundred forecasts, 100% of 1,200 blameless records fail it. Not most of them. All of them, at both of the two shortest lengths measured.
What a hundred honest forecasts look like
A calibration error is a single number summarising a reliability diagram: the count-weighted mean distance between each bin’s mean forecast and the share of that bin’s events that happened. Ten equal-count bins throughout what follows, which is the ordinary choice, and the bin count is a dial the number moves with rather than a detail.
The distribution has no mass anywhere near zero, and that is the first thing to notice about it. A calibration error is a mean of absolute values, so it cannot be negative and cannot cancel: every bin contributes its own sampling departure with a positive sign, and ten of them are added up. A quantity built that way has a positive expectation under a null that is exactly true, and no amount of averaging over records moves it, because averaging is what produced the number in the first place.
It is also a distribution rather than a number, which is why 1,200 records were drawn at this length rather than a summary computed once. A single record’s calibration error is one draw from the picture above, and a claim about what an honest forecaster’s diagram strays to cannot be made from a mean: the mean at a hundred forecasts is 0.0884 and the record at the 95th percentile reads 0.1269, and it is the second number that decides whether a threshold is cleared by anybody rather than on average.
Its 95th percentile at this length is 0.1269 and the worst single bin of the average record is off by 0.233. A reader handed one such diagram — a bin where the forecaster said 0.7 and the events happened about 47% of the time — would have a story about overconfidence, and the story would be about ten Bernoulli draws.
This is the same shape as twenty intervals producing one expected miss: a procedure behaving exactly as advertised, generating an observation that reads as a failure of it. The difference is that a miss is a discrete event that somebody has counted the expected number of, and a calibration error is a continuous number with no advertised value at all — so there is nothing to compare the observation against, and the comparison people actually make is against zero.
The closed form, and the constant it hides
The expectation is not something to be simulated and then trusted. A bin’s observed rate is a mean of n/K Bernoulli outcomes, its standard deviation is therefore √(v_k K/n) with v_k the mean of f(1 − f) inside the bin, and the mean absolute deviation of something nearly normal is √(2/π) of its standard deviation. Adding those over K bins with weights 1/K,
which falls as 1/√n, rises as √K, and is never zero.
The mean root bin variance is 0.352391, integrated over the ten bins’ own boundaries, so the closed form gives 0.12574, 0.08891, 0.03976, 0.01988 and 0.00889 at the five lengths. Counted against computed, the worst gap anywhere is 2.3%, at n = 2,000 where the count is over 250 records and its own sampling noise is largest. The two routes share nothing: one is a mean over simulated records and the other is an integral over the forecaster’s population law with a normal approximation to the mean absolute deviation on top.
The cleaner way to read the same agreement is as one constant. Multiplying each counted mean by √n gives 0.8853, 0.8841, 0.9006, 0.9096 and 0.8955 across a factor of two hundred in record length, against a closed form of 0.8891. One number, five lengths, no trend — which is the 1/√n law stated as a measurement rather than fitted to one.
That is the sense in which a calibration error is not interpretable on its own. It is approximately 0.889/√n times a quantity that depends on the forecaster’s own sharpness and on the bin count, and none of those three things is usually reported beside it. A calibration error quoted without its record length is a number whose largest component has been omitted — the same structural complaint a p-value quoted without its sample size attracts, arriving in a setting where nobody has thought to make it.
The threshold, and the record it needs
Setting the closed form equal to 0.02 and solving gives the record length at which an honest forecaster’s mean calibration error finally falls under the threshold. It is 1,976 forecasts. Solved, not searched: the expression is monotone in n and inverts directly.
The mean is the generous reading. Half of all honest records at 1,976 forecasts are above the threshold by construction, so a forecaster wanting to be confident of clearing it needs the 95th percentile under 0.02 instead, which the same law puts at about 4,111 forecasts. The share of blameless records failing the threshold runs 100%, 100%, 98.9%, 50.0% and 0.0% across the five lengths — that fourth figure being the definitional half at a length close to 1,976, and the fifth being what a threshold looks like once the record is long enough for it to mean something.
The gap between those two record lengths is itself worth reading. A factor of 2.08 in n separates the length at which half of honest records clear the threshold from the length at which nineteen in twenty do, and under a 1/√n law that factor has to be the square of the ratio between the 95th percentile and the mean. That ratio is 0.1789 against 0.1252 at fifty forecasts and 0.0576 against 0.0403 at five hundred — about 1.43 either way, and 1.43² is 2.04 against the 2.08 the two crossings give. So the distribution’s shape is close to constant across lengths and only its scale moves, which is what makes a single crossing point meaningful at all. A quantity whose shape changed with n would not have one.
Two thousand daily forecasts is five and a half years. Ten thousand is twenty-seven. Very few forecasting records anybody actually evaluates are in the range where a fixed threshold on a calibration error separates an honest forecaster from a distorted one, and every record shorter than that produces a number above the threshold whatever the forecaster does. The threshold is not too strict or too lenient; it is a threshold on a quantity whose scale is set by something the threshold does not mention.
At two thousand forecasts the worst bin of the average record is still off by 0.054, which is a bin where the forecaster said 0.35 and the events happened about 30% of the time. The distribution has narrowed and it has not moved anywhere near zero, and it never will: there is no record length at which an honest forecaster’s diagram is flat, only lengths at which the departure is smaller than whatever else is being argued about.
The squared summary is the same story, and it has a correction
The reliability term — the mean squared bin gap rather than the mean absolute one — behaves the same way and is worth reading beside it, because unlike the calibration error it has a correction that works.
For the same honest forecaster it reads 0.027742, 0.013795, 0.002894, 0.000727 and 0.000141 at the five lengths, against a closed form of K·E[f(1 − f)]/n which gives 0.028296, 0.014148, 0.002830, 0.000707 and 0.000141. Falling as 1/n rather than as 1/√n, because a square of a sampling departure is what it is, and never zero for the same reason.
The correction is computable from the forecasts alone, before any outcome is looked at, and subtracting it leaves an estimate inside its own standard error of zero at every bin count. Nothing equivalent exists for the calibration error, and the reason is exactly the absolute value: the expectation of a square decomposes into a bias plus a true value and the expectation of an absolute value does not, so the sampling term cannot be subtracted off, only approximated and then subtracted approximately. The summary that is easier to read is the one that cannot be repaired.
There is a second reason to prefer the squared summary, and it is about what happens when a real distortion is present. The correction subtracts a quantity that depends only on the forecasts, so it is the same subtraction whether the forecaster is honest or not: a distorted forecaster’s corrected reliability is an estimate of its true reliability rather than of zero. The calibration error has no such split available. Its expectation under a distortion is not the true error plus a sampling term — absolute values of a signal and a noise do not add — so even an approximate correction would be approximating a different quantity at every distortion size. That is the difference between a bias and an offset, and only one of the two can be taken off.
That is an argument for reporting the reliability term with its correction rather than the calibration error with a threshold, and it is worth being clear that it is an argument about instruments and not about forecasters. Both numbers are describing the same diagram. One of them has a known bias with a closed form and the other has a known bias without an estimator.
A test with the right reference is wrong in both directions at once
The disciplined answer to a biased summary is to stop reading it against zero and start reading it against its own null distribution. That is available here and it is honest arithmetic: under the null that each f_i is the true probability, each y_i is Bernoulli(f_i) independently, so within a bin the sum of y_i − f_i has mean zero and variance Σ f_i(1 − f_i), and
is a sum of K nearly-standard-normal squares to be read against a χ² on K degrees of freedom. No parameter has been fitted, which is what separates this from the same statistic computed after a logistic regression, and is why the degrees of freedom are K rather than K − 2.
The reference is right asymptotically and it is not right at fifty forecasts. Against a critical value of 18.3070 its error rate under a true null — the rate at which it rejects an honest forecaster — is 7.7% at a nominal 5%, then 7.2%, 6.4%, 4.8% and 4.0% as the record grows. The mean statistic under the null is 9.9, 9.9, 10.1, 10.4 and 9.7 against ten degrees of freedom, so the reference distribution’s centre is right at every length — what is wrong at short lengths is its upper tail, which is where the test is read. That is the ordinary situation and not a surprise: an approximation is worst exactly where it is used, and a mean that lands on its target says nothing about a 95th percentile.
The other half is worse. Against a real distortion — the same honest posterior pushed towards the ends by a factor of 1.25, with its mean report held at the base rate so that only the shape of its calibration curve is wrong — the test’s power is 27.7% at fifty forecasts and 29.3% at a hundred. It reaches 62.9% at five hundred and 100.0% at two thousand. The mean statistic under that alternative runs 17.3, 19.1, 25.6, 56.0 and 206.7, so the signal is there and growing at every length; what is missing at fifty forecasts is enough of it to clear a critical value set for ten degrees of freedom.
Reading the mean statistic beside the rejection rate is what makes that diagnosis available, and it is the same move as reading a whole distribution of p-values rather than counting how many fell under 0.05 — a rejection rate cannot tell a reference distribution that is wrong in the tail from one that is wrong everywhere, and a mean of 9.9 against ten degrees of freedom rules out the second immediately. What is left is a tail that is too heavy at short lengths, which is what a sum of ten squared quantities each built from five Bernoulli draws should be.
So at fifty forecasts the check rejects half again as often as it should when there is nothing there, and misses seven distortions in ten when there is. A short record is a record the check is wrong about in both directions at once, which is a specific and unusual failure — a test with the right size and no power has learned to say nothing, and a test with power and the wrong size has learned to say everything, and this one has neither excuse. Both halves are asserted together for that reason: measuring only the size would have made the test look mildly conservative to fix, and measuring only the power would have made it look merely underpowered.
What this is not: a forecaster that reads nothing
One reading has to be blocked, because everything above pushes towards it. The finding is not that calibration checks are useless and long records are the answer, and the evidence against that reading is a forecaster this field has already measured.
A forecaster that issues the base rate every time has a reliability of zero in population and, at every record length measured here, the smallest apparent calibration error of anything in this world — because its forecasts have no spread, its bins collapse, and there is nothing for sampling to scatter. It is also worth nothing: its resolution is exactly 0 and its score is 0.234237, which is the uncertainty term. So the instrument’s bias is not merely noise added to a good measurement; it rewards hedging, since a forecaster that says less has less to be caught being wrong about.
That is the sting in this essay and it is why it sits beside the measurement that a calibration check passes a table of base rates. The two failures compound: the check cannot distinguish a useful forecaster from a useless one in population, and on a finite record it actively prefers the useless one. A forecaster optimising for a calibration diagnostic on a short record has an available strategy, and it is the strategy that destroys the thing the forecast was for.
What is not measured, and it would change the numbers
Every outcome above is independent given its forecast. That is what makes the χ² reference the right one and what makes the closed form for the calibration error a sum of independent bin variances, and it is exactly the assumption a real forecasting record breaks. Neighbouring days are correlated; whatever made yesterday’s event likely is still there today.
The direction is not in doubt and the size is not measured here. Dependence inflates the variance of every bin’s observed rate, so the apparent calibration error grows and the test’s size grows with it — which is the same arithmetic that makes fifty correlated observations worth about six, applied to a bin rather than to a series. Turning that into numbers needs a dependence structure stated and defended rather than assumed, since the answer is a function of it, and nothing here does that. It is named as the next measurement rather than gestured at, because a correction of unstated size is worse than an honest independent-case figure with a warning attached.
Two smaller limits belong beside it. Everything is at ten equal-count bins, and the calibration error rises as √K, so every figure in this essay is a statement about a diagram drawn one particular way — which is the point the field’s second essay makes and is not repaired by any of this. And the power figures are against one alternative, a confidence multiplier of 1.25 with the mean report held fixed; a distortion of a different shape, or one that moves the level as well, would be found at a different rate, and the 27.7% at fifty forecasts should be read as one alternative’s number rather than as the test’s general ability to see anything.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A forecaster that rounds — both name probability forecast, reliability diagram, reliability term
- A taper and a critical value — both name critical value, error rate, statistical power
- An order that spends the error rate — both name error rate, null hypothesis, statistical power
- Estimating how many nulls are true — both name error rate, null hypothesis, statistical power
- How long a block a multiplier shares — both name critical value, error rate, null hypothesis
- The analysis after three arms — both name error rate, null hypothesis, statistical power
Named objects
A flat tag is an object no other essay names yet.
BinningCalibration testChi-squareClimatological forecastCritical valueDebiased estimatorError rateExpected calibration errorForecast calibrationNull hypothesisProbability forecastReliability diagramReliability termSampling variationStatistical power