Five values chosen by the forecasts
Worth reading first: An identity in three terms · What the model says next.
A forecaster that rounds priced three vocabularies that are chosen before anyone looks at the forecasts: a grid of equal steps, bins of equal width, and groups holding equal shares of the population. Each spends a fixed number of printed values, and each, once its printed values are replaced by its groups’ event rates, is calibrated by construction. What separates them is how much of the forecaster’s resolution survives the grouping, and the answer depended on where the cuts fell rather than on how many there were. That leaves the obvious question. Among every way of cutting the probability scale into groups, which loses least?
The question has an exact answer, and the answer is old. It belongs to signal processing rather than forecasting, and the forecasting literature rarely uses it. This essay finds the answer for the same world as before — the honest posterior at full signal, with an event rate of 37.4%, a resolution of 0.092758 and an area under the ROC curve of 0.8683 — and then asks the two questions the answer raises. A forecaster does not know the world, so it has to find the cuts from the forecasts it has issued. And the Brier score is one way of averaging over users, so other users may want other cuts.
The loss a calibrated vocabulary pays is variance inside its groups
Write the honest probability as and the printed value as the event rate of the group falls in. Because each printed value is its group’s event rate, the reliability term on the printed values is zero, and the whole excess Brier score over the honest forecaster is resolution lost. Resolution is the variance of the printed values about the base rate; the honest forecaster’s resolution is the variance of itself; and the difference, by the law of total variance, is the average variance of inside the groups. The identity that splits a score into three terms is doing all the work: a calibrated vocabulary is scored only on what it averages away.
That turns the choice of cuts into a familiar problem. Minimising the expected squared distance between a quantity and the value it is replaced by, over partitions into cells, is the design of a scalar quantiser. Its two necessary conditions were written down by Lloyd and by Max. Each printed value must sit at the mean of its cell, which here means at its group’s event rate, since the honest averages to the event rate inside any group. Each cut must sit halfway between the printed values on either side, so that every forecast goes to the printed value nearer to it. Neither condition alone fixes a vocabulary. Alternating them from any start never increases the loss, and the iteration used here, started from the equal-mass cuts, stops when no cut moves by more than . For five values it takes 180 rounds, and for eleven, 827. Each round is two closed forms, the same normal tails and bivariate orthants that priced the rounding grids, so nothing in the result is counted.
The best five cut at 0.145, 0.339, 0.552 and 0.777, and print 0.054, 0.236, 0.441, 0.662 and 0.892. Every cut is the midpoint of the two printed values beside it, to the eighth decimal, a condition the figure is checked against before it is drawn. Every printed value is its group’s event rate, so nothing about the vocabulary is miscalibrated. It loses 0.003193 of resolution. Equal-width bins lose 0.003454 and equal-mass groups 0.004046, so the best cuts save 7.6% against the bins and 21.1% against equal mass.
Where the best cuts sit, and why
The figure shows what the best cuts are doing. The honest forecasts pile up near zero, because the event is less common than not, and a forecaster with a strong signal says “almost certainly not” to most of the cases where nothing happens. Equal-mass groups follow that pile. Their first cut is at 0.066 and their second at 0.211, so three of the five groups sit below 0.422. The top group then has to stretch from 0.695 to 1 to collect its fifth of the forecasts, and a wide group has a large variance whatever is inside it. Equal-width bins do the opposite. They ignore the pile entirely, put 38.8% of all forecasts into the first bin, and space their cuts at 0.2, 0.4, 0.6 and 0.8 whether anything lands there or not.
The best cuts sit between the two. The first group holds 32.1% of the forecasts, fewer than the bins’ 38.8% and more than the equal-mass 20%, and the shares then fall gently — 20.7%, 17.2%, 15.3% and 14.7%. The cuts are closer together where forecasts are dense, as equal mass would have them. But they are pulled much less hard than equal mass pulls them, because a group’s contribution to the loss is its share times its variance, and halving a group’s width quarters its variance while halving its share only halves it. High-resolution quantisation theory makes that exact. For many cells the best density of cuts is proportional to the cube root of the density being quantised, which grows far more slowly than the density itself. Equal mass puts cuts in proportion to the density, and that is too steep a response for a loss measured in squared distance.
This matters because the two familiar choices are each defensible in isolation. Equal mass is what risk categories usually use, and it ranks well. Equal width is what a grid of round numbers resembles, and it resolves better than equal mass. Calibration says nothing about either, since every vocabulary here is calibrated. The best cuts are the answer for one stated purpose, the Brier score, and they are neither of the conventions.
How the loss falls with the number of values
The same theory says how fast the loss should fall. For a quantiser with many cells the mean squared error scales as , so multiplying the loss by should give a curve that flattens to a constant. That constant depends on the density being quantised and on how the cuts are placed, but not on .
The three curves flatten quickly. The best cuts settle near 0.078: they lose 0.021173 at two values, 0.003193 at five, 0.000782 at ten and 0.000541 at twelve, so each extra value is worth a little less than the one before, in the proportion the square law predicts. Equal-width bins settle near 0.085 and equal-mass groups near 0.101. So, at any number of values, equal mass loses about 30% more than the best cuts and equal width about 9% more. Those ratios barely move between five values and twelve. The ordering is not an accident of one ; it is a constant of the shape of this forecaster’s distribution.
Tenths sit on the same picture. Printed as their nominal values they add 0.000792 to the Brier score, of which 0.000708 is resolution lost and 0.000083 is the reliability of printing a band’s centre rather than its event rate. Relabelled by their event rates they lose only the 0.000708, which is the point drawn at eleven values, and sits almost on the equal-width curve — as it should, since tenths are equal-width bins with half-width bands at either end. The best eleven values lose 0.000645. Against tenths as usually printed that is a saving of 18.5%; against tenths relabelled it is 9.0%. The best ten values, one fewer than tenths, still lose slightly less than tenths printed as they are: 0.000782 against 0.000792.
The best eleven print 0.022, 0.094, 0.174, 0.261, 0.352, 0.446, 0.543, 0.643, 0.746, 0.849 and 0.954. They are roughly a tenths grid squeezed toward the middle, and most of the saving comes from the two ends. Tenths spend a whole printed value on a half-width band either side — 0 for everything below 0.05 and 1 for everything above 0.95 — and those bands are cheap in variance but wasted in vocabulary. The best eleven move those two values inward and use them on the dense lower region instead. The ranking improves as well, since the area under the ROC curve lost falls from 0.0033 for tenths to 0.0030. The best cuts for the Brier score are not the best for ranking, but here they happen to rank better than the grid they replace.
The cuts have to be found from a record
None of this is available to a forecaster in practice. The cuts above are computed from the distribution of the honest probability over the whole population, and a forecaster has instead a record of forecasts it has issued and outcomes it has seen. Finding the best cuts means running Lloyd’s iteration on that record: cut its own forecasts at its own quantiles, set each printed value to the mean forecast in its group, move each cut to the midpoint, and repeat. Labelling them means printing each group’s event rate as counted on the same record. Both steps are estimates, and both cost something the population vocabulary did not pay.
Cuts estimated away from the best ones lose more resolution than the best cuts. Labels estimated away from their groups’ true event rates add a reliability term, because a printed value that differs from the rate it claims is a miscalibration, however honest the procedure that produced it. The second cost is the one that matters, and it has a simple shape. A group holding a share of the forecasts has about of them on a record of , its counted rate has variance , and the reliability term weights that by . Summed over groups, the expected reliability is about . For the best eleven values that sum is 1.757, so a record of forecasts adds about to the Brier score through its labels alone. That is the same shape that made a perfect forecaster look miscalibrated on a finite record, and it is the same quantity: what binning a finite record costs, now charged to the forecaster rather than to its diagram.
The count agrees with the formula, and the formula is not kind. On two hundred records of 2,000 forecasts the fitted eleven add 0.001547 to the Brier score on average. Only 0.000710 of that is resolution, barely above the 0.000645 the exact cuts lose, and 0.000836 is labels, against 0.000879 from the formula. The fitted vocabulary is therefore about twice as costly as tenths, which need no fitting at all. At 500 forecasts it adds 0.003984, five times what tenths do. At 5,000 it adds 0.001014, still above tenths. At 20,000 it adds 0.000731, and at last beats them. The point where the label cost falls to the saving is where equals 0.000792 minus 0.000645, at about 12,000 forecasts.
The cuts themselves are cheap to find. On two thousand forecasts the fitted cuts lose only 10% more resolution than the exact ones, and on twenty thousand they lose 0.5% more. What is expensive is knowing what to print once the cuts are found. A grid of round numbers avoids that cost entirely: an honest forecaster rounding to tenths is reporting its own probability to the nearest step and is miscalibrated only by the small, fixed gap between a band’s centre and its rate. A fitted vocabulary that prints counted rates has replaced that small fixed bias with noise that shrinks only as .
That is the same lesson a chosen setting’s height taught in a different field. Choosing a thing from data and then reporting a number about it charges the data twice. Here it does so even though the cuts were chosen on forecasts and the labels counted on outcomes. The fitted vocabulary is the right target. Hitting it costs a record longer than many forecasters will ever accumulate on one product, and a vocabulary of five does not escape. Fitted five add 0.003615 on two thousand forecasts against the exact five’s 0.003193, a premium of 13%, smaller in proportion only because five groups hold more forecasts each.
Which users the Brier score speaks for
The forecaster that rounded showed a second way to read the same loss. A user with a cost ratio acts when the chance of the event exceeds . Acting on a coarse vocabulary instead of the honest probability costs that user a regret, which is zero when the printed value and the honest probability lead to the same decision and positive when they disagree. Twice the regret, averaged over spread evenly between 0 and 1, is the Brier score lost. So the Brier score speaks for a population of users whose cost ratios are spread evenly, and the best cuts above are the best cuts for them.
Few populations look like that. A vaccination campaign, an evacuation, a cheap precaution against an expensive loss — these are users who act at low chances, because acting is cheap and not acting is not. Take users whose cost ratios follow a Beta distribution with parameters 2 and 18, with a mean of 0.10 and most of its weight between 0.02 and 0.25. The vocabulary best for them is found the same way as Lloyd’s, with one change: each cut moves to the centroid, under the users’ own density, of the interval between its neighbouring printed values, rather than to the plain midpoint. The users’ density replaces the even spread that the Brier score assumes.
The vocabulary for those users cuts at 0.043, 0.093, 0.155 and 0.250, and prints 0.016, 0.067, 0.123, 0.201 and 0.597. Four of its five values sit below a quarter, and the fifth covers everything from 0.250 to 1 with a single number. Every user is well served where the users are, and badly served everywhere else. The regret curve shows it: low and tightly scalloped across the shaded region, then a single wide peak above 0.05 around a probability of 0.6, where a user acting at even odds is told 0.597 whether the honest chance is 0.3 or 0.95. The Brier-best vocabulary has smaller, even peaks across the whole axis, and its largest peaks sit in the very region the clustered users occupy, because its first group stretches from 0 to 0.145 and a user acting at 0.05 or 0.10 falls inside it.
The two populations disagree by an order of magnitude in each direction. Users clustered near one in ten expect a regret of 0.002286 from the Brier-best vocabulary and 0.000305 from their own, a factor of 7.5. Users spread evenly expect 0.001596 from the Brier-best vocabulary — half its 0.003193 of resolution lost, as the identity requires — and 0.013841 from the vocabulary cut for the clustered users, a factor of 8.7 the other way. As a Brier score, the vocabulary for clustered users loses 0.027683 of resolution, a little under nine times what the Brier-best one loses. A verification table would call it a poor forecast, and for the people it was built for it is the better of the two by a wide margin.
This is the same disagreement between two summaries that ran through comparisons of whole forecasters, and here it decides a format. A proper score is a choice of whom to average over. The Brier score averages evenly over cost ratios, and a format cut to minimise it is a format cut for nobody in particular. A format cut for a known population is better for that population and worse on the summary the field reports. Both are calibrated, so the reliability diagram cannot distinguish them.
What a vocabulary of printed probabilities should state
Whom it was cut for. The best cuts for the Brier score serve users spread evenly over cost ratios. A population that acts at low chances wants its cuts crowded where it acts, and the gap between the two is a factor of seven or eight in regret either way. A format that does not say whom it serves has chosen the even spread without saying so.
Whether the cuts and labels were fitted, and on how much. A fitted vocabulary pays about for its labels, which for eleven values in this world is 1.757 divided by the record’s length. Below about twelve thousand forecasts, tenths fixed in advance beat the best eleven fitted. A fitted format’s apparent saving is not a saving until the record is long enough to pay for it.
The loss at the stated number of values. For this forecaster the best cuts lose about , equal-width bins about and equal-mass groups about . Those constants are what a format designer trades against the number of values a reader will tolerate.
Proved, computed and counted
Proved. A vocabulary that prints its groups’ event rates has no reliability term, and its excess Brier score is the variance of the honest probability inside its groups; Lloyd’s two conditions are necessary for a vocabulary to minimise that; the excess Brier score is twice the regret of users with cost ratios spread evenly; and a label counted on forecasts adds to the reliability term on average.
Computed in closed form for this world. Every cut, printed value, resolution lost and area lost, from normal tails and bivariate orthants at the cuts. Lloyd’s iteration and its cost-weighted version run on those closed forms until no cut moves. Every regret is a difference of those tails, and averaged over a population it is integrated on a grid of cost ratios.
Counted. Eleven-value and five-value vocabularies fitted on two hundred records of 500, 2,000, 5,000 and 20,000 honest forecasts each, and then scored in closed form against the world. The counted label cost agrees with the formula to within ten per cent at every length, and within seven from two thousand forecasts up.
Still open: a vocabulary that can be refitted
A record grows. The vocabulary best for two thousand forecasts is tenths. The one best for twenty thousand is a fitted eleven, and between them lies a crossover near twelve thousand. That crossover is an average, so a forecaster switching at that length will switch too early on some records and too late on others. A format could start as tenths and move to fitted cuts once the record supports them, or shrink fitted labels toward the grid values they replace with a weight that grows with . That rule is the partial pooling of group estimates, applied to the groups of a vocabulary. Whether such a shrunk vocabulary beats both tenths and the fitted eleven at every record length, and by how much near twelve thousand forecasts, is a measurement these closed forms and records can make and have not yet made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The liar with two answers — both name brier score, closed form, discrimination, murphy's decomposition, probability forecast, proper scoring rule, reliability term, resolution term
- A score that rewards lying — both name brier score, discrimination, probability forecast, proper scoring rule
- A boundary for giving up — both name closed form, monte carlo
- A companion cut into many — both name closed form, monte carlo
- A companion with two coordinates — both name closed form, monte carlo
- A count that has to be estimated — both name closed form, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Brier scoreClosed formDiscriminationMonte CarloMurphy's decompositionProbability forecastProper scoring ruleReliability termResolution term