A forecaster that rounds
Worth reading first: An identity in three terms · What the model says next.
The liar with two answers threw away everything inside each half of the probability scale because an improper score paid it to. Most forecasters with few answers have them for a different reason. Precipitation probabilities are issued in tenths, risk scores in a handful of categories, credit in grades, and a forecaster working to such a format rounds its honest probability to the nearest printed value. Nothing pays it to — under a proper score an honest forecaster would rather say 0.34 than 0.3 — so it is not lying. It is speaking a coarse language, and a coarse language discards the ordering inside each of its words exactly as the liar discarded the ordering inside each half.
This essay prices that. The forecaster is the field’s honest posterior at full signal, the one with an area under the ROC curve of 0.8683 and a Brier score of 0.141479, and the only thing changed is its vocabulary. Every loss has a closed form from the two tails the two-valued forecaster was built from — the probability that the honest forecast exceeds a cut, and the probability that it does and the event happens — differenced between consecutive cuts. The losses that can be counted on records are counted as well.
Rounding to the nearest whole number was the liar
Rounding to the nearest multiple of 1/k gives a vocabulary of k + 1 values, and at k = 1 the vocabulary is 0 and 1 with the cut at a half — which is the liar, exactly. Its area under the ROC curve is 0.7684, the same number the two-valued forecaster was given, now computed as the first of a family. Rounding to the nearest half, which allows 0, ½ and 1, lifts the area to 0.8205. Rounding to tenths reaches 0.8650, against the honest 0.8683.
On the ROC square a vocabulary of v values is v points joined by segments, and its area counts every pair of an event and a non-event that the forecast separates, with the pairs it ties counted as half. That is where the ranking goes: two forecasts of 0.31 and 0.39 are both issued as 0.3 and no longer say which case is riskier. Counted by ranks on two hundred records of two thousand forecasts — ties sharing their rank, no bivariate normal anywhere — the tenths score 0.865504 ± 0.000559, and the tenths rank worse than the honest forecasts on 200 of 200 records. Twentieths still rank worse on 196 of 200.
What each step of coarseness costs
The losses fall fast. The area given up is 0.0999 for whole numbers, 0.0478 for halves, 0.0111 for fifths, 0.0033 for tenths, 0.0009 for twentieths and 0.0002 for fiftieths — from fifths on, close to a quarter each time the step is halved, which is the square of the step. Resolution follows the same law: 0.022154, 0.011638, 0.002544, 0.000708, 0.000188 and 0.000032 across the same grids.
The law has a simple origin. A band of width 1/k holds honest probabilities spread nearly evenly across it once k is large, and a uniform spread over a width 1/k has variance , which is 0.000833 at tenths. The exact figure is lower, 0.000708, because the two end bands of a rounding grid are only half as wide — everything below 0.05 is issued as 0 — and by fiftieths the two are 0.000033 against 0.000032.
Resolution is most of the damage to the Brier score, and the rest is small. The tenths’ reports are not quite the event rates of their own bands — the forecasts issued as 0.1 carry an event rate of 0.0963, and those issued as 1 carry 0.9768 — and that mismatch is a reliability term of 0.000083. Added to the resolution lost it gives the Brier score added, 0.000792, of which 89.5% is resolution. Relabelling each tenth by its band’s event rate would remove the 0.000083 and nothing else, for the same reason relabelling the liar’s two answers could not return its resolution.
The same term a reliability diagram already cannot see
The resolution lost has a second reading, and it is computed a second way. Resolution on a partition is the variance of the group event rates, so what a partition loses against the honest forecaster is the variance of the honest probability inside the groups, averaged over them. Integrating that variance on the signal band by band, by Simpson’s rule and without reference to the orthants, gives 0.000708 at tenths and matches the closed resolution lost at every grid to the ninth decimal.
That inside-the-band variance is also the within-bin term of a reliability diagram drawn on those bands. A diagram groups an unrounded forecaster’s forecasts into bins and reads the three terms of the decomposition off the bin averages, and the part of the Brier score those averages cannot carry is exactly the variance of the forecasts inside each bin. So a forecaster who rounds to the evaluator’s own bins loses precisely what the evaluator’s diagram was already discarding. On ten equal-width bins that is 0.000853. An evaluation that only ever looks at a ten-bin diagram cannot tell a forecaster issuing continuous probabilities from one issuing each bin’s event rate — and a finer diagram, or a Brier score computed forecast by forecast, can.
Users between two printed values
A Brier score averages over users who never appear in it. A user with cost–loss ratio t loses t for acting on a non-event and 1 − t for failing to act on an event, and acts whenever the forecast exceeds t. The squared error of a probability forecast is twice the integral of that loss over every t from 0 to 1, so the Brier score added by rounding is twice the average regret across users. Computed threshold by threshold from the same two tails, the average regret for tenths is 0.000396, and twice it is the 0.000792 found from the decomposition.
The average hides a sawtooth. A user whose threshold sits exactly at a cut point between two printed values — 0.05, 0.15, 0.25 and so on — acts on the tenths exactly when acting on the honest probability and loses nothing. A user whose threshold sits at a printed value loses most, and loses it on both sides of the value for opposite reasons.
Just past a printed value, the tenths make the user wait: a user acting on anything over 30% cannot act before the forecast reaches 0.4, which happens only when the honest probability passes 0.35, so every case between 0.30 and 0.35 goes unacted on. Just below a printed value, the tenths make the user act too soon: a user acting on anything over 29.9% acts on a printed 0.3, which is issued once the honest probability passes 0.25, so every case between 0.25 and 0.299 is acted on at a loss. The regret just below 0.1 is 0.002289 and just past it 0.001700; at 0.2, 0.001544 and 0.001334; at 0.5, 0.000984 and 0.000929; at 0.8, 0.000804 and 0.000788. The side below is the larger at every printed value from 0.1 to 0.8, and the gap closes up the scale, because the honest forecasts are densest at low probabilities and a band there holds more cases to act on wrongly. At 0.9 the two cross, 0.000785 below against 0.000802 past, where the band above is the half-width one issued as 100%.
So a forecast in tenths is least useful to exactly the users a table of tenths invites: the ones whose rule is written in the table’s own numbers, “act at 10%” or “act at 30%”. A user who could state a threshold of 0.25 instead would lose nothing at all.
The largest regret belongs to a user who acts at almost any chance of the event. That user can act on any printed value but 0, so the tenths cost it exactly the events issued as 0% that then happened — 0.003207 of all forecasts, eight times the average regret. At the other end a user who acts only on near-certainty is charged the forecasts issued as 100% on events that did not happen, 0.000858. Rounding costs little on average and most at the edges of the scale, which is where the users with the most extreme cost ratios are.
Whether a record can see the loss
A loss of 0.000792 is a population number. A record of forecasts is finite, and whether the rounding is visible on one record is a separate question with a counted answer. On two hundred records of two thousand forecasts, the honest record scores better than the same record rounded on 200 of 200 for whole numbers, 200 for halves, 199 for fifths, 189 for tenths and 164 for twentieths. The counted Brier score added for tenths is 0.000789 against the closed 0.000792.
The same records give the noise the loss has to be seen against. The difference between the rounded and the honest squared error on a single forecast has a standard deviation of 0.022645 for tenths, so a record needs about 3,296 forecasts before the rounding’s loss sits two standard errors from zero. For fifths it needs about 946, and for twentieths about 10,382. A year of daily forecasts is 365 of them, which is why a forecaster issuing tenths and one issuing continuous probabilities cannot be told apart by any single station’s record — and calibration alone never could tell them apart, since both are calibrated to within the reliability term of 0.000083.
Equal steps, equal bins or equal risk
A grid of equal steps is one vocabulary among several. Risk categories are usually cut so that each holds a similar share of the population and are then labelled with each group’s observed rate, which makes them calibrated by construction. With five groups of equal mass the forecaster loses 0.0144 of area against 0.0178 for five equal-width bins, and 0.004046 of resolution against 0.003454. With ten groups the same trade holds: 0.0036 against 0.0047 of area, 0.001006 against 0.000853 of resolution.
Equal mass ranks better and resolves worse, and the two summaries are measuring different things about the same cuts. Area lost counts tied pairs, and ties are spread most evenly when each group holds the same share of forecasts: ten equal-width bins put 25.7% of all forecasts, every one of them below a probability of 0.1, into a single tie, where the equal-mass groups split the same forecasts across three groups, the first of which ends at 0.021. Resolution lost is variance inside the groups, and the upper equal-mass groups have to stretch to collect their share of the rarer high forecasts — the top one runs from 0.850 to 1 — where every equal-width bin is a tenth wide, and a wide group has a large variance whatever it holds. A format chosen to rank patients and a format chosen to state their risks want different cuts, which is the distinction between two proper summaries arriving as a design decision about a table.
A zero a logarithmic score cannot forgive
A grid that includes 0 issues it whenever the honest probability is below half a step, and the honest probability down there is small but not zero. In tenths, 17.0% of all forecasts are issued as 0%, and those forecasts carry an event rate of 1.89%: 3.207 in every thousand forecasts are a 0% on an event that happened. The counted records agree at 3.18 per thousand. The rate falls with the step and never reaches zero — 121.980 per thousand for whole numbers, 41.101 for halves, 9.658 for fifths and 1.060 for twentieths.
On the Brier score each of those costs one, the most any single forecast can cost. On a logarithmic score each costs minus the logarithm of zero, which is unbounded, so the expected logarithmic score of every one of these grids is infinite, however fine, because every one of them prints 0% on events that happen at a positive rate. The two proper scores that disagreed about a loud and a hedged forecaster disagree here without limit. The repair is a vocabulary without its ends: issuing the bottom band’s event rate, 1.89%, instead of 0% costs nothing on the Brier score beyond what the band already cost and removes the infinite charge entirely.
Relabelling every printed value by its band’s event rate makes the cost finite everywhere, and then it can be read. The honest forecaster’s expected logarithmic score is 0.432276 nats. Tenths relabelled by their event rates score 0.435881, a charge of 0.003605 nats for the rounding, against the 0.000708 the same relabelled tenths add to the Brier score. Fifths relabelled add 0.010398 nats and twentieths 0.001214. The logarithmic score charges rounding about five times what the Brier score does in these units, and it charges most in the end bands, where a probability of 0.02 and one of 0.001 are nearly the same number to a square and very different numbers to a logarithm — which is the same reason it cannot forgive a printed zero at all.
A coarse vocabulary costs a weak forecaster more
The numbers so far are at full signal. At weaker signals the resolution lost to tenths hardly moves — 0.000719 at a signal correlation of a quarter, 0.000798 at a half and 0.000708 at full signal — because it is set by the width of the bands rather than by what the forecaster knows. What moves is how much of the forecaster’s resolution that is: 13.54% at a correlation of a quarter, 3.72% at a half and 0.76% at full signal.
Ranking tells the same story more sharply. At a correlation of a quarter the honest area is 0.5879, so the forecaster’s ranking skill above chance is less than nine points of area, and rounding to fifths gives up 0.0401 of it — 45.6% of the skill. Tenths give up 0.0119. A forecaster that knows a great deal can afford a coarse vocabulary; one that knows a little spends a large share of what it knows on the format.
What a probability format should state
The number of steps, and whether the ends are printed. Tenths lose 0.000792 of Brier score, which no record shorter than a few thousand forecasts can see, and they print 0% on 3.207 events in every thousand forecasts, which a logarithmic score cannot score at all.
The event rate inside each printed value. It is the band’s reliability, and relabelling by it removes the only part of the rounding’s cost that can be removed.
How the cuts were chosen. Equal steps, equal-width bins and equal-mass groups spend the same number of values differently, and the choice decides whether the format protects ranking or resolution.
Proved, computed and counted
Proved. Rounding’s resolution lost is the variance of the honest probability inside the rounding groups, which is the within-bin term of a diagram drawn on those groups; the excess Brier score is that plus the grid’s reliability term; and the Brier score is twice the integral over thresholds of a cost–loss user’s loss, so twice the average regret is the excess.
Computed in closed form for this world. Every area, resolution, reliability term, regret and zero rate, from normal tails and bivariate normal orthants at the cuts, with the resolution lost computed a second time by quadrature on the signal and the excess a second time by integrating the regret. The 3,296 forecasts are a ratio of a closed loss and a counted standard deviation.
Counted. Areas by ranks, Brier scores and zero reports on two hundred records of two thousand forecasts, each agreeing with its closed form within a few of its own standard errors.
Still open: the best few values to print
Equal steps, equal-width bins and equal-mass groups are all chosen before the forecaster’s distribution is looked at. The vocabulary of k values that loses least on the Brier score is not any of them: it is the quantiser that places each printed value at the event rate of its group and each cut halfway between neighbouring printed values, found by iterating those two conditions, and it depends on where the honest forecasts actually fall. What that best vocabulary of five or ten values looks like in this world, how much it saves against tenths — at most 0.000792 — and whether the cuts it chooses for the Brier score are anywhere near those a user’s cost ratio or a benchmark fixed in advance would choose, is a measurement these closed forms can make and have not made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A score that rewards lying — both name brier score, discrimination, probability forecast, proper scoring rule, roc curve
- The miscalibration a perfect forecaster shows — both name probability forecast, reliability diagram, reliability term
- A boundary for giving up — both name closed form, monte carlo
- A count that has to be estimated — both name closed form, monte carlo
- A covariate with no levels — both name closed form, monte carlo
- A coverage table with its own error — both name closed form, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Brier scoreClosed formDiscriminationMonte CarloMurphy's decompositionProbability forecastProper scoring ruleReliability diagramReliability termResolution termROC curve