A forecast that is a probability

A forecaster that rounds

An honest probability issued in tenths loses 0.0033 of ROC area and 0.000708 of resolution — the variance its bands average away, and 89.5% of the 0.000792 it adds to the Brier score. Two hundred records of two thousand forecasts show that loss on 189; it takes about 3,300 forecasts to put it two standard errors from zero. And 3.207 in every thousand forecasts in tenths are a 0% on an event that happened, which a logarithmic score charges without limit.

Worth reading first: An identity in three terms · What the model says next.

The liar with two answers threw away everything inside each half of the probability scale because an improper score paid it to. Most forecasters with few answers have them for a different reason. Precipitation probabilities are issued in tenths, risk scores in a handful of categories, credit in grades, and a forecaster working to such a format rounds its honest probability to the nearest printed value. Nothing pays it to — under a proper score an honest forecaster would rather say 0.34 than 0.3 — so it is not lying. It is speaking a coarse language, and a coarse language discards the ordering inside each of its words exactly as the liar discarded the ordering inside each half.

This essay prices that. The forecaster is the field’s honest posterior at full signal, the one with an area under the ROC curve of 0.8683 and a Brier score of 0.141479, and the only thing changed is its vocabulary. Every loss has a closed form from the two tails the two-valued forecaster was built from — the probability that the honest forecast exceeds a cut, and the probability that it does and the event happens — differenced between consecutive cuts. The losses that can be counted on records are counted as well.

The honest curve, and the same forecaster in three coarse vocabularies. The ROC curve of the honest probability at full signal, area 0.8683, beside the same forecaster rounded to the nearest whole number, area 0.7684 — the two-valued liar — to the nearest half, area 0.8205, and to the nearest tenth, area 0.8650. A vocabulary of v values is v points on the square joined by straight segments, and tied reports count half.
Fig. 1 The honest forecaster’s ROC curve beside the same forecaster rounded to the nearest whole number, the nearest half and the nearest tenth. A vocabulary of v values is v points on the square joined by straight segments.

Rounding to the nearest whole number was the liar

Rounding to the nearest multiple of 1/k gives a vocabulary of k + 1 values, and at k = 1 the vocabulary is 0 and 1 with the cut at a half — which is the liar, exactly. Its area under the ROC curve is 0.7684, the same number the two-valued forecaster was given, now computed as the first of a family. Rounding to the nearest half, which allows 0, ½ and 1, lifts the area to 0.8205. Rounding to tenths reaches 0.8650, against the honest 0.8683.

On the ROC square a vocabulary of v values is v points joined by segments, and its area counts every pair of an event and a non-event that the forecast separates, with the pairs it ties counted as half. That is where the ranking goes: two forecasts of 0.31 and 0.39 are both issued as 0.3 and no longer say which case is riskier. Counted by ranks on two hundred records of two thousand forecasts — ties sharing their rank, no bivariate normal anywhere — the tenths score 0.865504 ± 0.000559, and the tenths rank worse than the honest forecasts on 200 of 200 records. Twentieths still rank worse on 196 of 200.

What each step of coarseness costs

What rounding an honest probability gives up, by the number of steps in the grid. Rounded to the nearest multiple of 1/k, the honest forecaster at full signal loses 0.0999 of area and 0.022154 of resolution at k = 1; 0.0478 of area and 0.011638 of resolution at k = 2; 0.0259 of area and 0.006169 of resolution at k = 3; 0.0162 of area and 0.003777 of resolution at k = 4; 0.0111 of area and 0.002544 of resolution at k = 5; 0.0081 of area and 0.001829 of resolution at k = 6; 0.0049 of area and 0.001076 of resolution at k = 8; 0.0033 of area and 0.000708 of resolution at k = 10; 0.0024 of area and 0.000502 of resolution at k = 12; 0.0016 of area and 0.000328 of resolution at k = 15; 0.0009 of area and 0.000188 of resolution at k = 20; 0.0004 of area and 0.000086 of resolution at k = 30; 0.0002 of area and 0.000032 of resolution at k = 50. The resolution lost equals the variance of the honest probability inside the bands, computed a second way by quadrature, and approaches 1/(12k²). The Brier score added is the resolution lost plus the grid's reliability term. Open circles are the Brier score added counted on 200 records of 2,000 forecasts at k = 1, 2, 5, 10, 20.
Fig. 2 Against the number of steps in the grid, on log scales: the area under the ROC curve lost, the resolution lost and the Brier score added, with 1/(12k2)1/(12k^2) as a guide and the counted Brier score added as open circles.

The losses fall fast. The area given up is 0.0999 for whole numbers, 0.0478 for halves, 0.0111 for fifths, 0.0033 for tenths, 0.0009 for twentieths and 0.0002 for fiftieths — from fifths on, close to a quarter each time the step is halved, which is the square of the step. Resolution follows the same law: 0.022154, 0.011638, 0.002544, 0.000708, 0.000188 and 0.000032 across the same grids.

The law has a simple origin. A band of width 1/k holds honest probabilities spread nearly evenly across it once k is large, and a uniform spread over a width 1/k has variance 1/(12k2)1/(12k^2), which is 0.000833 at tenths. The exact figure is lower, 0.000708, because the two end bands of a rounding grid are only half as wide — everything below 0.05 is issued as 0 — and by fiftieths the two are 0.000033 against 0.000032.

Resolution is most of the damage to the Brier score, and the rest is small. The tenths’ reports are not quite the event rates of their own bands — the forecasts issued as 0.1 carry an event rate of 0.0963, and those issued as 1 carry 0.9768 — and that mismatch is a reliability term of 0.000083. Added to the resolution lost it gives the Brier score added, 0.000792, of which 89.5% is resolution. Relabelling each tenth by its band’s event rate would remove the 0.000083 and nothing else, for the same reason relabelling the liar’s two answers could not return its resolution.

The same term a reliability diagram already cannot see

The resolution lost has a second reading, and it is computed a second way. Resolution on a partition is the variance of the group event rates, so what a partition loses against the honest forecaster is the variance of the honest probability inside the groups, averaged over them. Integrating that variance on the signal band by band, by Simpson’s rule and without reference to the orthants, gives 0.000708 at tenths and matches the closed resolution lost at every grid to the ninth decimal.

That inside-the-band variance is also the within-bin term of a reliability diagram drawn on those bands. A diagram groups an unrounded forecaster’s forecasts into bins and reads the three terms of the decomposition off the bin averages, and the part of the Brier score those averages cannot carry is exactly the variance of the forecasts inside each bin. So a forecaster who rounds to the evaluator’s own bins loses precisely what the evaluator’s diagram was already discarding. On ten equal-width bins that is 0.000853. An evaluation that only ever looks at a ten-bin diagram cannot tell a forecaster issuing continuous probabilities from one issuing each bin’s event rate — and a finer diagram, or a Brier score computed forecast by forecast, can.

Users between two printed values

The regret of acting on a forecast in tenths, at every threshold a user might have. A user who acts when the event's probability exceeds a threshold t loses nothing by acting on the honest probability and something by acting on the tenths. The regret is zero for thresholds at the cut points 0.05, 0.15, …, 0.95, and largest on either side of each printed value — just below it the user acts too early, just past it too late: 0.002289 just below 0.1 and 0.001700 just past it, 0.001544 just below 0.2 and 0.001334 just past it, 0.000984 just below 0.5 and 0.000929 just past it, 0.000804 just below 0.8 and 0.000788 just past it; 0.003207 for a user acting at almost any chance, which is the share of forecasts that are a 0% on an event that happened. Averaged over thresholds it is 0.000396, and twice that is the tenths' excess Brier score, 0.000792.
Fig. 3 The expected regret of acting on a forecast in tenths rather than on the honest probability, for a user who acts when the probability exceeds a threshold, at every threshold, with its average across thresholds.

A Brier score averages over users who never appear in it. A user with cost–loss ratio t loses t for acting on a non-event and 1 − t for failing to act on an event, and acts whenever the forecast exceeds t. The squared error of a probability forecast is twice the integral of that loss over every t from 0 to 1, so the Brier score added by rounding is twice the average regret across users. Computed threshold by threshold from the same two tails, the average regret for tenths is 0.000396, and twice it is the 0.000792 found from the decomposition.

The average hides a sawtooth. A user whose threshold sits exactly at a cut point between two printed values — 0.05, 0.15, 0.25 and so on — acts on the tenths exactly when acting on the honest probability and loses nothing. A user whose threshold sits at a printed value loses most, and loses it on both sides of the value for opposite reasons.

Just past a printed value, the tenths make the user wait: a user acting on anything over 30% cannot act before the forecast reaches 0.4, which happens only when the honest probability passes 0.35, so every case between 0.30 and 0.35 goes unacted on. Just below a printed value, the tenths make the user act too soon: a user acting on anything over 29.9% acts on a printed 0.3, which is issued once the honest probability passes 0.25, so every case between 0.25 and 0.299 is acted on at a loss. The regret just below 0.1 is 0.002289 and just past it 0.001700; at 0.2, 0.001544 and 0.001334; at 0.5, 0.000984 and 0.000929; at 0.8, 0.000804 and 0.000788. The side below is the larger at every printed value from 0.1 to 0.8, and the gap closes up the scale, because the honest forecasts are densest at low probabilities and a band there holds more cases to act on wrongly. At 0.9 the two cross, 0.000785 below against 0.000802 past, where the band above is the half-width one issued as 100%.

So a forecast in tenths is least useful to exactly the users a table of tenths invites: the ones whose rule is written in the table’s own numbers, “act at 10%” or “act at 30%”. A user who could state a threshold of 0.25 instead would lose nothing at all.

The largest regret belongs to a user who acts at almost any chance of the event. That user can act on any printed value but 0, so the tenths cost it exactly the events issued as 0% that then happened — 0.003207 of all forecasts, eight times the average regret. At the other end a user who acts only on near-certainty is charged the forecasts issued as 100% on events that did not happen, 0.000858. Rounding costs little on average and most at the edges of the scale, which is where the users with the most extreme cost ratios are.

Whether a record can see the loss

On how many records of two thousand forecasts the rounding shows. Two hundred records of 2,000 forecasts from the honest forecaster at full signal, each scored by Brier as issued and rounded. The honest record scores better on 200 of 200 in 0 or 1, 200 of 200 in halves, 199 of 200 in fifths, 189 of 200 in tenths, 164 of 200 in twentieths. The mean Brier score added is 0.065504 counted against 0.066494 closed in 0 or 1; 0.018003 counted against 0.018024 closed in halves; 0.003039 counted against 0.003076 closed in fifths; 0.000789 counted against 0.000792 closed in tenths; 0.000206 counted against 0.000202 closed in twentieths.
Fig. 4 Two hundred records of two thousand forecasts, each scored by Brier as issued and rounded: on how many the honest record scores better, with the Brier score added counted beside its closed value.

A loss of 0.000792 is a population number. A record of forecasts is finite, and whether the rounding is visible on one record is a separate question with a counted answer. On two hundred records of two thousand forecasts, the honest record scores better than the same record rounded on 200 of 200 for whole numbers, 200 for halves, 199 for fifths, 189 for tenths and 164 for twentieths. The counted Brier score added for tenths is 0.000789 against the closed 0.000792.

The same records give the noise the loss has to be seen against. The difference between the rounded and the honest squared error on a single forecast has a standard deviation of 0.022645 for tenths, so a record needs about 3,296 forecasts before the rounding’s loss sits two standard errors from zero. For fifths it needs about 946, and for twentieths about 10,382. A year of daily forecasts is 365 of them, which is why a forecaster issuing tenths and one issuing continuous probabilities cannot be told apart by any single station’s record — and calibration alone never could tell them apart, since both are calibrated to within the reliability term of 0.000083.

Equal steps, equal bins or equal risk

Equal steps, equal-width bins and equal-mass groups, by what each gives up. The honest forecaster at full signal reported three ways: the fifths grid, six values loses 0.0111 of area and 0.002544 of resolution; five equal-width bins loses 0.0178 of area and 0.003454 of resolution; five equal-mass groups loses 0.0144 of area and 0.004046 of resolution; the tenths grid, eleven values loses 0.0033 of area and 0.000708 of resolution; ten equal-width bins loses 0.0047 of area and 0.000853 of resolution; ten equal-mass groups loses 0.0036 of area and 0.001006 of resolution. Groups of equal probability mass keep more of the ranking and less of the resolution than bins of equal width with the same number of groups.
Fig. 5 The honest forecaster reported on the rounding grid, in equal-width bins reporting their midpoints, and in groups of equal probability mass reporting their event rates, with five and ten groups: the area under the ROC curve lost, and the resolution lost beside it.

A grid of equal steps is one vocabulary among several. Risk categories are usually cut so that each holds a similar share of the population and are then labelled with each group’s observed rate, which makes them calibrated by construction. With five groups of equal mass the forecaster loses 0.0144 of area against 0.0178 for five equal-width bins, and 0.004046 of resolution against 0.003454. With ten groups the same trade holds: 0.0036 against 0.0047 of area, 0.001006 against 0.000853 of resolution.

Equal mass ranks better and resolves worse, and the two summaries are measuring different things about the same cuts. Area lost counts tied pairs, and ties are spread most evenly when each group holds the same share of forecasts: ten equal-width bins put 25.7% of all forecasts, every one of them below a probability of 0.1, into a single tie, where the equal-mass groups split the same forecasts across three groups, the first of which ends at 0.021. Resolution lost is variance inside the groups, and the upper equal-mass groups have to stretch to collect their share of the rarer high forecasts — the top one runs from 0.850 to 1 — where every equal-width bin is a tenth wide, and a wide group has a large variance whatever it holds. A format chosen to rank patients and a format chosen to state their risks want different cuts, which is the distinction between two proper summaries arriving as a design decision about a table.

A zero a logarithmic score cannot forgive

Forecasts of 0% on events that happened, per thousand forecasts, by grid. Rounded to the nearest multiple of 1/k, the honest forecaster reports exactly 0% whenever its probability is below 1/(2k), and some of those events happen: 121.980 per thousand forecasts in 0 or 1, 121.31 counted; 41.101 per thousand forecasts in halves, 40.87 counted; 9.658 per thousand forecasts in fifths, 9.53 counted; 3.207 per thousand forecasts in tenths, 3.18 counted; 1.060 per thousand forecasts in twentieths, 1.01 counted. A logarithmic score charges each of them without limit, so the expected logarithmic score of every one of these grids is infinite. Reports of 100% on events that did not happen: 85.992 in 0 or 1, 19.914 in halves, 3.245 in fifths, 0.858 in tenths, 0.232 in twentieths per thousand.
Fig. 6 Forecasts issued as exactly 0% on events that happened, per thousand forecasts, for each grid, with the counted rate and the forecasts of 100% on events that did not happen.

A grid that includes 0 issues it whenever the honest probability is below half a step, and the honest probability down there is small but not zero. In tenths, 17.0% of all forecasts are issued as 0%, and those forecasts carry an event rate of 1.89%: 3.207 in every thousand forecasts are a 0% on an event that happened. The counted records agree at 3.18 per thousand. The rate falls with the step and never reaches zero — 121.980 per thousand for whole numbers, 41.101 for halves, 9.658 for fifths and 1.060 for twentieths.

On the Brier score each of those costs one, the most any single forecast can cost. On a logarithmic score each costs minus the logarithm of zero, which is unbounded, so the expected logarithmic score of every one of these grids is infinite, however fine, because every one of them prints 0% on events that happen at a positive rate. The two proper scores that disagreed about a loud and a hedged forecaster disagree here without limit. The repair is a vocabulary without its ends: issuing the bottom band’s event rate, 1.89%, instead of 0% costs nothing on the Brier score beyond what the band already cost and removes the infinite charge entirely.

Relabelling every printed value by its band’s event rate makes the cost finite everywhere, and then it can be read. The honest forecaster’s expected logarithmic score is 0.432276 nats. Tenths relabelled by their event rates score 0.435881, a charge of 0.003605 nats for the rounding, against the 0.000708 the same relabelled tenths add to the Brier score. Fifths relabelled add 0.010398 nats and twentieths 0.001214. The logarithmic score charges rounding about five times what the Brier score does in these units, and it charges most in the end bands, where a probability of 0.02 and one of 0.001 are nearly the same number to a square and very different numbers to a logarithm — which is the same reason it cannot forgive a printed zero at all.

A coarse vocabulary costs a weak forecaster more

The numbers so far are at full signal. At weaker signals the resolution lost to tenths hardly moves — 0.000719 at a signal correlation of a quarter, 0.000798 at a half and 0.000708 at full signal — because it is set by the width of the bands rather than by what the forecaster knows. What moves is how much of the forecaster’s resolution that is: 13.54% at a correlation of a quarter, 3.72% at a half and 0.76% at full signal.

Ranking tells the same story more sharply. At a correlation of a quarter the honest area is 0.5879, so the forecaster’s ranking skill above chance is less than nine points of area, and rounding to fifths gives up 0.0401 of it — 45.6% of the skill. Tenths give up 0.0119. A forecaster that knows a great deal can afford a coarse vocabulary; one that knows a little spends a large share of what it knows on the format.

What a probability format should state

The number of steps, and whether the ends are printed. Tenths lose 0.000792 of Brier score, which no record shorter than a few thousand forecasts can see, and they print 0% on 3.207 events in every thousand forecasts, which a logarithmic score cannot score at all.

The event rate inside each printed value. It is the band’s reliability, and relabelling by it removes the only part of the rounding’s cost that can be removed.

How the cuts were chosen. Equal steps, equal-width bins and equal-mass groups spend the same number of values differently, and the choice decides whether the format protects ranking or resolution.

Proved, computed and counted

Proved. Rounding’s resolution lost is the variance of the honest probability inside the rounding groups, which is the within-bin term of a diagram drawn on those groups; the excess Brier score is that plus the grid’s reliability term; and the Brier score is twice the integral over thresholds of a cost–loss user’s loss, so twice the average regret is the excess.

Computed in closed form for this world. Every area, resolution, reliability term, regret and zero rate, from normal tails and bivariate normal orthants at the cuts, with the resolution lost computed a second time by quadrature on the signal and the excess a second time by integrating the regret. The 3,296 forecasts are a ratio of a closed loss and a counted standard deviation.

Counted. Areas by ranks, Brier scores and zero reports on two hundred records of two thousand forecasts, each agreeing with its closed form within a few of its own standard errors.

Still open: the best few values to print

Equal steps, equal-width bins and equal-mass groups are all chosen before the forecaster’s distribution is looked at. The vocabulary of k values that loses least on the Brier score is not any of them: it is the quantiser that places each printed value at the event rate of its group and each cut halfway between neighbouring printed values, found by iterating those two conditions, and it depends on where the honest forecasts actually fall. What that best vocabulary of five or ten values looks like in this world, how much it saves against tenths — at most 0.000792 — and whether the cuts it chooses for the Brier score are anywhere near those a user’s cost ratio or a benchmark fixed in advance would choose, is a measurement these closed forms can make and have not made.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Brier scoreClosed formDiscriminationMonte CarloMurphy's decompositionProbability forecastProper scoring ruleReliability diagramReliability termResolution termROC curve