A forecast that is a probability

The liar with two answers

The forecaster an absolute-error score pays for says only 0 or 1, and on the ROC square it is a single point: its area is (TPR + TNR)/2 = 0.7684, against the honest forecaster's 0.8683, and it falls below the honest one on 200 of 200 counted records. No relabelling of its two answers returns what it threw away — the best recovers a Brier score short of the honest one by exactly the 0.022154 of resolution lost — and below a signal correlation of 0.7332 the same score prefers saying no every time to an honest forecast.

Worth reading first: An identity in three terms · What the model says next.

A score that rewards lying showed that an absolute-error score is minimised by reporting 0 or 1 — the nearer end of the interval to the forecaster’s honest probability — and that a liar doing exactly that beats a truthful forecaster on that score on every one of two hundred records. It then stated, structurally, what the lie costs: a step function discards the ordering of events inside each half of the interval, so the liar’s ability to rank events must be worse. And it said, correctly, that the liar’s own ranking was not among the numbers it had computed.

This essay computes it, and everything else the lie gives up. The liar here is the precise one the score instructs: a forecaster with the field’s full signal that says 1 wherever its honest probability exceeds a half and 0 everywhere else. Every quantity has a closed form in the bivariate normal that generates the field’s world, and the ones that can be counted on records are counted too.

One curve, and two forecasters that are each a single point. The ROC curve of an honest forecaster whose signal has correlation 1 with the latent state — every threshold on its probability, from the bivariate normal — and two forecasters that only ever say 0 or 1. The one an absolute-error score pays for says 1 wherever the honest probability exceeds a half: it has a true-positive rate of 0.6742 and a false-positive rate of 0.1375, and its "curve" is the two straight segments through that point, with area (TPR + TNR)/2 = 0.7684. The honest curve's area is 0.8683. Thresholding at the base rate of 0.3744 instead of at a half gives the largest area any two-valued forecaster can have here, 0.7818, and it is still 0.0866 short.
Fig. 1 The ROC curve of the honest forecaster with the full signal, traced through every threshold on its probability, beside two forecasters that only ever say 0 or 1: the one that thresholds at a half, which the absolute-error score pays for, and the one that thresholds at the base rate.

A forecaster with two answers is a point

An ROC curve is traced by moving a threshold along a forecaster’s reports and plotting, at each threshold, the share of events that happened that it flags against the share of non-events it flags. An honest probability has a continuum of reports and traces a curve. A forecaster that says only 0 or 1 has one threshold that means anything, and so one point: a true-positive rate and a false-positive rate. Its “curve” is the two straight segments from the origin through that point to the far corner, and the area under two segments is exactly

AUC=TPR+(1FPR)2,\text{AUC} = \frac{\text{TPR} + (1 - \text{FPR})}{2},

the average of the sensitivity and the specificity — the two numbers a screening test is described by, which is what a two-valued forecaster is.

For the liar, the true-positive rate is 0.6742 and the false-positive rate 0.1375, so its area is 0.7684. The honest forecaster’s area, from the same signal, is 0.8683. Both are closed forms. The joint probability that the liar says 1 and the event happens is an integral of the signal’s density against its honest probability past the threshold, which reduces to a bivariate normal orthant, and computing the same integral a second time by Simpson’s rule agrees with it to 1.8×10131.8 \times 10^{-13}.

The counted route shares nothing with either. On two hundred records of two thousand forecasts — the records the earlier essay counted the liar winning on — the area is computed by ranks, as the Mann–Whitney statistic, with tied reports sharing their rank and no threshold or normal distribution anywhere in the arithmetic. The liar’s counted area is 0.768093 ± 0.000682 against the closed 0.768387; the honest forecaster’s is 0.868066 ± 0.000551 against 0.868312. And the liar’s area is below the honest forecaster’s on 200 of 200 records. The same records that gave the liar a unanimous win on the absolute-error score give it a unanimous loss on ranking.

Resolution, which no relabelling returns

Ranking is one of two things a forecast carries. The other is resolution: how far the event rates in its own groups spread away from the base rate, which is the credit the three-term decomposition gives a forecaster for telling one case from another.

What each score charges the liar, and what cannot be bought back. Expected scores for a forecaster whose signal has correlation 1, lower being better. On the Brier score the honest probability is charged 0.141479 and the liar that says 0 or 1 by which side of a half it falls 0.207973. Replacing the liar's two answers by the event's frequency among the forecasts that carried them — 0.1844 and 0.7459 — is the best any relabelling can do, and it reaches 0.163633: short of the honest score by exactly the resolution the liar threw away, 0.022154. On the absolute-error score the order reverses, 0.207973 for the liar against 0.282958 for the honest probability, and for a forecast of 0 or 1 the absolute and Brier scores are the same number, the share of events it gets wrong.
Fig. 2 Expected Brier and absolute-error scores for the honest probability, the liar, the liar’s two answers relabelled by the event’s frequency among the forecasts that carried each, and the base rate issued every time. Lower is better.

The liar’s two reports define the partition by distinct forecast value, which is the one partition on which the decomposition is an identity rather than an approximation, and the identity holds exactly: on the liar’s two values the three terms sum to its Brier score with a residual of zero. Its resolution is 0.070604 against the honest forecaster’s 0.092758. The difference, 0.022154, is information.

The obvious rescue is to recalibrate: replace the liar’s 0 by the frequency of the event among the forecasts that said 0, which is 0.1844, and its 1 by the frequency among those that said 1, 0.7459. That relabelling removes the liar’s entire reliability term, because the new reports are the event rates of their own groups, and it is the best any relabelling of two answers can do. It reaches a Brier score of 0.163633. The honest forecaster’s Brier score is 0.141479. The gap between them is 0.022154 — exactly the resolution the liar lost, to the last digit, because recalibration repairs reliability and cannot touch resolution, and a record of two values has only two groups to resolve events into.

The two scores rank the pair in opposite orders, and the reason is visible in the same table. For a report of 0 or 1 the Brier score and the absolute-error score are the same number, the share of events the report gets wrong, 0.207973 for the liar. The honest forecaster is charged 0.141479 by Brier and 0.282958 by absolute error, because absolute error charges an honest 0.3 on an event that happened 0.7 and Brier charges it 0.49. Squaring rewards being close; the absolute value rewards being at an end.

Why the square cannot be gamed the same way

It is worth being exact about why the Brier score does not produce this liar, because the reason is also why the liar’s two answers cannot be repaired into an honest forecast. The Brier score charges a report r on an event of probability p an expected p(1r)2+(1p)r2p(1-r)^2 + (1-p)r^2. That is a parabola in r, and its minimum is where its slope is zero, at r = p — the truth, for every p. The absolute-error score’s expected charge, p(1r)+(1p)rp(1-r) + (1-p)r, is a straight line in r, and a straight line on an interval has its minimum at an end. The whole difference between a score that elicits a probability and one that elicits a certainty is whether the expected charge curves.

The curvature is also what makes the Brier score reward the information the liar discards. Among forecasts that are all calibrated, the Brier score is lower by exactly the resolution — the honest forecaster’s 0.092758 against the relabelled liar’s 0.070604 — because separating events into more groups with more different rates always reduces an expected squared charge. A straight-line charge has no such property. It rewards moving each report to the nearer end whatever that does to the separation of events, which is why it can pay a forecaster for converting a well-resolved forecast into a coarse one, and why the choice between two proper scores is a different kind of choice from the choice between a proper score and this one.

What a reliability diagram shows of a two-valued record

A reliability diagram groups forecasts by their value and plots each group’s mean forecast against the share of its events that happened. For a record with two values the diagram has two points and nothing else, and both of them are far from the diagonal: the forecasts of 0 carry an event rate of 0.1844, and the forecasts of 1 carry 0.7459. The liar’s reliability term — the charge for saying the wrong probability, averaged over the two groups — is 0.044340, where the honest forecaster’s is zero.

So the liar is easy to catch on a diagram and on the decomposition, and the earlier essay said as much. What this adds is how the two points are placed. The liar says 1 on 33.8% of its forecasts against a base rate of 37.4%, so its mean forecast is below the base rate even though it says “certainly” a third of the time: it is overconfident about the events it flags and silent about the rest. A diagram drawn with a binning chosen by the analyst would put both of these points into its first and last bins and report them as the two most miscalibrated groups on the page — correctly, and without any indication that the miscalibration was the score’s instruction rather than the forecaster’s mistake.

Three summaries, three best thresholds

The liar thresholds at a half because that is what the absolute-error score asks for. A forecaster that is going to issue only two answers anyway has a choice of threshold, and the choice reveals that the summaries of a two-valued forecast disagree about what it should be.

Three summaries of a two-valued forecast, three best thresholds. A forecaster with a signal of correlation 1 says 1 wherever its honest probability exceeds a threshold and 0 elsewhere. Each curve is how far one summary of that two-valued forecast falls short of its own best as the threshold moves, so each touches zero at the threshold it prefers. The share of events misjudged — which is the absolute-error score, and the Brier score too, on answers of 0 and 1 — is smallest at 0.5000, exactly a half. The area under the ROC curve is largest at 0.3744, which is the base rate, because there the signal's likelihood ratio is one. Resolution is largest at 0.4350, between the two, where it reads 0.071585. The liar the absolute score pays for is the best two-valued forecaster by one of these and not by the other two.
Fig. 3 For a two-valued forecaster that says 1 above a threshold on its honest probability, how far three summaries fall short of their own best as the threshold moves: the share of events misjudged, the area under the ROC curve, and resolution. Each touches zero at the threshold it prefers.

The share of events misjudged — which is the absolute-error score and the Brier score on answers of 0 and 1 — is smallest at a threshold of 0.5000, exactly a half, found by golden section on the closed form. The area under the ROC curve is largest at 0.3744, which is the base rate: at that threshold the signal’s likelihood ratio is one, so moving the threshold trades a true positive for a false positive at par, and the area stops improving. The best two-valued area there is 0.7818, still 0.0866 short of the honest forecaster’s. Resolution is largest at 0.4350, strictly between the two, where it reads 0.071585.

So the liar is the best two-valued forecaster by one summary and not by the other two, and the summary it wins on is the one that charges an honest probability for not being at an end. A threshold is a decision about which errors to prefer — the positive predictive value a screening cut buys is the same arithmetic — and an improper score makes that decision for the forecaster, silently, in favour of the errors it happens to count.

The ranking lost at every signal strength

The numbers so far are at the field’s full signal. The world lets the signal’s correlation with the latent state be dialled from nothing to full, and the loss can be read at every setting.

The ranking the liar gives up, at every signal strength. The area under the ROC curve against the correlation between a forecaster's signal and the latent state, for the honest probability, the two-valued forecaster that thresholds at the base rate, and the one that thresholds at a half — the liar an absolute-error score pays for. At a correlation of 1 the three read 0.8683, 0.7818 and 0.7684; at 0.5, 0.6773, 0.6272 and 0.5931. The liar gives up 0.0999 of area at the strongest signal and 0.0718 at 0.25, where its honest probability crosses a half on only 4.8% of forecasts and it says 0 nearly every time.
Fig. 4 The area under the ROC curve against the correlation between a forecaster’s signal and the latent state, for the honest probability, the two-valued forecaster that thresholds at the base rate, and the liar that thresholds at a half.

At full signal the honest, best two-valued and lying forecasters read 0.8683, 0.7818 and 0.7684; at a correlation of a half, 0.6773, 0.6272 and 0.5931. The liar gives up 0.0999 of area at full signal and 0.0718 at a correlation of a quarter, where something else happens: its honest probability crosses a half on only 4.8% of forecasts, so it says 0 nearly every time and ranks almost nothing. A liar with a weak signal is not a coarse forecaster; it is very nearly a constant.

That ordering — honest above best two-valued above liar — holds at every informative signal strength, and the gap between the last two is the price of the threshold the score chose, on top of the price of having two answers at all.

The score pays most for knowing nothing

If the lie costs information, a natural hope is that the score at least pays it less as the information it destroys grows. It does the opposite.

What the absolute-error score pays for the lie, by what the liar knows. The absolute-error score of the honest probability minus the score of the forecaster that says 0 or 1 by which side of a half that probability falls, at six signal strengths. The payment is 0.094025, 0.086486, 0.087571, 0.088263, 0.084406, 0.074986 at correlations 0, 0.25, 0.5, 0.7, 0.85, 1. It is largest, 0.094025 = 2p̄(1 − p̄) − p̄, for a forecaster whose signal carries nothing, where the lie is to say 0 every time and the forecast ranks nothing at all. The note beside each bar is the liar's area under the ROC curve against the honest forecaster's.
Fig. 5 What the absolute-error score pays for the lie — the honest probability’s score minus the liar’s — at six signal strengths, with the liar’s area against the honest forecaster’s beside each bar.

The payment for lying is 0.094025 for a forecaster whose signal carries nothing, 0.086486 at a correlation of a quarter, 0.087571 at a half, 0.088263 at 0.7, 0.084406 at 0.85 and 0.074986 at full signal. It is largest where the forecaster knows nothing, and there it has a closed form: an uninformed honest forecaster issues the base rate p̄ every time and is charged 2p̄(1 − p̄), while saying 0 every time is charged p̄, so the payment is 2p̄(1 − p̄) − p̄ exactly, 0.094025 at this base rate. The forecaster that earns the largest reward for gaming the score is the one whose lie destroys the least, because it had the least to destroy — and the reward is paid for converting a calibrated base rate into a constant that is wrong about every event that happens.

An honest forecast that loses to saying no

The last consequence follows from the one before, and it is the one that would matter most to a reader choosing a forecaster by this score.

An honest forecast that loses to saying no. The expected absolute-error score of an honest forecaster, E[2q(1 − q)], against how much of the latent state its signal carries, beside the score of a forecaster that says 0 every time, which is the base rate, 0.3744. The honest forecaster only beats saying no once its signal's correlation passes 0.7332, by which point its area under the ROC curve is already 0.7636 and its resolution 0.047012. Below that the score prefers a forecast that ranks nothing to one that ranks events genuinely well: at a correlation of a half the honest score is 0.425634, with an area of 0.6773.
Fig. 6 The expected absolute-error score of an honest forecaster against how much of the latent state its signal carries, beside the score of a forecaster that says 0 every time, with the signal strength at which the two tie marked.

A forecaster that says 0 every time is charged the base rate, 0.3744, and ranks nothing. An honest forecaster is charged E[2q(1q)]\mathrm{E}[2q(1-q)], which falls as its signal improves. The two tie at a signal correlation of 0.7332, found by bisection. Below that, the absolute-error score prefers the constant to the honest forecast — and at the tie the honest forecaster’s area under the ROC curve is already 0.7636 and its resolution 0.047012. A forecaster that ranks events well enough to be genuinely useful loses, on this score, to one that has never looked at the data.

The counted version is as unanimous as the earlier essay’s. On two hundred records of two thousand forecasts from an honest forecaster with a correlation of a half — an area of 0.6773 — the constant zero scores 0.374405 against the honest forecaster’s 0.425499, closed forms 0.374449 and 0.425634, and the constant wins on 200 of 200 records. An evaluation built on this score would not merely fail to reward the better forecaster; it would select against it with certainty, in the region where calibrated forecasts differ most in usefulness.

Which of these numbers are proved and which are counted

Proved. A two-valued forecaster’s ROC area is the mean of its two rates, which is the area of two line segments. The decomposition is an identity on the partition by distinct forecast value, so it is exact on the liar’s two answers, and recalibration removes reliability and leaves resolution, so the best relabelling falls short of the honest Brier score by exactly the resolution lost. The misjudged share is minimised at a half and the area at the threshold where the likelihood ratio is one, and the payment to an uninformed forecaster is 2p̄(1 − p̄) − p̄; each is a line of algebra, and the file checks each at machine precision rather than taking it on trust.

Computed in closed form for this world. Every rate, area, resolution and score at every threshold and signal strength, from bivariate normal orthants and one-dimensional quadrature against the world’s own law — including the tie at 0.7332 and the resolution-optimal threshold at 0.4350, which are properties of this world’s base rate and signal family rather than general constants.

Counted. The areas by ranks, the misjudged share and the win counts on two hundred records of two thousand forecasts, each agreeing with its closed form within a few of its own standard errors.

What an evaluation of probability forecasts should carry

A proper score, and not only one number. The earlier essay settled that the absolute-error score must go. What this essay adds is that even a proper score is one number, and the liar’s two failures — lost ranking and lost resolution — are visible only when discrimination and resolution are reported beside it.

The number of distinct values a forecaster issues. A liar’s record has two. It is the cheapest diagnostic available, it needs no outcomes at all, and a record of a thousand forecasts with two distinct values is a record whose ranking has been capped whatever its score says.

The comparison against a constant. A forecaster that cannot beat saying the base rate every time on a proper score has shown no skill; one that cannot beat saying 0 every time on an improper score may have shown nothing but the score’s preference, and a benchmark whose place in the comparison is fixed is what distinguishes the two.

Where this goes next: a forecaster that rounds

The liar has two answers because an improper score asked for them. Most real forecasters have a few answers for a different reason: forecasts are issued on a coarse scale — probabilities of precipitation in tenths, risk categories in five bands, credit grades — and rounding to a grid is a step function too, only with more steps. It is not a lie, because under a proper score the forecaster rounds its honest probability rather than pushing it to an end, but it discards the ordering inside each band exactly as the liar discards it inside each half.

The measurement that would settle how much coarse reporting costs is this essay’s table with the two answers replaced by k equal-width bands: the ROC area and resolution lost against k, the Brier score of the rounded forecast against the honest one, and the band count at which the loss falls under what a reliability diagram’s own binning already costs to read. It is distinct from this essay because the rounding is incentive-compatible — no score pays for it — so the question is purely what a reporting format throws away, and the answer is a number of bands.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Base rateBrier scoreClosed formDiscriminationImproper scoring ruleMurphy's decompositionProbability forecastProper scoring ruleRecalibrationReliability termResolution termROC curveSensitivity and specificity