The liar with two answers
Worth reading first: An identity in three terms · What the model says next.
A score that rewards lying showed that an absolute-error score is minimised by reporting 0 or 1 — the nearer end of the interval to the forecaster’s honest probability — and that a liar doing exactly that beats a truthful forecaster on that score on every one of two hundred records. It then stated, structurally, what the lie costs: a step function discards the ordering of events inside each half of the interval, so the liar’s ability to rank events must be worse. And it said, correctly, that the liar’s own ranking was not among the numbers it had computed.
This essay computes it, and everything else the lie gives up. The liar here is the precise one the score instructs: a forecaster with the field’s full signal that says 1 wherever its honest probability exceeds a half and 0 everywhere else. Every quantity has a closed form in the bivariate normal that generates the field’s world, and the ones that can be counted on records are counted too.
A forecaster with two answers is a point
An ROC curve is traced by moving a threshold along a forecaster’s reports and plotting, at each threshold, the share of events that happened that it flags against the share of non-events it flags. An honest probability has a continuum of reports and traces a curve. A forecaster that says only 0 or 1 has one threshold that means anything, and so one point: a true-positive rate and a false-positive rate. Its “curve” is the two straight segments from the origin through that point to the far corner, and the area under two segments is exactly
the average of the sensitivity and the specificity — the two numbers a screening test is described by, which is what a two-valued forecaster is.
For the liar, the true-positive rate is 0.6742 and the false-positive rate 0.1375, so its area is 0.7684. The honest forecaster’s area, from the same signal, is 0.8683. Both are closed forms. The joint probability that the liar says 1 and the event happens is an integral of the signal’s density against its honest probability past the threshold, which reduces to a bivariate normal orthant, and computing the same integral a second time by Simpson’s rule agrees with it to .
The counted route shares nothing with either. On two hundred records of two thousand forecasts — the records the earlier essay counted the liar winning on — the area is computed by ranks, as the Mann–Whitney statistic, with tied reports sharing their rank and no threshold or normal distribution anywhere in the arithmetic. The liar’s counted area is 0.768093 ± 0.000682 against the closed 0.768387; the honest forecaster’s is 0.868066 ± 0.000551 against 0.868312. And the liar’s area is below the honest forecaster’s on 200 of 200 records. The same records that gave the liar a unanimous win on the absolute-error score give it a unanimous loss on ranking.
Resolution, which no relabelling returns
Ranking is one of two things a forecast carries. The other is resolution: how far the event rates in its own groups spread away from the base rate, which is the credit the three-term decomposition gives a forecaster for telling one case from another.
The liar’s two reports define the partition by distinct forecast value, which is the one partition on which the decomposition is an identity rather than an approximation, and the identity holds exactly: on the liar’s two values the three terms sum to its Brier score with a residual of zero. Its resolution is 0.070604 against the honest forecaster’s 0.092758. The difference, 0.022154, is information.
The obvious rescue is to recalibrate: replace the liar’s 0 by the frequency of the event among the forecasts that said 0, which is 0.1844, and its 1 by the frequency among those that said 1, 0.7459. That relabelling removes the liar’s entire reliability term, because the new reports are the event rates of their own groups, and it is the best any relabelling of two answers can do. It reaches a Brier score of 0.163633. The honest forecaster’s Brier score is 0.141479. The gap between them is 0.022154 — exactly the resolution the liar lost, to the last digit, because recalibration repairs reliability and cannot touch resolution, and a record of two values has only two groups to resolve events into.
The two scores rank the pair in opposite orders, and the reason is visible in the same table. For a report of 0 or 1 the Brier score and the absolute-error score are the same number, the share of events the report gets wrong, 0.207973 for the liar. The honest forecaster is charged 0.141479 by Brier and 0.282958 by absolute error, because absolute error charges an honest 0.3 on an event that happened 0.7 and Brier charges it 0.49. Squaring rewards being close; the absolute value rewards being at an end.
Why the square cannot be gamed the same way
It is worth being exact about why the Brier score does not produce this liar, because the reason is also why the liar’s two answers cannot be repaired into an honest forecast. The Brier score charges a report r on an event of probability p an expected . That is a parabola in r, and its minimum is where its slope is zero, at r = p — the truth, for every p. The absolute-error score’s expected charge, , is a straight line in r, and a straight line on an interval has its minimum at an end. The whole difference between a score that elicits a probability and one that elicits a certainty is whether the expected charge curves.
The curvature is also what makes the Brier score reward the information the liar discards. Among forecasts that are all calibrated, the Brier score is lower by exactly the resolution — the honest forecaster’s 0.092758 against the relabelled liar’s 0.070604 — because separating events into more groups with more different rates always reduces an expected squared charge. A straight-line charge has no such property. It rewards moving each report to the nearer end whatever that does to the separation of events, which is why it can pay a forecaster for converting a well-resolved forecast into a coarse one, and why the choice between two proper scores is a different kind of choice from the choice between a proper score and this one.
What a reliability diagram shows of a two-valued record
A reliability diagram groups forecasts by their value and plots each group’s mean forecast against the share of its events that happened. For a record with two values the diagram has two points and nothing else, and both of them are far from the diagonal: the forecasts of 0 carry an event rate of 0.1844, and the forecasts of 1 carry 0.7459. The liar’s reliability term — the charge for saying the wrong probability, averaged over the two groups — is 0.044340, where the honest forecaster’s is zero.
So the liar is easy to catch on a diagram and on the decomposition, and the earlier essay said as much. What this adds is how the two points are placed. The liar says 1 on 33.8% of its forecasts against a base rate of 37.4%, so its mean forecast is below the base rate even though it says “certainly” a third of the time: it is overconfident about the events it flags and silent about the rest. A diagram drawn with a binning chosen by the analyst would put both of these points into its first and last bins and report them as the two most miscalibrated groups on the page — correctly, and without any indication that the miscalibration was the score’s instruction rather than the forecaster’s mistake.
Three summaries, three best thresholds
The liar thresholds at a half because that is what the absolute-error score asks for. A forecaster that is going to issue only two answers anyway has a choice of threshold, and the choice reveals that the summaries of a two-valued forecast disagree about what it should be.
The share of events misjudged — which is the absolute-error score and the Brier score on answers of 0 and 1 — is smallest at a threshold of 0.5000, exactly a half, found by golden section on the closed form. The area under the ROC curve is largest at 0.3744, which is the base rate: at that threshold the signal’s likelihood ratio is one, so moving the threshold trades a true positive for a false positive at par, and the area stops improving. The best two-valued area there is 0.7818, still 0.0866 short of the honest forecaster’s. Resolution is largest at 0.4350, strictly between the two, where it reads 0.071585.
So the liar is the best two-valued forecaster by one summary and not by the other two, and the summary it wins on is the one that charges an honest probability for not being at an end. A threshold is a decision about which errors to prefer — the positive predictive value a screening cut buys is the same arithmetic — and an improper score makes that decision for the forecaster, silently, in favour of the errors it happens to count.
The ranking lost at every signal strength
The numbers so far are at the field’s full signal. The world lets the signal’s correlation with the latent state be dialled from nothing to full, and the loss can be read at every setting.
At full signal the honest, best two-valued and lying forecasters read 0.8683, 0.7818 and 0.7684; at a correlation of a half, 0.6773, 0.6272 and 0.5931. The liar gives up 0.0999 of area at full signal and 0.0718 at a correlation of a quarter, where something else happens: its honest probability crosses a half on only 4.8% of forecasts, so it says 0 nearly every time and ranks almost nothing. A liar with a weak signal is not a coarse forecaster; it is very nearly a constant.
That ordering — honest above best two-valued above liar — holds at every informative signal strength, and the gap between the last two is the price of the threshold the score chose, on top of the price of having two answers at all.
The score pays most for knowing nothing
If the lie costs information, a natural hope is that the score at least pays it less as the information it destroys grows. It does the opposite.
The payment for lying is 0.094025 for a forecaster whose signal carries nothing, 0.086486 at a correlation of a quarter, 0.087571 at a half, 0.088263 at 0.7, 0.084406 at 0.85 and 0.074986 at full signal. It is largest where the forecaster knows nothing, and there it has a closed form: an uninformed honest forecaster issues the base rate p̄ every time and is charged 2p̄(1 − p̄), while saying 0 every time is charged p̄, so the payment is 2p̄(1 − p̄) − p̄ exactly, 0.094025 at this base rate. The forecaster that earns the largest reward for gaming the score is the one whose lie destroys the least, because it had the least to destroy — and the reward is paid for converting a calibrated base rate into a constant that is wrong about every event that happens.
An honest forecast that loses to saying no
The last consequence follows from the one before, and it is the one that would matter most to a reader choosing a forecaster by this score.
A forecaster that says 0 every time is charged the base rate, 0.3744, and ranks nothing. An honest forecaster is charged , which falls as its signal improves. The two tie at a signal correlation of 0.7332, found by bisection. Below that, the absolute-error score prefers the constant to the honest forecast — and at the tie the honest forecaster’s area under the ROC curve is already 0.7636 and its resolution 0.047012. A forecaster that ranks events well enough to be genuinely useful loses, on this score, to one that has never looked at the data.
The counted version is as unanimous as the earlier essay’s. On two hundred records of two thousand forecasts from an honest forecaster with a correlation of a half — an area of 0.6773 — the constant zero scores 0.374405 against the honest forecaster’s 0.425499, closed forms 0.374449 and 0.425634, and the constant wins on 200 of 200 records. An evaluation built on this score would not merely fail to reward the better forecaster; it would select against it with certainty, in the region where calibrated forecasts differ most in usefulness.
Which of these numbers are proved and which are counted
Proved. A two-valued forecaster’s ROC area is the mean of its two rates, which is the area of two line segments. The decomposition is an identity on the partition by distinct forecast value, so it is exact on the liar’s two answers, and recalibration removes reliability and leaves resolution, so the best relabelling falls short of the honest Brier score by exactly the resolution lost. The misjudged share is minimised at a half and the area at the threshold where the likelihood ratio is one, and the payment to an uninformed forecaster is 2p̄(1 − p̄) − p̄; each is a line of algebra, and the file checks each at machine precision rather than taking it on trust.
Computed in closed form for this world. Every rate, area, resolution and score at every threshold and signal strength, from bivariate normal orthants and one-dimensional quadrature against the world’s own law — including the tie at 0.7332 and the resolution-optimal threshold at 0.4350, which are properties of this world’s base rate and signal family rather than general constants.
Counted. The areas by ranks, the misjudged share and the win counts on two hundred records of two thousand forecasts, each agreeing with its closed form within a few of its own standard errors.
What an evaluation of probability forecasts should carry
A proper score, and not only one number. The earlier essay settled that the absolute-error score must go. What this essay adds is that even a proper score is one number, and the liar’s two failures — lost ranking and lost resolution — are visible only when discrimination and resolution are reported beside it.
The number of distinct values a forecaster issues. A liar’s record has two. It is the cheapest diagnostic available, it needs no outcomes at all, and a record of a thousand forecasts with two distinct values is a record whose ranking has been capped whatever its score says.
The comparison against a constant. A forecaster that cannot beat saying the base rate every time on a proper score has shown no skill; one that cannot beat saying 0 every time on an improper score may have shown nothing but the score’s preference, and a benchmark whose place in the comparison is fixed is what distinguishes the two.
Where this goes next: a forecaster that rounds
The liar has two answers because an improper score asked for them. Most real forecasters have a few answers for a different reason: forecasts are issued on a coarse scale — probabilities of precipitation in tenths, risk categories in five bands, credit grades — and rounding to a grid is a step function too, only with more steps. It is not a lie, because under a proper score the forecaster rounds its honest probability rather than pushing it to an end, but it discards the ordering inside each band exactly as the liar discards it inside each half.
The measurement that would settle how much coarse reporting costs is this essay’s table with the two answers replaced by k equal-width bands: the ROC area and resolution lost against k, the Brier score of the rounded forecast against the honest one, and the band count at which the loss falls under what a reliability diagram’s own binning already costs to read. It is distinct from this essay because the rounding is incentive-compatible — no score pays for it — so the question is purely what a reporting format throws away, and the answer is a number of bands.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The test is a point somebody chose — both name base rate, roc curve, sensitivity and specificity
- The miscalibration a perfect forecaster shows — both name probability forecast, reliability term
- The prevalence the test has to estimate — both name base rate, sensitivity and specificity
- The second test that is not a second opinion — both name base rate, sensitivity and specificity
Named objects
A flat tag is an object no other essay names yet.
Base rateBrier scoreClosed formDiscriminationImproper scoring ruleMurphy's decompositionProbability forecastProper scoring ruleRecalibrationReliability termResolution termROC curveSensitivity and specificity