Concept

Base rate — where it appears

How common a condition is before any test is run, which is what a positive result has to be combined with before it means anything. At a low enough base rate most positives from an accurate test are false, and the arithmetic is Bayes' rule rather than a fallacy.

Named by 10 essays across 4 fields — each of them below, with the objects they name alongside it.

What each forecaster says, and what is true. The true probability of the event given a forecaster's signal, and what three forecasters report. The truth is Φ(-0.5 + 1.2w) and the honest forecaster reports it, so its curve and the truth are the same line. The loud forecaster pushes every probability towards the ends and the hedged one pulls every probability towards the middle; both have their mean report held at the base rate of 0.3744, so each crosses the truth exactly once and neither can be caught by checking its average. Their reliability terms are 0.008500 and 0.013025 against the honest forecaster's zero, and all three have the same area under the ROC curve, 0.868312.

An identity in three terms

Reliability minus resolution plus uncertainty is quoted as a rewriting of a probability score. It is an identity to 2.6·10⁻¹⁵ on the one grouping where reliability is the whole score and resolution exactly cancels uncertainty, and it is out by 0.004125 on the coarsest grouping anybody would actually draw.

calibrate · Calibration
What a second positive is worth, prevalence 0.10%. A 90% sensitive, 95% specific test. One positive gives 1.77%. Two independent positives give 24.49%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 10.16%, and at 0.5 it is 3.16%.

The second test that is not a second opinion

Two positives from a 90/95 test on a one-in-a-thousand condition give a 24.49% chance of disease if the tests are independent. At a correlation of 0.1 between their errors it is 10.16%, and at 0.5 it is 3.16% — barely more than the 1.77% one positive was worth.

screening · Baserate
What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.

What a positive test is worth

A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.

paradox · Baserate
One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.

A curve that is a binning

A forecaster with no miscalibration in it at all reads 0.001429 at five bins and 0.014100 at fifty, on the same five hundred forecasts. The closed form is K/n times the forecaster's own irreducible score, and subtracting it returns zero.

calibrate · Calibration
The operating characteristic, and the point that minimises harm at 0.10% prevalence. The published pair — 90% sensitive, 95% specific — is the open mark. With a miss costing 100 times a false alarm and a prevalence of 0.10%, the threshold that minimises expected cost sits at 74.7% sensitivity and 98.81% specificity, with a predictive value of 5.9%.

The test is a point somebody chose

A test reported as 90% sensitive and 95% specific is not two properties of a test. It is one property read at a threshold, and the threshold that minimises harm runs from 3.05 standard deviations of the score at a prevalence of one in ten thousand to −0.12 at one in two — 45% of cases detected at one end and 99.9% at the other.

screening · Baserate
Six forecasters, all calibrated, not equally useful. The resolution of six forecasters that are all perfectly calibrated, each reporting the true probability of the event given a signal that carries more or less of the latent state. The largest reliability anywhere in the family is 2.0e-33, so a calibration check passes every one of them. They are not equally good: resolution runs from exactly 0.00 for the forecaster that issues the base rate every time to 0.092758 for the one that sees everything, and their Brier scores run from 0.234237 — which is the world's own uncertainty, and the score of a table of base rates — to 0.141479. Calibration is a necessary condition that a constant forecast satisfies exactly.

Calibrated and useless

Six forecasters that are calibrated to 2·10⁻³³ run from resolution exactly 0 to 0.092758, and three forecasters with reliabilities from 0 to 0.013025 have areas under the ROC curve identical to every bit a double carries. Each measure is exactly blind to what the other one sees.

calibrate · Calibration
Estimating a prevalence of 0.10% from 1,000 tests. The positive rate reads 5.09%, which is 50.9 times the truth. The Rogan-Gladen correction averages 0.100% — unbiased — with a standard deviation of 0.820 points against the positive rate's 0.695, and it comes out negative on 48.6% of samples.

The prevalence the test has to estimate

Every predictive value takes a prevalence as given, and the prevalence is usually estimated from the same test's positive rate. At a true prevalence of one in a thousand that rate reads 5.09% — fifty times the truth — and the correction that inverts it is unbiased, 18% more variable, and negative on 48.6% of samples of a thousand.

screening · Baserate
What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.

The base rate was always Bayes

The screening arithmetic everybody finds counter-intuitive is a posterior update with a prior of one in a thousand. Naming it that way turns a famous puzzle into an instance of a rule, and makes the sequential version obvious.

bayes · Baserate
Where each rule says to put the number. The expected score of reporting each probability on the axis when the event's true probability is 0.25, for four scoring rules. The Brier, logarithmic and spherical scores each bottom out at 0.25 — found by search rather than assumed, to 8 decimal places — which is what makes them proper: a forecaster with a genuine belief cannot improve its expected score by reporting anything else. The absolute-error score is a straight line in the reported value, p + r(1 − 2p), so it has no interior minimum at all; its optimum is 0.0, a distance of 0.250 from the truth, and taking it saves 0.125.

A score that rewards lying

An absolute-error score pays a forecaster exactly ⅛ of a point to replace a true quarter with a zero, and over two hundred records a liar beats a truthful forecaster on 200 of 200. A skill score against the forecaster's own average buys 0.012633 of reported skill for 0.002035 of real score.

calibrate · Calibration
One curve, and two forecasters that are each a single point. The ROC curve of an honest forecaster whose signal has correlation 1 with the latent state — every threshold on its probability, from the bivariate normal — and two forecasters that only ever say 0 or 1. The one an absolute-error score pays for says 1 wherever the honest probability exceeds a half: it has a true-positive rate of 0.6742 and a false-positive rate of 0.1375, and its "curve" is the two straight segments through that point, with area (TPR + TNR)/2 = 0.7684. The honest curve's area is 0.8683. Thresholding at the base rate of 0.3744 instead of at a half gives the largest area any two-valued forecaster can have here, 0.7818, and it is still 0.0866 short.

The liar with two answers

The forecaster an absolute-error score pays for says only 0 or 1, and on the ROC square it is a single point: its area is (TPR + TNR)/2 = 0.7684, against the honest forecaster's 0.8683, and it falls below the honest one on 200 of 200 counted records. No relabelling of its two answers returns what it threw away — the best recovers a Brier score short of the honest one by exactly the 0.022154 of resolution lost — and below a signal correlation of 0.7332 the same score prefers saying no every time to an honest forecast.

calibrate · Calibration

Named alongside it

The objects these essays reach for when they reach for this one.

Brier scorePredictive valueProbability forecastScreeningSensitivity and specificityForecast calibrationReliability termROC curveClimatological forecastDiscriminationMurphy's decompositionProper scoring rule

All concepts