Reversals that are not errors

What a positive test is worth

A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.

“The test is 90% accurate, and it came back positive.” The natural conclusion is that there is a 90% chance the condition is present. The actual figure, for a condition affecting one person in a thousand, is under 2%.

A thousand people, a condition one in a hundred has, a test that is 90% and 95%9 people have it and test positive. 50 do not have it and test positive anyway. So of the 59 positive results, 15% are right — and that is with a test most people would call accurate.9 have it, test positive50 do not have it, test positive1 have it, test negative15% of positives are trueeach dot is one personthe arithmetic is not in dispute
Fig. 1 A thousand people, a condition one in a hundred of them has, and a test most people would call accurate. Each dot is a person.

Counting people

The dot figure is the whole argument and it needs no algebra.

Out of a thousand people, ten have the condition. The test catches nine of them — it is 90% sensitive. Of the 990 who do not have it, the test wrongly flags about 50 — it is 95% specific, so it gets 5% of them wrong.

So there are 59 positive results, and 9 of them are right. The chance that a positive result is correct is 15%.

Nothing has gone wrong with the test. It performed exactly to specification. The arithmetic is not in dispute and can be checked by counting the dots.

Which allocations reverse the overall comparisonThe per-group success rates are held fixed; only the split of each group between treatment and control changes. 32% of the allocations reverse, and the worst reverses by 13.1 percentage points.fraction of group A given the treatmentgroup B0.00.51.032% reverseshaded: the overall comparison reversesthe rates are identical everywhere here
Fig. 2 A related reversal, swept the same way: the region of the parameter space where the aggregate contradicts its parts.

Why the intuition fails

The intuition swaps two conditional probabilities that are not the same.

The test’s sensitivity is P(positive | condition) — of the people who have it, how many does the test catch. What a patient wants is P(condition | positive) — of the people who test positive, how many have it.

Bayes’ theorem relates them, and the relation involves the prevalence. When the condition is rare, the pool of healthy people is enormous compared with the pool of sick ones, so even a small false-positive rate produces a large false-positive count — and that count swamps the true positives.

The confusion has a name, the base rate fallacy, and naming it does not stop it. The dot picture stops it, which is why it is worth drawing.

The whole curve

Fixing the prevalence at one value gives one number and hides the structure. Sweeping it across four orders of magnitude shows the shape.

What a positive test means, sensitivity 90%, specificity 95%At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in ten, 33 are. The test has not changed.00.2500.5000.7501how common the condition ischance a positive result is true1 in 10,0001 in 1,0001 in 1001 in 101 in 11.8%66.7%one test, every prevalencethe base rate outweighs the test
Fig. 3 The predictive value of a positive result against how common the condition is. The test is unchanged along the whole curve.

It is a cliff. At a prevalence of one in ten a positive result is worth something; at one in a thousand it is worth very little; and the transition is fast.

The measured comparison that makes the point hardest: a worse test on a commoner condition beats a better test on a rare one. At 90% sensitivity and 95% specificity with prevalence 0.1, the predictive value is 67%. At 99% and 99% with prevalence 0.001, it is 9%.

Improving the test by a lot loses to a difference in base rate.

Two groups, one treatment, and both readings of the same numbersThe treatment wins in group A (93.0% against 87.0%) and in group B (73.0% against 69.0%), and loses overall (78.1% against 82.6%). Nothing here is a trick; the allocation differs between the groups.treatedcontrolverdictgroup A93.0%87.0%treatmentgroup B73.0%69.0%treatmentboth together78.1%82.6%controlgroup A: 88 treated of 351 · group B: 263 of 351wins in both groups, loses overallthe rates never changeonly the allocation does
Fig. 4 The same lesson about weighting: an aggregate is a weighted average, and the weights are not the thing being measured.

What this means for screening

The consequence is not that testing is useless. It is that screening a low-prevalence population is a different activity from testing someone who presented with symptoms, and the same test means different things in the two settings.

Someone with relevant symptoms has, in effect, a much higher pre-test probability — they are drawn from a subpopulation where the condition is far commoner. The same positive result is worth much more for them than for a randomly screened person, because the prevalence in the denominator has changed.

Move the specificity slider and a second point emerges: specificity is what matters for a rare condition. Raising sensitivity from 90% to 99% barely moves the curve at low prevalence, because the problem is not the cases being missed — it is the healthy people being flagged. Raising specificity from 95% to 99.9% transforms it.

That is a design conclusion with a number behind it: for a screening test aimed at a rare condition, specificity is the parameter to buy.

Two measurements of the same thing, correlated 0.60Pick the worst 15% on the first measurement and their average rises by 0.55 on the second. Pick the best and theirs falls by 0.59. No treatment was given to anybody.first measurementsecond+0.55-0.59predicted 0.64 from the correlation alonenobody was treated
Fig. 5 A third case where a correct calculation predicts an outcome intuition does not.

The harm that follows

Worth stating plainly, since this is one of the places where a statistical misunderstanding has direct consequences.

A false positive is not a neutral event. It leads to further investigation, which carries its own risks; to treatment that may be unnecessary; and to a period of fear that is real whatever the eventual outcome. When 98% of positives are false, those harms are being delivered overwhelmingly to people who do not have the condition.

That is the argument behind screening guidelines that recommend against testing low-risk populations, which are routinely reported as rationing and are usually this calculation. Whether a given programme is worth it depends on the harms and benefits of each outcome, which is a question the arithmetic does not settle — but the arithmetic sets out what the outcomes will be, and it is not what intuition predicts.

What a one-sample t test at n = 20 can detectAt an effect of 0.5 standard deviations the test finds it 56% of the time. Below that, a non-significant result is the expected outcome of a real effect — which is why "no significant difference" is not evidence of no difference.00.2500.5000.750100.5001true effect, in standard deviationsprobability of rejecting80%, the usual target56% at d = 0.54,000 simulated studies per pointthe second number a p-value needs
Fig. 6 The same reversal in inference: the probability of the data given no effect is not the probability of no effect given the data.

The general rule

A conditional probability read in the wrong direction is not an approximation to the right one. P(A|B) and P(B|A) can differ by orders of magnitude, and the factor between them is the ratio of base rates.

The same reversal underlies the misreading of a p-value: P(data | no effect) is not P(no effect | data), and converting between them needs a prior probability that the p-value does not contain. That is a separate essay, and it is the same mistake in different clothes.

The odds form, which is easier to carry

Bayes’ theorem in probabilities is hard to compute in the head. In odds it is a multiplication, and the multiplication makes the base rate’s dominance obvious.

Posterior odds = prior odds × likelihood ratio.

For a positive result the likelihood ratio is sensitivity ÷ (1 − specificity). At 90% and 95% that is 0.9 / 0.05 = 18.

So a positive result multiplies the odds by eighteen, whatever they were. At a prevalence of 1 in 1000 the prior odds are 1:999, and eighteen times that is 18:999 — about 1.8%. At 1 in 10 the prior odds are 1:9, and eighteen times that is 2:1 — about 67%.

The same test, the same multiplier, wildly different conclusions, because the thing being multiplied differs by a factor of a hundred. Stated this way the fallacy is hard to commit: a multiplier cannot say where the odds end up without saying where they started.

Why specificity is the parameter to buy

The odds form also settles the design question the slider raises.

The likelihood ratio for a positive is sensitivity ÷ (1 − specificity). Sensitivity appears in the numerator, where it can at most double the ratio by going from 0.5 to 1. Specificity appears in the denominator as (1 − specificity), which can go from 0.1 to 0.001 — a factor of a hundred.

So for ruling a condition in, specificity does almost all the work, and improving it is worth far more than improving sensitivity. For ruling a condition out the reverse holds, because the negative likelihood ratio is (1 − sensitivity) ÷ specificity.

That is the whole of the classic mnemonic — a specific test rules in, a sensitive test rules out — derived rather than memorised, and it explains why screening programmes use a sensitive test first and a specific one to confirm.

The screening sequence, which is what actually happens

Real programmes rarely rest on one test, and the two-stage structure is a direct response to the arithmetic above.

The first test is chosen to be sensitive: it must not miss cases, and it is allowed a high false-positive rate. Everyone positive goes to a second, more specific test.

The point is that the second test is applied to a population with a completely different base rate. Among first-test positives the prevalence is no longer 1 in 1000 — it is the predictive value of the first test, which might be 15%. The second test operates on prior odds a hundred times better, and its positive result is worth correspondingly more.

That is the base-rate argument used constructively rather than as a warning, and it is why the sequence works when neither test alone would.

What the numbers do not decide

The essay’s arithmetic says what a positive result is worth. It does not say whether the programme is worthwhile, and the gap between those matters.

Deciding requires the consequences of each of the four outcomes: a true positive caught early, a false positive investigated unnecessarily, a false negative reassured wrongly, a true negative. Those are measured in health, in anxiety, in money and in risk from the follow-up procedures, and weighing them is not a statistical exercise.

What the arithmetic supplies is the counts — how many of each will occur per thousand screened. That is the input the weighing needs, and it is the part that is routinely wrong in public discussion, where a test described as 95% accurate is assumed to produce overwhelmingly correct positives.

Getting the counts right does not settle the argument. It makes the argument about the right thing.

The same arithmetic outside medicine

The essay uses a medical test because the numbers are familiar. The structure appears wherever a rare thing is screened for, and the consequences are the same.

Fraud detection. A model flagging transactions at 99% specificity, applied to a population where fraud is one in ten thousand, produces overwhelmingly false alerts. The usual response is to lower the alert rate, which lowers sensitivity and misses fraud.

Security screening. Rare threat, enormous population, and a base rate that makes almost every alarm false however good the detector.

Automated content moderation, rare-disease genetic screening, predictive policing — all the same shape, and in each case the false positives fall on people who did nothing.

The arithmetic is not in dispute in any of these. What differs is whether the harm of a false positive is counted, and the base-rate calculation is what makes the count possible.

When the base rate is unknown

The essay assumes the prevalence is known. Often it is not, and the honest treatment of that matters more than the arithmetic.

Two responses are available and one of them is bad.

Bad: assume a prevalence that makes the test look useful, usually implicitly, by not mentioning it. That is how a test’s accuracy comes to be reported as though it were the predictive value.

Good: report the predictive value across a range of plausible prevalences, which is the curve in the figure. A reader who knows their own population’s rate can then read off the answer, and a reader who does not can see how much it matters.

That is a general habit for any quantity depending on an unknown parameter: sweep it and show the dependence, rather than choosing a value and reporting a single number that hides the choice.

The individual against the population

A distinction that does real work when this arithmetic is applied to a person.

The prevalence in the calculation should be the rate in the population the individual is drawn from, and that is rarely the general population. Someone with symptoms, a family history, or an occupational exposure belongs to a subgroup with a much higher rate.

So the same test result genuinely means different things for two people, and the difference is not a subjective adjustment — it is a different denominator.

This is what clinicians mean by pre-test probability, and it is the reason a screening result in an asymptomatic person and a diagnostic result in a symptomatic one are treated differently. Both are Bayes’ theorem with different priors, and the priors come from clinical knowledge rather than from the test.

Which is the honest limit of the arithmetic here: it says exactly how to combine a prior with a test result, and it supplies no priors.