The test is a point somebody chose
Worth reading first: What a positive test is worth.
The predictive value’s own sweep holds a test fixed at 90% sensitive and 95% specific and sweeps the prevalence across four orders of magnitude. That is one axis of a plane, and the other one is the test.
The test is not fixed either. It is a continuous score with a line drawn on it, and the two numbers are where the line is.
One parameter, not two
Almost every diagnostic test is a measurement — a concentration, an optical density, a count, a model score — compared with a cut-off. Above the cut-off is positive.
So sensitivity and specificity are not independent characteristics that a test happens to have. They are
two readings of one distribution pair at one value of c. Moving c moves both, in opposite directions, along a curve that is fixed by how far apart the two score distributions are.
With normal scores separated by d standard deviations, that separation is the whole of the test’s quality, and every operating point is available. The 90/95 quoted throughout corresponds to d = 2.926 with the cut at 1.645; the same assay reporting a different pair is the same assay with the line moved.
Nothing about the published pair says which point was chosen or why. A validation study reports the operating characteristics at the manufacturer’s threshold, and the threshold was chosen by the manufacturer, against a population and a cost ratio that are not the reader’s.
Which point is best
The threshold is a decision, so it needs a loss. Let a missed case cost k times a false alarm. Expected cost per person tested is
Minimising that over the cut gives a different answer at every prevalence, and the answers are a long way apart.
| prevalence | best cut | cases detected | healthy cleared | predictive value |
|---|---|---|---|---|
| 1 in 10,000 | 3.05 | 45.0% | 99.89% | 3.8% |
| 1 in 1,000 | 2.26 | 74.7% | 98.81% | 5.9% |
| 1 in 100 | 1.47 | 92.8% | 92.89% | 11.6% |
| 1 in 20 | 0.91 | 97.8% | 81.76% | 22.0% |
| 1 in 5 | 0.38 | 99.5% | 64.73% | 41.3% |
| 1 in 2 | −0.12 | 99.9% | 45.32% | 64.6% |
At a miss costing a hundred false alarms, the best test on a one-in-ten-thousand condition detects 45% of the cases it is looking for. That is not a defect of the test; it is the correct answer to the question of how many false alarms a rare condition’s cases are worth.
At one in two, the best test clears 45% of healthy people and flags the rest. It detects essentially every case, and it is useless as a rule-out.
The direction is the one people find backwards
A rare condition wants a stricter threshold and therefore detects a smaller share of its cases. The instinct runs the other way — a rare and serious condition feels like one to be vigilant about — and the arithmetic disagrees.
The reason is the denominator. At one in ten thousand, every percentage point of false-positive rate costs a hundred false alarms per case in the population. Loosening the threshold to catch one more case in ten thousand buys one detection and costs hundreds of alarms, so it is worth doing only when a miss costs hundreds of alarms.
That is exactly the same asymmetry the predictive value’s own curve shows, arriving as a design decision rather than as an interpretation. There the reader is told that a positive on a rare condition means little; here the consequence is that the test should be designed differently, and a test designed for a common condition is the wrong test for a rare one at any threshold.
The cost ratio moves it too, and less than one expects. Taking the miss from 3 false alarms to 1,000 — a factor of three hundred — moves the best cut at one in ten thousand from 4.24 to 2.26 standard deviations, which is about two. The prevalence, over the same table’s range, moves it by three. The base rate matters more than the loss function, which is the standing result of every screening calculation appearing in the one place a reader would expect the loss to dominate.
What the published point costs
The gap between the two marks on the first figure is a number, and it is worth having because it decides whether any of this is actionable.
At a prevalence of one in a thousand with a miss costing a hundred false alarms, the published 90/95 operating point costs 0.0600 per person tested on the scale where a false alarm costs 1. The harm-minimising point costs 0.0372. So using the threshold the manufacturer chose rather than the one this population wants costs 61% more harm, and the fix is to report a different cut-off on a score that has already been measured.
That is an unusually cheap improvement, and it is worth saying what it is not. It is not a better test, a larger study or a new method: it is the same measurement on the same people, compared with a different number. The only thing standing between the two rows is that the threshold is embedded in the assay’s instructions rather than in the analysis.
The size of the gap depends on how far the published point is from the right one, and the right one depends on a prevalence that varies between populations by orders of magnitude. A threshold set for a clinical population and applied to a screening population is the common case, and it is the case where the gap is largest — a clinical population has a prevalence in the percents and a screening one in the tenths of a per cent, which is three rows of the table apart.
What a screening programme is doing
The table is the design argument for two-stage screening, and it makes the design obvious rather than clever.
A population screening programme faces a low prevalence, so the harm-minimising single test is strict and detects under half the cases. That is unacceptable for a condition worth screening for. The way out is not a better threshold; it is a different structure.
First stage: a loose threshold. Detect nearly every case and accept a large false-positive rate. This stage’s job is to reduce the population, not to diagnose.
Second stage: a confirmatory test, on a much higher prevalence. Among first-stage positives the prevalence is the first stage’s predictive value — 2% rather than 0.1% — and the table says a test on a 2% prevalence can be operated at a much looser threshold while still being useful.
The arithmetic of the second stage is where the independence assumption lives, and this is why it matters so much: the entire design rests on the second test’s prevalence being the first one’s predictive value, which is only true if the second test is genuinely new information.
What the curve cannot be improved past
The separation d is the test’s quality and the threshold is the only thing a user controls. That gives a clean division of what can be fixed by whom.
An analyst can move along the curve. Any pair of (sensitivity, specificity) on it is available by reporting a different cut-off, at no cost and with no new data. A published pair that is wrong for a population is a one-line fix if the underlying score is reported.
Only a better assay moves the curve. Getting 95% sensitivity and 95% specificity out of a test with d = 2.926 is impossible — the curve does not pass through that point — and no threshold, no reporting convention and no statistical method changes it.
The practical consequence is that the score should be reported and not only the verdict. A test that returns “positive” has thrown away the reader’s ability to choose their own threshold; one that returns the score has kept it. That is a reporting decision with no cost, and the reason it is not universal is that a verdict is easier to act on — which is the argument for a threshold in every other part of this subject, and it loses here for the same reason it loses there: the threshold that is right for the person who set it is not right for the person reading it.
Where the same structure appears without a disease in it
The threshold is a decision rule on a score, and nothing above used the word “disease” in an essential way. The same arithmetic governs every rule of that shape, and the cases differ only in what the two costs are.
A spam filter has a prevalence, two costs that are wildly asymmetric — a missed spam is an annoyance and a blocked letter is a disaster — and a threshold chosen by the vendor for an average user. The right threshold for someone whose incoming mail is 95% spam is not the right one for someone whose incoming mail is 5% spam.
A fraud rule on card transactions faces a prevalence near one in a thousand and costs that differ by a factor of hundreds, which by the table above wants a very strict threshold and a detection rate under half. That is what fraud rules do, and it is routinely reported as a failing.
A p-value threshold is the same object one level up: a score, a cut, and two error rates traded against each other, with the family it belongs to deciding where the cut should be. What a p-value does not say is partly this — 0.05 is a threshold chosen once for a general case, and the base rate of true hypotheses in a field is its prevalence.
That last analogy is worth taking seriously rather than as a flourish. A field in which one hypothesis in a thousand is true is a field screening a rare condition, and the table says its threshold should be very strict and its detection rate low. A field in which half the hypotheses are true can afford 0.05. Neither field’s threshold is derived from its own base rate, and both use the same number.
A summary of the curve is not a substitute for the curve
The standard summary is the area under the operating characteristic, and it deserves a paragraph because it is the number that gets compared between tests.
The area is a property of the separation d alone here, so it is exactly the test’s quality with the threshold integrated out — which is its virtue and its defect. Two tests with the same area are equally good averaged over all thresholds, including the thresholds nobody would use, and one of them can be much better than the other over the range a particular application operates in.
That matters most in the corner this page is about. A screening application lives at very high specificity, where the curve is nearly vertical and small differences in shape are large differences in detection. Two tests with identical areas can detect 45% and 60% of cases at a specificity of 99.9%, and the summary reports them as equivalent.
It is the same complaint as R-squared one field over: a single number standing in for a curve, comparable across studies, and silent about the part of the curve anyone is standing on. The repair is the same too — report the operating characteristic, or at least the point that matters.
What is claimed here, and what is not
Two statements, both over the sweep rather than at a point.
The minimising threshold costs no more than the published operating point, at every prevalence the slider takes. That is nearly a tautology — it is a minimum — and it is worth stating because the plausible slip is a cost evaluated at the wrong threshold, which would put the “minimum” above the published point and still look reasonable.
A rarer condition wants a stricter threshold and detects a smaller share of its cases, as a monotone relationship across the whole prevalence range. It would reverse if the prevalence entered the cost with the wrong sign, which is the other plausible slip and would turn the table over without making any individual number impossible.
The reading that does not survive is a published operating point taken as a property of the test. That reading needs the harm-minimising threshold to be the published one, and at a prevalence of one in a thousand the published 90/95 is beaten by a 74.7/98.81 operating point costing 0.0372 against its 0.0600 — 38% less harm from moving a line. A published pair is an answer to a question with two inputs, and neither input is printed beside it.
What to ask of a test report
The practical content of this page is four questions, and all four are answerable from a validation study that already exists.
What is the score, and what is the cut-off? If the report gives a verdict without a score, every choice on this page has been made for the reader and none of them was made with their population in mind.
What prevalence was the cut-off chosen against? Usually the validation population’s, which is usually enriched for cases — a validation study recruits known positives deliberately, so its prevalence is often 30% or 50% rather than the field’s 0.1%. A threshold chosen there is in the bottom rows of the table and is being applied in the top rows.
What were the two costs? Almost never stated, and frequently implicit at 1:1, which the table says is right for a prevalence near a half.
What does the curve look like near the operating point the reader needs? The area under the curve does not answer this, and the curve itself does.
A reader who has those four can recompute the threshold in a line. A reader who has a verdict has nothing, and the difference between the two reports is a column of numbers that the laboratory already has.
The same four questions apply to any automated decision that arrives as a label rather than a score, which is most of them. A model that returns “high risk” has embedded a prevalence and a cost ratio in a place nobody can inspect, and the label is only the right label for whoever chose them — which is the rule being part of the data in the form it takes when the rule is somebody else’s.
Still open: the score distributions that are not normal
Every curve here comes from two normal score distributions with equal variances, which makes the operating characteristic a one-parameter family and the arithmetic clean.
Real score distributions are neither normal nor equally spread. The diseased group is usually more variable — a mixture of severities — and often skewed, and the healthy group is frequently bounded below by a detection limit. Unequal variances make the curve cross the diagonal somewhere, which means there are thresholds at which the test is worse than a coin, and asymmetric distributions make the harm-minimising cut move in ways no closed form describes.
None of that changes the qualitative argument: the threshold is still a choice, the best choice still depends on the prevalence, and the published pair is still one point. What it changes is whether the table transfers, and it probably does not — a test whose diseased scores are a two-component mixture has an operating characteristic with a shoulder in it, and the harm-minimising threshold can jump discontinuously as the prevalence moves.
The measurement that would settle it needs real score distributions rather than a model, which is the one input that cannot be generated from a rule.
What can be done from here is to sweep the unequal-variance case, which is a two-parameter family rather than a one-parameter one and is still generated from a stated rule. The question it would answer is whether the harm-minimising threshold moves continuously in the prevalence when the variances differ, or whether it jumps — and a threshold that jumps is a threshold that cannot be recommended as a formula, which would change the practical advice on this page from “compute it” to “compute it and check the second-best point”.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The base rate was always Bayes — both name base rate, predictive value, screening
- A score that rewards lying — both name base rate, roc curve
- Calibrated and useless — both name base rate, roc curve
- The measurement that got them enrolled — both name screening, threshold
Named objects
A flat tag is an object no other essay names yet.
Base rateDecision thresholdExpected lossPredictive valueROC curveScreeningSensitivity and specificityThreshold