The threshold a validation study can see
Worth reading first: An identity in three terms.
A threshold that jumps found that when a test’s diseased scores are more spread than its healthy ones, the threshold that minimises expected harm does not move smoothly with the prevalence. With a miss costing a hundred false alarms and diseased scores twice as spread, the best cut flags 81.1% of healthy people just below a prevalence of 25.69% and every one of them just above it. The cause was a hook in the operating curve: flagging the last fifth of healthy people still catches diseased people whose scores sit far below the healthy mean, and at a high enough prevalence those catches are worth it.
Every curve in that essay was known exactly. A real threshold is chosen from a validation study of a few hundred people, whose empirical operating curve is a staircase. The question the essay left was how often such a study would recommend the wrong one of the two regimes. The answer turns on where the hook lives, and it lives where a validation study has almost no data.
What a validation study’s curve is made of
A validation study scores a set of people known to have the condition and a set known not to. Its empirical operating curve is traced by sliding a threshold down through the pooled scores: each healthy score passed moves the curve right by one over the number of healthy people, each diseased score passed moves it up by one over the number of diseased. The recommended threshold is the one whose empirical sensitivity and false-positive rate give the least expected harm at the programme’s prevalence.
Nothing in that procedure is wrong, and away from the extremes of the curve it works: the staircase follows the true curve within the noise of a few hundred people. The difficulty is at the top right, which is exactly where the jump happens.
The true curve is still rising there: at 40% of healthy people flagged it catches 97.60% of diseased people, and it reaches 100% only as the last healthy person is flagged, because diseased scores twice as spread as healthy ones have a long lower tail. The study’s curve reaches 100% the moment its lowest diseased score is passed — here with 67.3% of its healthy people flagged — and is flat from there on. A hundred diseased scores rarely include one from the far lower tail, so the study does not know the tail exists.
That flat top is where the jump was decided. The best threshold jumps to “everyone” because flagging the last third of healthy people still buys sensitivity worth a hundred false alarms per diseased person caught; the study’s curve says it buys nothing. So the study recommends stopping at its lowest diseased score, or a little above it, and never sees the reason to go further.
What the studies recommend
Just past the jump, at a prevalence of 26%, the best threshold flags everyone. Of a thousand validation studies with a hundred diseased and three hundred healthy people, 7.7% recommend a threshold that flags nearly everyone. The rest recommend cuts spread across the whole range: 34.9% of all the studies flag fewer than half of healthy people, which on the true curve misses a few per cent of diseased people at a hundred times the cost of each false alarm saved. Averaged over the studies, the recommended thresholds cause 15.4% more harm than the best threshold would.
In a programme’s own units the excess is concrete. Per thousand people tested at a prevalence of 26%, the recommended thresholds miss 4.00 more diseased people on average than the best threshold does, and spare 285.9 healthy people a false alarm. At a hundred false alarms to a miss that is a bad trade by a factor of almost one and a half, and it is the trade the flat top of the staircase makes invisible: the study saw no diseased person beyond its cut and so counted the misses as zero.
The problem is not confined to the prevalence where the jump happens.
At a prevalence of 15%, where the best threshold is an interior cut flagging 48.6% of healthy people and the hook is irrelevant, the recommended thresholds still cost 10.1% more than the best, and 2.1% of studies recommend flagging everyone. A staircase of a hundred diseased people is a noisy guide to any threshold that trades sensitivity above 97% against specificity, because every step of it is one person.
Why the staircase is reliable in the middle and not at the top
A step of the empirical curve is one person. In the middle of the curve, where a threshold trades an 80% sensitivity against a 90% specificity, a hundred diseased people estimate the sensitivity to within about four percentage points, and the harm at a cut is a smooth function of estimates that are good to that precision. At the top of the curve the relevant quantity is the share of diseased people not caught, which is a few per cent, and a hundred diseased people estimate a few per cent as zero, one, two or three people. The relative error is enormous exactly where the decision is most sensitive to it, because each missed person costs a hundred false alarms.
That is the same arithmetic an upper limit is the finding met from the other side: a rare event counted in a modest sample is better described by an upper limit than by an estimate, and “no diseased person beyond this cut” in a study of a hundred is compatible with a true miss rate of up to about 3% — the rule of three that what a reference distribution costs to sample used to bound an acceptance rate that no draw had reached. A threshold chosen by harm treats the zero as a zero.
It is also a cousin of the cut that fitted best, which found a baseline threshold chosen from a trial’s own data flattering the precision it reported. Here the flattery is in sensitivity: the recommended threshold is, by construction, the one at which the study’s own diseased people all happened to score above it, so its apparent sensitivity of 100% is the best-looking reading of a noisy staircase.
A model for the scores sees the tail
The repair is not a larger study of the same kind but a model of the same data. Fit a normal distribution to the healthy scores and another to the diseased scores — two means, two spreads — and choose the threshold that minimises the fitted model’s harm. The fitted diseased distribution has a lower tail whether or not the sample happened to include anyone from it, and its spread, estimated from all hundred diseased scores, says how long that tail is.
At a prevalence of 26%, 47.4% of studies using the fitted model recommend flagging everyone, and the recommended thresholds cause 2.5% more harm than the best on average — a sixth of the empirical curve’s excess. That 47% is not a failure to find the right answer: at the jump the two minima of the true harm are almost exactly level, so a method that sees both should split between them, and whichever it picks costs almost the same. At 15%, the fitted model’s excess is 3.6%, against 10.1% for the staircase.
The model is doing what the staircase cannot: extrapolating into the tail from the shape of the body. That is an assumption — that the diseased scores are normal far below where any were observed — and it is exactly the assumption that produced the jump in the first place. A test whose diseased scores had a lighter lower tail than a normal would have no hook, and the fitted model would invent one. Which of the two failures is more likely is a question about the test, and a validation study of a hundred diseased people can say little about it.
The people a study needs more of
A validation study is usually sized by its total, or by the precision of its sensitivity and specificity at a pre-specified cut. Neither is the right frame for choosing the cut.
The excess harm falls with the number of diseased people and barely moves with the number of healthy ones. From the study’s own curve with three hundred healthy people it is 81.0% at twenty-five diseased, 15.4% at a hundred and 2.2% at a thousand; with fifteen hundred healthy people and a hundred diseased it is 15.1% — the extra twelve hundred healthy people bought nothing. From the fitted model it is 9.3%, 2.5% and 0.5%.
The asymmetry follows from where the information is needed. The jump is decided by the lower tail of the diseased scores, and only diseased people carry information about it. Healthy people pin down the false-positive rate, which is already estimated well by a few hundred; adding more sharpens a quantity that was not in doubt. A programme validating a test for a threshold it will set by expected harm should spend its validation budget on confirmed cases, which are usually the scarce and expensive half.
Where the two published numbers come from
A test’s published sensitivity and specificity come from the same kind of study, at one threshold chosen before it began, and they are the numbers every predictive-value calculation starts from — what a positive test is worth turned a 90/95 pair into a predictive value of 1.8% at a prevalence of one in a thousand. The study that produced the pair can report it well: sensitivity and specificity at a fixed cut are proportions with ordinary binomial precision. What it cannot report well is the rest of the curve, and in particular the part where a programme at a high prevalence would want to operate.
The second test that is not a second opinion found that a second positive is worth far less than independence promises when the two tests’ errors are correlated. A threshold moved into the flat top of a validation study’s curve has the same kind of hidden dependence: the people it additionally catches are the ones whose scores sit in the far lower tail, and whether they exist in the population is exactly what the study could not observe.
Why the prevalence makes it worse
The jump sits at a prevalence of about a quarter, and a quarter is not an exotic prevalence for the settings where a threshold is chosen by harm: a referral clinic, a symptomatic population, a second-stage test after a positive screen. Those are also the settings where the prevalence itself is known only roughly — an upper limit is the finding found that even a large survey often bounds a rare prevalence rather than estimating it — so a programme choosing its threshold faces two uncertainties at once, and they interact. Near the jump a small change in the assumed prevalence moves the best threshold from one regime to the other; the validation study’s staircase then adds its own noise on top.
A report that gives one recommended threshold hides both. A report that gives the two candidate regimes — the interior cut and the flag-everyone rule — with the prevalence at which they trade, and says how well its data pin down the tail that decides between them, lets a programme make the choice with its own prevalence in hand. That is a longer report and a more honest one, and it is the natural companion of a test that is a point somebody chose: the published sensitivity and specificity describe one point on a curve, and the point that should be used depends on things the validation study does not know.
The prevalence the programme has to supply
The recommended threshold depends on the prevalence at which harm is computed, and the validation study does not supply it: its mix of diseased and healthy people is set by recruitment, not by the population the test will be used in. The programme supplies it, usually from the same test’s positive rate in the field, which the prevalence the test has to estimate found reads fifty times the truth at a prevalence of one in a thousand before it is corrected for the test’s own errors.
Near the jump that dependence compounds the staircase’s. A programme whose corrected prevalence estimate is 22% sits on the interior side, one at 30% on the everyone side, and the validation study’s recommendation is noisy on both. The fitted model at least makes the dependence explicit: it gives a harm curve for every prevalence from one fit, and the programme can read off where the jump sits for its own test and how far its own prevalence estimate is from it. The empirical curve gives one staircase, and a different recommended cut for every prevalence, with no sign of the jump between them because the part of the curve that makes the jump is flat.
What the validation of a threshold should report
How many diseased people the curve’s top rests on. The part of the operating curve where more than 95% of diseased people are caught is traced by the last few diseased scores. A study that recommends a threshold there should say how many of its diseased people lie beyond it, because that number, not the study’s total, is its sample size for the decision.
The fitted curve beside the empirical one. A binormal fit extrapolates, and says so; an empirical curve refuses to extrapolate, and says nothing about what it cannot see. Reporting both makes the dependence on the tail visible.
The harm of the two regimes at the programme’s prevalence. Where a hook exists, the decision is between two local minima, not the location of one, and the confidence interval for “the optimal threshold” that a smooth-optimum analysis would report does not describe a choice between two separated answers.
What is measured here and what is not
Just past the jump prevalence, 7.7% of validation studies of a hundred diseased and three hundred healthy recommend flagging everyone, and their recommended thresholds cause 15.4% more harm than the best; fitted binormal models recommend it in 47.4% of studies at an excess of 2.5%.
The excess harm is set by the number of diseased people validated, not the number of healthy ones.
Every rate is counted over a thousand simulated validation studies a point — five hundred at five hundred diseased or more — scored by the test of the essay on the jump: healthy scores standard normal, diseased scores normal with twice the spread, the pair 90/95 at the published cut, and a miss costing a hundred false alarms.
Not measured: tests whose diseased scores are not normal, where the fitted model’s extrapolation is itself the error; smoothed empirical curves, which sit between the two methods here; and the harm of a threshold chosen at an estimated rather than known prevalence, which compounds both uncertainties.
Still open: a test whose tail is not normal
The fitted model won here because the truth was binormal. The case that matters in practice is the one where it is not: a diseased population that is a mixture of mild and severe cases, whose scores are bimodal, or a test that saturates, whose scores pile up at a ceiling. In both the binormal fit extrapolates a tail that is not there, or misses one that is, and the empirical curve’s refusal to extrapolate becomes a virtue.
How much the fitted model’s advantage survives a plausible departure from binormality, whether a more flexible model of the diseased scores keeps the advantage without inventing hooks, and how a study can tell from a hundred diseased scores which kind of test it has, are the measurements that would decide between the two methods in general. None of them has been made here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The liar with two answers — both name base rate, roc curve, sensitivity and specificity
- A score that rewards lying — both name base rate, roc curve
- Calibrated and useless — both name base rate, roc curve
- The level a limit should be set at — both name expected loss, sample size
Named objects
A flat tag is an object no other essay names yet.
Base rateDecision thresholdExpected lossExtrapolationPrevalenceROC curveSample sizeSensitivity and specificity