Two tests, a threshold, and the rate they are read against

A threshold that jumps

When a test's scores are equally spread in the healthy and the diseased, the harm-minimising threshold slides smoothly as the prevalence changes. Make the diseased scores twice as spread — the ordinary shape of a group that mixes mild and severe cases — and keep the published 90/95 pair, and the best threshold flags 81.1% of healthy people at a prevalence just under 25.7% and every healthy person just above it. At three times the spread it jumps from 45.7% to everyone at 13.0%. A threshold that jumps cannot be given as a formula; it has to be given with the second-best point beside it.

Worth reading first: What a positive test is worth.

The test is a point somebody chose found that a test reported as 90% sensitive and 95% specific is one point on a curve, chosen by a threshold, and that the threshold minimising harm moves with the prevalence — from 3.05 standard deviations at one in ten thousand to −0.12 at one in two. That calculation assumed the scores of healthy and diseased people are normal with the same spread, and it ended by saying what that assumption hides. The diseased group is usually more variable, being a mixture of mild and severe cases, and with unequal spreads the harm-minimising threshold might not move continuously at all. A threshold that jumps is one that cannot be recommended as a formula.

It does jump. This essay keeps everything else from that calculation — binormal scores, the published 90/95 pair at its published cut, a miss costing a hundred false alarms — changes only the spread of the diseased scores, and measures where and why the best threshold stops sliding.

The share of healthy people the harm-minimising threshold flags, against the prevalence, for three spreads of the diseased scores. A miss costs a hundred false alarms. With diseased scores as spread as healthy ones the best threshold flags more healthy people smoothly as the prevalence rises. With them twice as spread it flags 81.1% just below a prevalence of 25.7% and everyone just above it; three times as spread, 45.7% and then everyone at 13.0%.
Fig. 1 The share of healthy people flagged by the harm-minimising threshold, against the prevalence on a logarithmic scale, for diseased scores as spread as the healthy ones and twice and three times as spread. Every test is 90% sensitive and 95% specific at its published cut. The dashed lines mark where the best threshold jumps.

Three tests with one published pair

Healthy scores are normal with mean 0 and spread 1. Diseased scores are normal with spread ss and a mean dd set so that the cut at the healthy 95th percentile, 1.645, catches exactly 90% of them: d=1.645+1.282 sd = 1.645 + 1.282\,s. At s=1s = 1 that is the test of the essay on the threshold, with d=2.93d = 2.93. At s=2s = 2 the diseased mean is 4.21 and at s=3s = 3 it is 5.49. All three tests would be reported the same way — 90% sensitive, 95% specific — and all three are, at that one cut.

A threshold is judged by the expected harm per person tested: the prevalence times the miss rate times a hundred, plus the rest of the population times the false-alarm rate. Lowering the threshold trades misses for false alarms, and the best cut is where the trade stops paying.

Where the best cut stops sliding

At equal spreads the grey curve in the hero figure rises smoothly: as the prevalence goes up, misses become more expensive in total, and the best threshold comes down to catch more cases at the cost of more false alarms. It flags about half the healthy population at a prevalence of one in two, and at no prevalence does it change by more than a little for a little change in prevalence.

With the diseased scores twice as spread, the best threshold follows much the same path until the prevalence reaches 25.7%. Just below that it sits at a cut of −0.88 and flags 81.1% of healthy people while catching 99.45% of the diseased. Just above it, the best single threshold is no threshold: flag everyone. At three times the spread the same happens earlier and further: at a prevalence of 13.0% the best cut, which flags 45.7% of healthy people, gives way to flagging everyone.

diseased spread jump at a prevalence of healthy flagged just before diseased caught just before
1.5 × 70.35% 99.46% 100.00%
2 × 25.69% 81.11% 99.45%
2.5 × 16.24% 59.87% 97.93%
3 × 12.98% 45.70% 96.36%

At one and a half times the spread the jump is from almost everyone to everyone, which is a jump in name only. From twice the spread onwards it is a real discontinuity: a policy that flagged 46% of healthy people becomes, with a change in prevalence of a fraction of a point, a policy that flags them all. Nothing about the test changed and nothing about the costs changed. The prevalence crossed a line.

Two minima trading places

The mechanism is visible in the harm itself, plotted against the threshold.

Expected harm against the threshold for diseased scores 2 times as spread, at three prevalences around the jumpAt a prevalence of 25.69% the harm curve's interior minimum, at a cut of −0.88, costs exactly what flagging everyone costs, which is the curve's value at the far left. A little below that prevalence the interior cut is better; a little above, flagging everyone is.0.8001-6-4-2024threshold on the score, healthy standard deviationsexpected harm per personprevalence 21.8%prevalence 25.7%prevalence 29.5%flag everyoneexact binormal harm, a miss worth a hundred alarmstwo minima trade places
Fig. 2 Expected harm per person against the threshold, for diseased scores twice as spread as healthy ones, at prevalences 15% below, at, and 15% above the jump. The flat stretch on the left is the harm of flagging almost everyone; the dip near zero is the interior threshold.

At equal spreads the harm curve has one minimum. At twice the spread it has two: a dip in the interior, near a cut of −0.9, and the far left, where the threshold is so low that everyone is flagged and the harm is simply the false-alarm cost of the whole healthy population. Between them the curve rises, which is the part that matters. Lowering the threshold from the interior dip at first increases the harm, because the cases it catches are few and the healthy people it flags are many; only when it has gone far enough to flag nearly everyone does it come back down.

As the prevalence rises both minima fall, at different rates. The prevalence of the jump is the one at which they are exactly level, and on either side of it the best threshold is one or the other with nothing in between. A search for the best threshold that starts from the published cut and moves downhill finds the interior dip, and above the jump that is the wrong answer — the second-best point of the curve, reported as the best.

Three times the spread

Expected harm against the threshold for diseased scores 3 times as spread, at three prevalences around the jump. At a prevalence of 12.98% the harm curve's interior minimum, at a cut of 0.11, costs exactly what flagging everyone costs, which is the curve's value at the far left. A little below that prevalence the interior cut is better; a little above, flagging everyone is.
Fig. 3 The same harm curves for diseased scores three times as spread as healthy ones, at prevalences around 13%. The interior dip is shallower and sits further to the right, and the flag-everyone stretch on the left is reached through a longer rise.

At three times the spread the interior minimum is shallower and the rise between it and the flag-everyone edge is longer. That is why the jump comes at a lower prevalence and spans more: the interior policy at the jump flags fewer than half the healthy population, so crossing to “flag everyone” means adding more than half of it at once. A programme that sits on the interior side of that line and one that sits on the other are running policies that differ in the treatment of more than half of all healthy people tested, and they would be recommended by the same formula on either side of a prevalence difference of a few tenths of a point.

The published pair gives no warning of any of it. The three tests are indistinguishable at their published cut, and the first sign that the spreads differ is how the operating curve behaves far from that cut, where most validation studies report little. The ratio the jump turns on is the likelihood ratio, which the odds form of the screening arithmetic uses at one point — the published one — to turn a prior into a posterior. The threshold question uses it at every point, and that is exactly where two tests with one published pair stop being the same test.

The prevalence the jump depends on has its own difficulty. Estimated from the test’s positive rate, it is uncertain by more than the distance between a prevalence of 12% and one of 14% in any survey of ordinary size, so a programme cannot know which side of a jump it is on from its own data. And a second test used to confirm positives changes the effective prevalence among those it is applied to, which moves the jump for the second test by an amount that depends on how correlated the two tests’ errors are. A threshold that jumps is sensitive to every input the threshold arithmetic has, and to the one it most often lacks.

A curve with a hook

The two minima come from a property of the operating curve that equal spreads rule out. The slope of the curve at any threshold is the likelihood ratio of a score at that threshold — how much more common that exact score is among the diseased than among the healthy — and the harm-minimising threshold is where that slope equals the ratio of costs weighted by the prevalence.

The slope of the operating curve — the likelihood ratio at the threshold — against the share of healthy people flagged, for three spreads. At equal spreads the likelihood ratio at the threshold falls all the way as the threshold is lowered. With the diseased scores twice as spread it bottoms out where 92.0% of healthy people are flagged and rises beyond; three times as spread, at 75.4%. A cut on the rising stretch flags more healthy people for a better ratio, which no harm-minimising single threshold would do.
Fig. 4 The likelihood ratio of a score at the threshold, on a logarithmic scale, against the share of healthy people flagged, for three spreads of the diseased scores. Rings mark where the ratio stops falling and starts to rise.

At equal spreads the likelihood ratio falls all the way down as the threshold is lowered: every lower score is less typical of disease than the one above it, the operating curve bends one way throughout, and for any cost ratio there is exactly one point where its slope matches. At unequal spreads the ratio falls, bottoms out and rises again. The logarithm of the ratio is a parabola in the score opening upwards, with its lowest point at −d/(s2−1)-d/(s^2 - 1), and below that score a lower reading is more typical of disease than a middling one, because the wider diseased distribution has more mass far out on both sides.

At twice the spread the ratio turns where 92.0% of healthy people are flagged; at three times, where 75.4% are. Beyond the turn the operating curve bends the wrong way — it has a hook near its top — and no single threshold on the hook is ever optimal, because any cost ratio that a point on the hook matches is matched better by a line from the hook’s start to the corner where everyone is flagged. The jump in the hero figure is the optimum crossing the hook in one step.

The stretch that is worse than a coin

The essay on the threshold warned that unequal spreads make the operating curve cross the diagonal, with thresholds at which the test is worse than a coin. That is true and it is worth knowing how little it matters. The curve is below the diagonal where the diseased are less likely than the healthy to score above the cut, which for these tests happens only at cuts below −d/(s−1)-d/(s - 1).

At twice the spread that is a cut of −4.21, where 99.9987% of healthy people are flagged. At three times it is −2.74, flagging 99.70%. The coin-worse stretch exists, and it lies entirely inside the hook, in a region no harm-minimising threshold visits. The part of the unequal-spread curve that changes practice is not the part below the diagonal; it is the much larger hook above it, which starts at 75% or 92% of healthy people flagged rather than at 99.7%.

What a second threshold buys

The rising likelihood ratio has a direct reading: very low scores are evidence for disease. The rule that uses the likelihood ratio rather than a single cut flags scores above one threshold and below another — the two roots of the parabola — and it removes the hook, since it never has to pass through the middling scores to reach the low ones.

spread prevalence best single cut two-cut rule, lower cut harm saved
2 × 5% 0.898 −3.703 0.02%
2 × 20% −0.338 −2.468 0.41%
3 × 5% 1.038 −2.410 2.79%
3 × 10% 0.450 −1.823 5.70%

At twice the spread the second threshold is almost never worth drawing: it flags scores more than 3.7 healthy standard deviations below the mean at a prevalence of 5%, which catches 0.004% of the diseased and saves 0.02% of the harm. At three times the spread and a prevalence of one in ten, it catches 0.74% of the diseased at the cost of flagging 3.42% of the healthy, and saves 5.70% of the harm. The two-cut rule is the right rule in principle and a small improvement in practice, except for a very variable diseased group at moderate prevalence — which is also where the single threshold is closest to its jump.

It has a cost the harm calculation does not show. A rule that calls a low score positive is one no clinician would accept without an explanation, and the explanation — the diseased group is so heterogeneous that a very low reading is more typical of it than a middling one — is a claim about the score distributions that the model imposes and the data must support. A binormal model is fitted mostly to the middle of each distribution, and the far lower tail of the diseased scores is the part of it least constrained by data — a validation study of a few hundred cases holds a handful of scores out there, which is a tail the sample never saw in another setting.

What the practical advice becomes

The essay on the threshold ended with the advice to compute the harm-minimising cut for the prevalence in hand. With equal spreads that is enough. With unequal spreads it needs two additions.

Compute the second-best point. Where the harm curve has two minima, report both and the prevalence at which they trade places. A screening programme running at a prevalence near its jump is one in which a small shift in who is tested — a referral pathway that sends more high-risk patients, a seasonal rise in incidence — changes the right policy from flagging half the healthy population to flagging everyone, and a recommendation that gave only the current best cut would say nothing about how close that is.

Check the spreads before trusting the formula. The published sensitivity and specificity cannot distinguish the three tests here; all three are 90/95. What distinguishes them is the spread of the diseased scores, which is visible in a validation study’s raw data and almost never in its abstract. The predictive value arithmetic needs only the pair and the prevalence. The threshold arithmetic needs the whole curve, and the curve needs both spreads.

Keep the score. Every threshold turns a measured score into a yes or a no, and cutting a measurement in two throws information away in a way that can be priced. Near a jump the loss is sharper than usual, because the right cut is itself uncertain; a report that carries the score, and lets the threshold be recomputed as the prevalence and the costs are updated, survives a jump that a report of positives and negatives cannot.

There is also a reading for the high-prevalence end, where the jump lands. “Flag everyone” means that, at that prevalence and those costs, the test is not worth using — its best contribution is to be ignored and everyone treated. That conclusion is sometimes right, and a single-threshold formula that returns an interior cut above the jump will hide it by recommending a test that does more harm than not testing.

Where the jump sits for each spread, and what the binormal model assumes

With binormal scores, a miss costing a hundred false alarms, and every test 90/95 at its published cut, the harm-minimising single threshold moves continuously with the prevalence at equal spreads and jumps at unequal ones: at twice the spread from flagging 81.1% of healthy people to flagging all of them at a prevalence of 25.69%, at three times from 45.70% to all at 12.98%.

The cause is a hook in the operating curve, where the likelihood ratio at the threshold turns and rises — at 92.0% of healthy people flagged for twice the spread and 75.4% for three times — and the stretch below the diagonal lies wholly inside it, beyond 99.70% flagged.

A second threshold for very low scores removes the hook and saves at most 5.70% of the harm in the cases computed, and 0.41% or less at twice the spread.

Every quantity is exact under the binormal model: harms from normal tail areas, the jump by bisection on the difference between the interior minimum and the flag-everyone harm, and the two-cut rule from the roots of the log likelihood ratio’s parabola.

Not claimed: that real score distributions are binormal, or that their diseased spreads are two or three times their healthy ones. The spreads were chosen to show the phenomenon’s range; the prevalence of the jump depends on the cost ratio as well as the spread, and a different ratio of harms moves it. A diseased group that is a mixture of two severities, rather than one wide normal, would produce a curve with a shoulder rather than a hook, and whether its threshold jumps in the same way has not been computed.

Still open: the threshold chosen from a validation study

Every curve here is known exactly. A real threshold is chosen from a validation study of a few hundred people, whose empirical operating curve is a staircase with the true curve’s hook blurred into it. Near a jump, a small change in the estimated curve moves the estimated best threshold from one minimum to the other, so the chosen threshold is itself unstable in a way a confidence interval for a smooth optimum would not capture.

How often a validation study of a given size would recommend the wrong one of the two minima, and whether reporting both — the interior cut and the flag-everyone alternative, with the prevalence at which they trade — protects a programme better than reporting either, is a calculation over resampled validation studies that the upper limit on a prevalence suggests is worth doing, since the prevalence that decides between them is itself usually known only as a bound.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Base rateDecision thresholdExpected lossLikelihood ratioPrevalenceROC curveScreeningSensitivity and specificity