Two tests, a threshold, and the rate they are read against

A threshold reported with its interval

Resample a validation study of a hundred diseased and three hundred healthy people, refit the binormal model each time, and read off the threshold each resample recommends. Far from the prevalence at which the best threshold jumps to flagging everyone, the interval is tight — flagging 12% to 24% of healthy people at a prevalence of 5%. Near the jump, the median study's interval runs from flagging under half of the healthy to flagging everyone, and 91% of intervals cross the jump. The two ends differ by fifteen false alarms' worth of harm per hundred people, and no rule for choosing inside the interval beats the study's own choice at every prevalence.

Worth reading first: What a positive test is worth.

A threshold that jumps found that when a screening test’s diseased scores are twice as spread as its healthy scores, the threshold that minimises harm does not move smoothly with prevalence. Below about a quarter it is an interior cut; at a prevalence of 0.257 it jumps to flagging everyone, because the true operating curve keeps climbing in a tail where few diseased people score. The threshold a validation study can see showed that a study’s own staircase curve cannot see that tail and that a binormal fit can, and a test whose tail is not normal showed what the binormal fit does when the tail it imagines is not there.

Every one of those essays had each study recommend one threshold. The last ended on the report a study should make instead: when the models disagree about the tail, the honest output is not a threshold but an interval of thresholds consistent with the study’s data, and a statement of how much harm each end of it would cause if the other were true. That interval can be built and counted, and it is wide exactly where the threshold matters most.

The interval, built

The test is the earlier essays’ unequal-spread test: healthy scores normal with unit spread, diseased scores normal with spread two, placed so that the published pair — 90% sensitive, 95% specific — holds at the published cut. A miss costs a hundred false alarms. A validation study scores a hundred diseased and three hundred healthy people, fits a normal curve to each group, and recommends the threshold its fitted model says minimises harm at the programme’s prevalence — an interior cut, or “everyone”.

The interval comes from resampling. The study’s diseased scores are resampled with replacement, its healthy scores likewise, the binormal model is refitted, and the recommended threshold is read off — two hundred times per study. The thresholds those resamples recommend, expressed as the share of healthy people each would flag, give a 95% interval from their 2.5th to their 97.5th percentile. A recommendation to flag everyone counts as flagging all of the healthy, so an interval can run up to one.

The thresholds one validation study supports, read from four hundred resamples of it. Prevalence 0.26, just past the jump to flagging everyone. The study's own binormal fit recommends a threshold flagging 65.8% of healthy people; its resamples' recommendations run from 36.6% to 100.0% at the 2.5th and 97.5th percentiles, and 18.5% of them flag everyone. The true best threshold flags everyone.
Fig. 1 One validation study at a prevalence of 0.26, just past the jump: the share of healthy people flagged by the threshold each of four hundred bootstrap resamples recommends, with the study’s own recommendation and the 95% interval beneath it. The red bar is the resamples that flag everyone.

The study drawn above is typical of the region near the jump. Its own fit recommends a threshold that flags 65.8% of healthy people. Its resamples recommend thresholds from 36.6% to 100% at the interval’s ends, and 18.5% of them recommend flagging everyone. The true best threshold at this prevalence flags everyone. The study’s single recommendation is wrong; its interval contains the right answer and also a threshold that flags barely a third of the healthy.

Tight far from the jump, everything near it

The thresholds validation studies support, against the prevalence they are chosen for. Four hundred studies a prevalence, two hundred resamples each. Median interval and true best threshold, as false-positive rates: 0.05: 12.3% to 24.4%, truth 18.5%, holding it in 94.0% of studies, crossing the jump in 0.0%; 0.1: 21.6% to 47.6%, truth 33.9%, holding it in 92.3% of studies, crossing the jump in 15.3%; 0.2: 36.7% to 100.0%, truth 63.2%, holding it in 91.5% of studies, crossing the jump in 80.0%; 0.26: 45.3% to 100.0%, truth 100.0%, holding it in 92.5% of studies, crossing the jump in 91.0%; 0.35: 57.0% to 100.0%, truth 100.0%, holding it in 98.3% of studies, crossing the jump in 91.3%.
Fig. 2 Over four hundred studies at each prevalence: the median 95% interval of thresholds the study supports, as shares of healthy people flagged, and the true best threshold (red). Above the dashed line a threshold flags nearly everyone.

At a prevalence of 5% the true best threshold flags 18.5% of healthy people, and the median study’s interval runs from 12.3% to 24.4% — a band twelve points wide that holds the truth in 94.0% of studies and never comes near flagging everyone. At 10% the truth flags 33.9%; the median interval runs from 21.6% to 47.6%, holds the truth in 92.3% of studies, and crosses the jump in 15.3% of them.

At 20% the truth is still an interior cut, flagging 63.2%, and the median interval already runs from 36.7% to 100% — crossing the jump in 80.0% of studies. At 0.26, just past the jump, the truth flags everyone and the median interval runs from 45.3% to 100%, crossing in 91.0% of studies. At 0.35, well past it, the median interval runs from 57.0% to 100%, crossing in 91.3%.

So a validation study of this size can locate the threshold when the prevalence is far from the jump and cannot when it is near. Near the jump the interval is not a refinement of the recommendation; it is the whole range from a moderate interior cut to the everyone rule, and the study’s data do not distinguish between them.

Coverage, and what it means here

The intervals hold the true best threshold in 91.5% to 98.3% of studies at the five prevalences — near but mostly short of their 95%, as resampling intervals for a minimiser often are, since the recommended threshold is a jumpy function of the data and its bootstrap distribution is lumpy. That is not where the interval’s value lies. Its value is in showing, for a given study, whether the recommendation can be trusted to the nearest few points or whether it is a coin between two different policies.

The crossing rate is the number that answers that. Below the jump by a factor of two and more, intervals do not cross: the study knows the threshold is interior and roughly where. From a prevalence of 0.2 upwards, four studies in five or more produce an interval that holds both an interior cut and the everyone rule. That is a finding about the test, not about the study: a test with this unequal spread, validated on a hundred diseased people, cannot be given a threshold near the jump. The base rate that decides the policy is a number the programme knows; the tail of the diseased scores that decides it as well is a number the study only half sees.

What each end costs

What acting on each end of a study's interval of thresholds costs, against the best threshold, by prevalence. Average excess harm per hundred people, misses costing a hundred false alarms. 0.05: lower end 1.91, upper end 1.31, the study's own choice 0.32; 0.1: lower end 4.52, upper end 7.29, the study's own choice 0.94; 0.2: lower end 10.77, upper end 5.34, the study's own choice 2.76; 0.26: lower end 15.29, upper end 0.11, the study's own choice 1.97; 0.35: lower end 24.38, upper end 0.13, the study's own choice 3.40.
Fig. 3 The true excess harm, per hundred people screened, of acting on the lower end of a study’s interval, on its upper end, and on the study’s own recommendation, against the harm of the best threshold, averaged over four hundred studies at each prevalence.

The earlier essay asked how much harm each end of the interval would cause if the other end were right. Measured against the true best threshold, the answer at a prevalence of 5% is small at both ends: acting on the lower end costs 1.91 false alarms’ worth of harm per hundred people above the best, the upper end 1.31, and the study’s own choice 0.32. The band is tight and either end is close.

Near the jump the ends separate. At 0.26 the lower end costs 15.29 per hundred above the best and the upper end — which is the everyone rule in most studies, and the right one here — 0.11. At 0.35 the lower end costs 24.38 and the upper 0.13. At 0.2, where the truth is still interior, the roles reverse in part: the lower end costs 10.77 and the upper end 5.34, because flagging everyone overshoots a threshold that should flag 63%.

That is the statement a programme deciding a screening policy needs from a validation study near the jump. The study cannot say which end is right; it can say that one end is expensive if wrong by roughly ten to twenty-five false alarms per hundred people screened, and the other by a fraction of one to five. A programme that cannot tolerate the larger of those has a reason to flag more people than the study’s own recommendation does, and the interval is what shows it.

Why crossing the jump is not itself the failure

There is a subtlety in the crossing rate that the cost of each end resolves. At a prevalence of 0.26 the best interior threshold flags 82.3% of the healthy and costs 74.22 false alarms’ worth of harm per hundred people; the everyone rule costs 74.00. The two policies are 0.22 apart. That is what a jump is: at the prevalence where the best threshold jumps, the interior minimum and the everyone rule are exactly level, and just past it they are nearly level. An interval that holds both is reporting, correctly, that near the jump it barely matters which.

What matters is the rest of the interval. Its lower end, flagging 45% of the healthy in the median study, is not the best interior cut; it is a much worse one, and acting on it is what costs fifteen per hundred. The study’s uncertainty near the jump is not mostly about which side of the jump to be on. It is about how far down the interior side a resample’s fitted tail can pull the threshold, and that is where more data helps.

More diseased people

With the same three hundred healthy people and more diseased ones — two hundred studies a size, a hundred resamples each, at a prevalence of 0.26 — the median interval’s lower end rises from 43.2% of the healthy with a hundred diseased to 56.6% with three hundred and 65.9% with a thousand. The excess harm of acting on the lower end falls from 16.15 per hundred to 5.71 and then 2.23, and the study’s own choice from 2.05 to 0.84 and 0.38.

The crossing rate does not fall at all: 91.0%, 91.0% and 90.5%. A thousand diseased people still cannot say which side of the jump a prevalence of 0.26 is on, because the two sides are 0.22 apart in harm and no sample of a feasible size resolves a difference that small. What a larger study buys is an interval whose interior end is a good interior cut, so that crossing the jump stops costing anything.

That separates two questions a validation study is asked to answer. Which policy — interior cut or everyone — is unanswerable near the jump and also unimportant there. How costly the worst defensible policy would be is answerable, falls quickly with the number of diseased people validated, and is the number a programme needs. An upper limit is the finding made the same move for a prevalence survey whose interval included zero: when the point is uninformative, the end of the interval that bounds the harm is the result.

Choosing inside the interval

Given the interval, two rules suggest themselves besides the study’s own choice. One averages the harm over the resamples’ fitted models and picks the threshold that minimises the average — a bagged recommendation. The other picks the threshold whose worst excess over any resample’s own best is smallest — a minimax-regret recommendation, which hedges against the resample whose model is most unlike the threshold chosen.

Three rules for choosing inside a study's interval of thresholds, by the excess harm they cause. 400 studies a prevalence. the study's own choice: 0.32, 0.94, 2.76, 1.97, 3.40; harm averaged over resamples: 0.33, 0.94, 3.00, 1.96, 3.01; worst excess over resamples: 0.39, 2.20, 2.66, 1.30, 3.53 at prevalences 0.05, 0.1, 0.2, 0.26, 0.35.
Fig. 4 The true excess harm per hundred people of three rules for choosing inside a study’s interval — the study’s own choice, the threshold minimising harm averaged over the resamples, and the threshold minimising the worst excess over any resample — at each prevalence, over four hundred studies.

The averaged rule is almost exactly the study’s own choice: 0.33 against 0.32 at 5%, 0.94 against 0.94 at 10%, 3.00 against 2.76 at 0.2, 1.96 against 1.97 at 0.26 and 3.01 against 3.40 at 0.35. Averaging the resamples’ harm curves smooths them, and the smoothed curve’s minimum sits where the original fit’s did.

The minimax rule is different and its record is mixed. Just past the jump it does best of the three — 1.30 against the study’s 1.97 — though it recommends the everyone rule outright in fewer studies, 26.5% against 43.8%: its other recommendations flag 78.6% of the healthy on average, where the plug-in’s interior cuts flag 65.1%, which is a hedge towards the rule that is right there. At 10% it does worst, 2.20 against 0.94, because hedging against the occasional resample that wants a much lower or higher threshold moves a threshold that was already right. At 0.35, 3.53 against 3.40.

No rule inside the interval beats the study’s own choice at every prevalence. The interval does not supply a better point; it supplies the range, the cost of each end, and therefore the basis for a decision the study alone cannot make.

Why the study cannot do better

The jump is a property of a tail the study barely samples. With diseased scores spread twice as widely as healthy ones, the true operating curve keeps rising among the few diseased people who score low, and whether flagging everyone is worth it depends on how many of them there are below the healthy range. A hundred diseased people include, on average, fewer than two who score below the healthy people’s average. The binormal fit extrapolates their spread into the tail, and the extrapolation’s uncertainty is the interval’s width: resampling which handful of low scorers the study happened to include moves the fitted spread, and near the jump the fitted spread decides which side the threshold falls.

A test whose tail is not normal showed the other half of the same fact: the fit is only as good as the normal shape assumed in the tail, and a skewed test with no lower tail makes the binormal fit recommend everyone almost always. The interval built here assumes the shape is right. Under a wrong shape the interval inherits the fit’s bias and is narrow around the wrong answer, which is a reason to report the normality check of the earlier essay beside it.

What a validation study should report

The published pair is not the test. The test is a point somebody chose found that a sensitivity and specificity describe one point on a curve; this essay’s tests all share the published pair and differ in where the best threshold sits, by up to the whole range of the interval. What a positive test is worth priced a positive result from the pair alone; a threshold priced from the pair alone inherits none of the uncertainty measured here.

The interval, in the units the policy is decided in. The share of healthy people a threshold flags is what a programme budgets for, and an interval from 45% to everyone is a different finding from a threshold of 66%, even though both come from the same study.

Whether the interval crosses the jump. That is the one-line verdict on whether the study has located the threshold. At this test’s spreads it has far below the jump and has not near it, and four studies in five near the jump cannot tell an interior cut from the everyone rule.

The cost of each end if the other is right. Fifteen false alarms per hundred people screened for acting on the lower end when the truth is everyone, a tenth of one for the reverse, at a prevalence of 0.26. A policy can be chosen against those numbers; it cannot be chosen against a point recommendation that hides them.

More diseased people, not more healthy ones. The threshold a validation study can see found that the excess harm falls with the diseased people validated and not with the healthy; the interval’s width is set by the same tail, and narrowing it near the jump means sampling more of the people who score low despite being diseased.

Counted here

Every rate is over four hundred simulated validation studies at each prevalence, a hundred diseased and three hundred healthy people each, two hundred bootstrap resamples per study; the single study drawn first has four hundred. The best threshold and every true harm are computed from the test’s exact distributions. Harm is measured per hundred people screened, a miss costing a hundred false alarms. The comparison across numbers of diseased people uses two hundred studies of each size with a hundred resamples apiece, so its row for a hundred diseased people differs by a point or two from the four-hundred-study figures above; the direction and size of every change it reports are well outside that difference.

Not claimed: that resampling is the only way to build the interval. A parametric interval from the fitted model’s own uncertainty — the four fitted means and spreads, propagated to the threshold — would be cheaper and might cross the jump less erratically; it was not built. Nor that the binormal shape is right for a real test; the interval is honest about sampling error under the shape assumed and silent about the shape.

Still open: a threshold for a programme that will learn

A screening programme does not run once. Each round produces new diseased cases — some found by the screen, some not — and those cases, especially the ones the screen missed, are exactly the low-scoring diseased people the validation study had too few of. A programme that starts at one end of the interval and updates its threshold as missed cases accumulate is a sequential design.

Which end it should start at, how many rounds of missed cases it takes for the interval to stop crossing the jump, and how much excess harm the learning costs against a programme that knew the tail from the start, are measurable with the same test and the same costs, and have not been measured here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Base rateBootstrapDecision thresholdExpected lossExtrapolationPrevalenceRegretROC curveSensitivity and specificity