A test whose tail is not normal
Worth reading first: An identity in three terms.
The threshold a validation study can see chose a diagnostic threshold from a validation study of a hundred diseased and three hundred healthy people, for a test whose diseased scores are twice as spread as its healthy ones. Just past the prevalence at which the best threshold jumps to flagging everyone, the study’s own staircase curve could not see the tail that made the jump: 7.7% of studies recommended flagging everyone, and their thresholds caused 15.4% more harm than the best. Fitting a normal curve to each group saw the tail — 47.4% recommended everyone — and cut the excess to 2.5%.
The essay ended on why the fitted model won: the truth was binormal. The case that matters is the one where it is not — a diseased group that is a mixture of mild and severe cases, or a test whose diseased scores simply do not reach below some level — and in those the binormal fit extrapolates a tail that is not there or misses one that is. How much of its advantage survives, whether a more flexible model keeps it, and whether a hundred diseased scores can tell a study which kind of test it has, were left open.
Three tests with one published pair
The comparison needs tests that a published sensitivity and specificity cannot tell apart, so all three here are 90% sensitive at the cut where healthy scores, standard normal, are 95% specific.
The binormal test is the earlier essay’s: diseased scores normal with twice the healthy spread, a long lower tail that reaches down through the healthy scores. The mixture test has nine in ten diseased people scoring like severe cases, normal around 3.54, and one in ten mild cases scoring around 1, well inside the upper healthy range: its lower tail is a lump rather than a long slope. The skewed test has the binormal’s mean and spread but an exponential shape, and no diseased score falls below 1.43 at all.
The best single threshold, with a miss costing a hundred false alarms, differs sharply between them at a prevalence of 26%. For the binormal test it flags everyone. For the mixture it flags 77.6% of healthy people, reaching down far enough to catch most of the mild lump. For the skewed test it flags 7.6% — a cut just below the lowest possible diseased score catches every diseased person and flags almost no one else, at a harm of 5.61 false alarms’ worth per hundred people.
A tail that is not there
The skewed test is where the binormal fit fails worst. It fits a normal curve to a hundred diseased scores whose mean and spread are the binormal test’s, and so it believes in the binormal test’s long lower tail — the one that made the earlier essay’s jump. It then recommends what the binormal truth would call for: 89.7% of studies recommend flagging everyone. On the true test those recommendations add 67.08 false alarms’ worth of harm per hundred people to the best threshold’s 5.61, about twelve times the harm the programme needed to cause.
The study’s own curve does far better, because it believes only what it saw. With no diseased score below 1.43, its staircase reaches full sensitivity near that point and stops there; no study recommends flagging everyone, and the excess harm is 16.26 per hundred — still several times the best threshold’s own harm, because a hundred diseased scores rarely include one right at the floor and the staircase stops at the lowest one it has, slightly above it.
The relative figures show how completely the ordering has turned over. In the earlier essay the staircase caused 15.4% more harm than the best threshold and the binormal fit 2.5% more. Here the staircase causes 289.9% more and the binormal fit 1196.0% more. The model that repaired the staircase’s blindness to a tail is the one that invents a tail when there is none.
Why the staircase stops a little high
The staircase’s excess on the skewed test has a distribution-free explanation. The chance that a new diseased person scores below the lowest of a study’s hundred diseased scores is whatever the shape of the scores — the arithmetic of order statistics that a range allowed a few misses built a whole specification on. A staircase that stopped exactly at its lowest diseased score would miss about one diseased person in a hundred, and at a prevalence of 26% with a miss costing a hundred false alarms that is about twenty-six false alarms’ worth per hundred people. The staircase’s cut sits halfway between its lowest diseased score and the next healthy score below it, which catches part of that gap, and its excess comes out at 16.26.
So the staircase’s error on a test with a floor is the price of a finite sample’s lowest value, and it shrinks as the study grows. The binormal fit’s error is the price of a shape, and a larger study only makes the fitted normal more confident about a tail that does not exist. A tail the sample never saw found the same division for a sum’s tail read from thirty draws: an approximation that was nearly exact given the source read the tail at a median of 0.061 of the truth given only a sample, because thirty draws hold nothing past their largest value. Past the data, a model’s tail is its assumption.
The skewed test never jumps
The best threshold for the skewed test flags 7.6% of healthy people at a prevalence of 10%, at 26% and at 40% alike: with no diseased score below the floor, flagging anyone further down catches no one, and no prevalence makes it worthwhile. The jump a threshold that jumps found is a property of the diseased scores’ lower tail, and a test without one has no jump to find. A binormal fit to its scores reports a jump that is not there — recommending that everyone be flagged in 32.9% of studies at 10% prevalence, 89.7% at 26% and 97.4% at 40% — because the fitted curve has the tail and the test does not.
A tail that is a lump
The mixture test fails the binormal fit in the opposite direction. One diseased person in ten scores around 1, and a normal curve fitted to all hundred diseased scores has a mean of about 3.3 and a spread only a little above one — it treats the mild cases as the lower edge of a single bell and gives them almost no probability below the upper healthy scores. The fitted model therefore stops early: it recommends flagging everyone in 0.1% of studies, where the best threshold flags 77.6% of healthy people, and its excess harm at 26% prevalence is 15.55 per hundred.
The staircase sees the mild cases directly — about ten of them sit among the healthy scores in every study — and does better, at 12.24. A two-normal mixture fitted to the diseased scores does better still, at 7.38, because the shape it can fit is the shape the truth has. But its flexibility has a cost of its own: in 23.5% of studies it decides the lump is part of a much wider lower tail and recommends flagging everyone, which on this test is never right.
Across prevalences the ordering holds. At 10% prevalence the fitted mixture’s excess is 1.54 per hundred against the binormal fit’s 2.10 and the staircase’s 4.86; at 40%, 8.58 against 24.04 and 21.88. The binormal fit is never the best model for a test with a lump, and at high prevalence it is the worst.
The model that was right costs little when kept
The two-normal mixture is not a free improvement, because on the binormal test — where it has one component too many — it gives back part of the binormal fit’s gain. At 26% prevalence its excess is 3.87 per hundred against the binormal fit’s 1.84, and it recommends flagging everyone in 33.9% of studies rather than 47.4%. On the skewed test it lands between the two others, at 22.71: a mixture of two normals cannot represent a hard floor either, only put less weight below it.
So no single model of the diseased scores wins on all three tests. The binormal fit wins when the scores are normal and loses badly when they are not; the staircase never wins and never loses badly; the mixture wins when there is a lump and is middling otherwise. The choice between them is a choice about the shape, and the shape is exactly what a published sensitivity and specificity leave out.
Whether a hundred diseased scores can say which
A validation study has its own diseased scores, and a hundred of them carry some information about their shape. The question is how much.
A straightness test — one minus the squared correlation of the sorted scores with their expected normal quantiles, the quantile plot read as a number — holds its level on binormal scores, rejecting 5.7% of samples of a hundred. On the skewed test it rejects 100.0%: a missing lower tail is plain in a hundred scores, and 90.3% of samples of only twenty-five show it. The lump is harder. A hundred diseased scores reveal it 60.4% of the time, twenty-five scores 19.2%, and it takes two hundred to see it in 87.6% of studies. A number for the shape found that no one summary reads every departure from a quantile plot, and here the missing tail and the lump are two departures of very different visibility.
That suggests a rule, and it can be priced. A study that tests its diseased scores for normality and fits the binormal model only when normality is not rejected — falling back on its own curve when it does not — recommends thresholds with an excess harm of 2.72 per hundred on the binormal test, against the binormal fit’s 1.84; 12.24 on the mixture test, where it mostly falls back on the staircase; and 16.26 on the skewed test, where it always does. It is never the worst of the four models on any of the three tests: it ties for best on the skewed test, where it always falls back on the staircase, and comes second on the other two.
What the published pair cannot carry
The test is a point somebody chose found that a sensitivity and specificity are one point on a curve, chosen by someone for some purpose, and say nothing about the rest of the curve. The three tests here share that point exactly, and the rest of their curves differ in precisely the region — the top right, where most healthy people are flagged — that decides whether a programme should flag everyone. A threshold that jumps showed that region deciding the best threshold for the binormal test; the measurement here shows the same region deciding which model of the scores to trust.
What a positive test is worth found that the value of a result depends on the prevalence as much as on the test; the value of a threshold, it turns out, depends on the tail as much as on the published pair.
A validation study that reports only its recommended threshold, or only a fitted binormal curve, has reported the conclusion of a modelling choice without the evidence for the choice. The diseased scores’ lower tail is the evidence, and it is the part of the data a hundred diseased people describe least well. Ten of them, on the mixture test, carry the whole of the lump; one or two, on the binormal test, carry the long tail.
What a validation study can report
The lower tail of the diseased scores, as a picture. A quantile plot of the diseased scores, or their lowest ten values beside the healthy distribution, shows a floor at once and a lump sometimes. It is the single display that separates the three tests here.
The threshold each model recommends, side by side. Where the staircase and the binormal fit agree, the tail is not doing the deciding. Where they disagree — one flagging everyone and the other stopping at the lowest diseased score — the disagreement is the finding, and it says the programme’s decision rests on a tail the study barely saw.
A normality check on the diseased scores before a binormal fit is used to extrapolate. It costs 0.88 per hundred on the binormal test and saves 50.82 on the skewed one.
More diseased people, if the decision is near a jump. The earlier essay found that the number of diseased people, not healthy ones, decides how well the tail is seen; here it also decides whether the shape can be told at all — a lump seen by 19.2% of studies of twenty-five diseased people is seen by 87.6% of studies of two hundred.
A validation study near a jump is in the position an upper limit is the finding described for a survey of a rare condition: what it can honestly report is bounded more tightly on one side than the other. A study whose diseased scores show a clear floor can say with confidence that flagging below the floor buys nothing. A study whose lowest diseased scores thin out gradually cannot say whether they stop or go on, and the choice between the staircase’s answer and a fitted curve’s is then a choice about an assumption the study cannot test with the people it has. Saying so, and saying which of the two answers the programme adopted, is the part of the report the published sensitivity and specificity will never carry.
Counted here
Counted over a thousand validation studies of a hundred diseased and three hundred healthy people at each setting, with a miss costing a hundred false alarms: at a prevalence of 26%, excess harms per hundred people of 11.41, 1.84, 3.87 and 2.72 for the staircase, binormal fit, two-normal mixture and checked rule on the binormal test; 12.24, 15.55, 7.38 and 12.24 on the mixture test; 16.26, 67.08, 22.71 and 16.26 on the skewed test.
Exact: the best thresholds and their harms for each shape, from its distribution function. Counted over two thousand samples at each size: the normality test’s rejection rates.
Not claimed: that the three shapes exhaust what diseased scores do. A test that saturates at a ceiling piles its diseased scores at the top, which leaves the lower tail — and the jump — untouched; a diseased group with two lumps, or a lump inside the healthy mean, would each need its own measurement. Not claimed either that a mixture of two normals is the best flexible model; a smoothed empirical distribution for the lower tail might keep the staircase’s refusal to invent and add some of the fit’s reach.
Still open: a threshold that admits it does not know
Every model here recommends one threshold. When the models disagree about the tail — the staircase stopping at the lowest diseased score, a fitted curve flagging everyone — the honest output of a validation study is not a threshold but an interval of thresholds consistent with its data, and a statement of how much harm each end of it would cause if the other were true.
A threshold with an uncertainty interval, built by resampling the study or by propagating the fitted models’ uncertainty, would say how often a study’s recommendation lies on the wrong side of the jump and how wide the range of defensible thresholds is. Whether such an interval is narrow enough to be useful at a hundred diseased people, how it behaves when the shape is wrong, and whether a programme should choose within it by minimising the worst case rather than the expected harm, have not been measured here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The liar with two answers — both name roc curve, sensitivity and specificity
- The prevalence the test has to estimate — both name prevalence, sensitivity and specificity
- The relation the table has to estimate — both name extrapolation, model misspecification
Named objects
A flat tag is an object no other essay names yet.
Decision thresholdExpected lossExtrapolationModel misspecificationNormalityPrevalenceROC curveSensitivity and specificity