Two tests, a threshold, and the rate they are read against

The second test that is not a second opinion

Two positives from a 90/95 test on a one-in-a-thousand condition give a 24.49% chance of disease if the tests are independent. At a correlation of 0.1 between their errors it is 10.16%, and at 0.5 it is 3.16% — barely more than the 1.77% one positive was worth.

Worth reading first: What a positive test is worth.

The base rate was always Bayes ends on the sequential version: a second positive multiplies the odds by the likelihood ratio again, so a test that is worth 18 to 1 on its own is worth 324 to 1 twice. The arithmetic is right and it has an assumption in it that the page names in one clause and does not price.

What a second positive is worth, prevalence 0.10%A 90% sensitive, 95% specific test. One positive gives 1.77%. Two independent positives give 24.49%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 10.16%, and at 0.5 it is 3.16%.00.1000.20000.2000.4000.600correlation between the two tests' errorschance of disease after two positiveswhat independence would giveone positive90% sensitive, 95% specificthe multiplication assumes the axis is zero
Fig. 1 What two positive results are worth, as the correlation between the two tests’ errors moves. The dashed line is what multiplying the likelihood ratios says; the lower line is what the tests are actually worth.

The assumption, written out

Multiplying likelihood ratios requires the two results to be conditionally independent given the disease status: among people who have the condition, the second test’s result must carry no information about the first’s, and the same among people who do not.

That is not the same as the two tests being unrelated. Of course they are related — they are both detecting the same thing. Conditional independence says that once the truth is fixed, the remaining randomness in the two results is unconnected.

For it to hold, the two tests must fail for entirely separate reasons. In practice they fail for overlapping ones:

The same sample. A blood draw that was haemolysed, a swab that missed, a specimen stored warm — anything wrong with the material is wrong for both assays run on it.

The same mechanism. Two antibody tests detecting the same epitope cross-react with the same other antigen. A person whose blood contains that antigen is a false positive on both, every time.

The same person. Some people simply sit near a threshold for reasons of their own physiology, and they do so on Tuesday as well as on Monday.

The same laboratory. A calibration drift, a reagent lot, an operator’s technique — all of these are shared by two tests run in the same place at the same time.

Each of those makes the errors positively correlated, and positive correlation is the direction that destroys the calculation.

What it costs

At a prevalence of one in a thousand with a 90% sensitive, 95% specific test:

correlation between the tests’ errors chance of disease after two positives
0 (independent) 24.49%
0.05 14.33%
0.10 10.16%
0.20 6.46%
0.30 4.76%
0.50 3.16%
0.70 2.39%

One positive is worth 1.77%. So at a correlation of 0.5 the second positive has taken the answer from 1.77% to 3.16% — it has roughly doubled the odds, where the independent arithmetic promised to multiply them by eighteen.

The collapse is steep at the left-hand end and that is the part worth understanding. Going from no correlation to 0.05 — a correlation most people would call negligible — takes the answer from 24.49% to 14.33%, losing 41% of it.

Why a small correlation does so much

The arithmetic is short and it explains the steepness.

Under independence, two false positives happen with probability (1 − spec)² = 0.05² = 0.0025. That is a very small number, and it is the entire reason two positives are convincing: a healthy person almost never produces them.

With a correlation ρ between the two results among the healthy, the probability becomes

(1spec)2+ρ(1spec)spec=0.0025+0.0475ρ.(1-\text{spec})^2 + \rho\,(1-\text{spec})\,\text{spec} = 0.0025 + 0.0475\,\rho .

The added term has 0.0475 in front of it, which is nineteen times the base term. So a correlation of 0.05 adds 0.0024 — it doubles the rate of double false positives — and a correlation of 0.1 triples it.

Meanwhile the numerator barely moves: two true positives go from 0.81 to 0.81 + 0.09ρ, which at ρ = 0.1 is a 1% increase.

So a small correlation multiplies the false-positive pairs and leaves the true-positive pairs alone, and the predictive value is a ratio of the two. The asymmetry is structural: the term that grows is proportional to spec·(1 − spec), which is large exactly when the false-positive rate is small, and a small false-positive rate is the whole basis of the sequential argument.

What a second positive is worth, prevalence 1.00%. A 90% sensitive, 95% specific test. One positive gives 15.38%. Two independent positives give 76.60%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 53.29%, and at 0.5 it is 24.76%.
Fig. 2 The same collapse at a prevalence of one in a hundred, where the independent answer is 76.6% and a correlation of 0.2 takes it to 41.1%. The relative damage is smaller because the prior is larger.

What the correlation is not

Two confusions are worth heading off, because both make the assumption sound more defensible than it is.

It is not the correlation between the two tests’ results. Those are strongly correlated whatever happens, because both are detecting the same condition — a population of a thousand people will show a thousand pairs of results with a large positive association in them, and that association is mostly the disease. The quantity here is the correlation within the healthy group and within the diseased group separately, and it can be zero while the overall association is enormous.

That distinction is the same one a stratified comparison draws, and it fails the same way: a marginal association says nothing about a conditional one, and only the conditional one enters the arithmetic.

And it is not removed by the tests being run by different people. Independence of operators removes one source and leaves the mechanism and the specimen. Two laboratories running the same assay on splits of one sample share everything except the operator, and the shared specimen is usually the largest term.

The practical consequence is that a report of confirmation by a second test says very little on its own. What matters is which of the four shared sources was broken, and a repeat of the same assay on the same sample breaks none of them.

Where the damage is worst

The prevalence changes how much of this matters, and the direction is the unhelpful one.

At one in a hundred, two independent positives give 76.6% and a correlation of 0.2 gives 41.1% — a substantial loss, and both numbers are still in a range where the result is informative.

At one in ten thousand, two independent positives give 3.1% and a correlation of 0.2 gives 0.7%. The sequential test was the only hope of making a positive mean anything at that prevalence, and the correlation removes it.

So the calculation is least robust exactly where it is most needed. A common condition does not need a second test to make a positive interpretable; a rare one does, and a rare one is where the second test’s contribution is most fragile.

The same shape appeared in the predictive value’s own sweep: the base rate does more work than the test’s characteristics, and here the base rate also decides how much a violated assumption costs.

How large a correlation is plausible

The table above is only useful with some sense of where on it a real pair of tests sits, and the four shared sources give an argument about that even without a measurement.

Consider a healthy person’s chance of a false positive. The 5% specificity failure rate is an average over the population, and it is not the same for everyone: a person with a cross-reacting antibody, a borderline baseline value or an unusual physiology has a much higher personal rate, and most people have a much lower one. That heterogeneity is the correlation — two draws from the same person are two draws from that person’s own rate rather than from the population’s.

The arithmetic follows from the variance of the personal rates. If a small share of the population has a personal false-positive probability near 1 and the rest near 0, the correlation is near 1 even though the average is 5%. If everyone has exactly 5%, the correlation is exactly 0.

Which of those is nearer the truth is a question about biology, and the fact that some individuals are known to false-positive repeatedly on particular assays — that being why a person is told their result is “a known interference” rather than retested — is direct evidence for the first. A population in which repeat false positives happen at all is a population with a positive correlation, and the arithmetic above says a correlation of 0.05 is already enough to lose two fifths of the second test’s value.

So the honest prior is that the correlation is not negligible, and the burden is on a calculation that assumes it away.

What can be done

Four things, in descending order of how often they are available.

Use a second test with a different mechanism. Conditional independence is a property of the pair, not of either test. An antibody test followed by a nucleic acid test have almost nothing in common except the organism, and their errors have very little reason to correlate. That is the standard design for confirmatory testing and it is the whole reason confirmatory tests are usually a different kind of test rather than a repeat.

Take a new sample. Much of the correlation lives in the specimen. A second draw removes the shared-material component and leaves the shared-mechanism one.

Measure the correlation. It is estimable — run both tests on a panel of known negatives and count the double positives against the product of the single rates. That number is what the calculation needs and it is almost never reported, because a validation study reports each test’s characteristics separately.

Report the answer as a range. If the correlation is unknown, the predictive value after two positives is not a number; it is an interval running from the independent value down to something near the single-test value — the same honest move as sweeping an unknown prevalence rather than choosing one, applied to the parameter one level up. Reporting 24.49% when the honest statement is “somewhere between 2% and 24% depending on a quantity nobody measured” is the failure this page is about.

A thousand people, a condition one in 1,000 has, a test that is 90% and 95%. 1 people have it and test positive. 50 do not have it and test positive anyway. So of the 51 positive results, 2% are right — and that is with a test most people would call accurate.
Fig. 3 The counting the independence assumption is applied to. Fifty false positives in a thousand people, and the sequential argument turns on how many of those fifty would test positive again.

The same assumption everywhere else

Conditional independence is not a screening-specific assumption; it is the assumption under every multiplication of likelihoods, and it fails in the same direction in every one.

A naive Bayes classifier multiplies the likelihood of each feature given the class. Features correlate, the multiplication overstates the evidence, and the classifier’s posterior probabilities are notoriously overconfident — often at 0.999 when it is right two thirds of the time. The classification is frequently still good, because the ordering survives; the probability does not.

Combining independent studies multiplies their likelihoods, and studies run by the same group on the same population with the same instrument share errors.

A panel of tests is the extreme case: ten correlated markers multiplied together produce posteriors near certainty from evidence worth a fraction of that — which is the effective-number problem with the sign reversed, since there correlation made a correction too severe and here it makes a conclusion too strong.

The general statement is the one this page measures in a single case: multiplying evidence assumes the evidence is separate, and the failure is always in the confident direction. Nothing on the page is a criticism of Bayes’ rule, which is correct; it is a criticism of the factorisation, which is an assumption.

The arithmetic in odds, where the failure is one number

The odds form is what makes the sequential calculation portable, and it also makes the failure easy to state.

Prior odds times likelihood ratio gives posterior odds. One positive from a 90/95 test has a likelihood ratio of 0.90/0.05 = 18. Two independent positives have 18 × 18 = 324.

With correlated errors the pair’s likelihood ratio is not 324. It is

P(both+D)P(both+¬D)=0.81+0.09ρ0.0025+0.0475ρ,\frac{P(\text{both} + \mid D)}{P(\text{both} + \mid \neg D)} = \frac{0.81 + 0.09\rho}{0.0025 + 0.0475\rho},

which is 324 at ρ=0\rho = 0, 167 at ρ=0.05\rho = 0.05, 69 at ρ=0.2\rho = 0.2 and 32.6 at ρ=0.5\rho = 0.5.

The last number is the one to carry, and it is best read against 18. Two positives from tests correlated at 0.5 are worth a factor of 32.6, where a single positive is worth 18 — so the second test has not doubled the evidence, it has added about eighty per cent of one more positive. The independent arithmetic promised to square the first test’s contribution and it has delivered less than one and a quarter of it.

A likelihood ratio is the quantity a test should be reported with, and the sequential likelihood ratio of a pair is a different quantity from the product of the two singles. Nothing in current practice reports the first, and the second is what gets used.

What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.
Fig. 4 The single test’s predictive value against prevalence, which is the curve the sequential argument is trying to escape. The escape works only when the second result is genuinely new information.
What a second positive is worth, prevalence 5.00%. A 90% sensitive, 95% specific test. One positive gives 48.65%. Two independent positives give 94.46%, which is what multiplying the likelihood ratios says. At a correlation of 0.1 between the tests' errors it is 85.60%, and at 0.5 it is 63.16%.
Fig. 5 Where the assumption barely matters. At a prevalence of one in twenty, two independent positives give 94.5% and a correlation of 0.2 gives 78.4% — worse, and still a result worth having, which is the regime the sequential calculation is safe in.

What is claimed here, and what is not

Two statements, and the first is what makes the second a measurement rather than a different model.

At zero correlation the arithmetic here is the independent arithmetic, to 10⁻¹². That is the reason the correlated form can be trusted as a generalisation: a parameterisation that did not reduce would be a different model whose numbers happen to be lower, and the comparison with the independent value would mean nothing.

A correlation of 0.5 takes most of what the second positive was worth, stated as a factor rather than a value, so it holds at every prevalence the slider takes.

The reading that does not survive is the multiplication itself. Multiplying the two likelihood ratios gives 24.49% after two positives where the correlated arithmetic gives 3.16%, and the two would agree only if the correlation were not entering at all. That is the plausible slip here: a parameter accepted, stored and never used produces a picture indistinguishable from one whose parameter does nothing, and the reduction at zero is what tells the two apart.

What a third test adds, and it is less than the second

The natural extension is a third positive, and the extension of the failure is worse than proportional.

Under independence three positives multiply the odds by 18³ = 5,832, and the predictive value at one in a thousand reaches 85.4%. Under correlation the third result is, by the same argument as the second, largely predictable from the first two: a person whose physiology or specimen produces false positives has produced two and will produce a third.

So the sequence’s evidence saturates. Each additional correlated test adds less than the last, and in the limit of perfect correlation the whole sequence is worth exactly one test however long it runs. The independent calculation has no saturation in it at all — it multiplies indefinitely, and its answer approaches certainty after four or five positives from any test better than a coin.

That difference is qualitative rather than numerical, and it is the reason the assumption deserves more attention than a clause. A model that saturates and a model that does not are different models of what repeated testing is for, and the choice between them is being made silently by a factorisation.

The practical form is a rule about protocols. A protocol that repeats a test until it agrees with itself is measuring the correlation rather than the condition, and the number of repeats it specifies is the strongest evidence available that nobody computed what the repeats were worth.

Still open: the correlation nobody has measured

Everything on this page is parameterised by a quantity that is estimable and is not estimated.

A test’s validation study reports sensitivity and specificity against a reference standard. Running two candidate tests on the same validation panel would give the correlation between their errors directly, at no extra cost in samples — the panel is already assembled and the tests are already being run.

It is not reported. The literature on combining diagnostic tests knows the assumption matters and mostly proceeds by assuming it away, occasionally with a sensitivity analysis over a range of correlations chosen without evidence — which is a parameter swept because nobody measured it rather than because the sweep is the answer.

What is not known here is how large the real correlations are. If they are typically 0.02, the sequential calculation is roughly right and this page is a caution. If they are typically 0.3, most published two-test predictive values are wrong by a factor of five. The measurement is a day’s work on any existing validation panel, and the number would settle which of those two sentences is the right one.

There is a second, harder question behind it. The model here gives the two tests one correlation, the same in both status groups, and the real structure is almost certainly richer: the errors among the healthy come from cross-reactivity and the errors among the ill come from low titre, and there is no reason those should share a number. Whether the single parameter is an adequate summary, or whether the two groups have to be modelled separately — which a shared summary standing in for two quantities is exactly the situation a shared summary warns about — is not settled by anything here.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Base rateConditional independenceLikelihood ratioPosteriorPredictive valueScreeningSensitivity and specificitySequential testing