Reversals that are not errors

The measurement that got them enrolled

Enrol the top tenth of one screening reading and give them nothing, and they fall by 0.702 standard deviations at follow-up. Measured from a fresh reading taken after enrolment they fall by nothing. Averaging ten screening readings still leaves 0.101, and it takes twenty-one to get under 0.05.

Worth reading first: Regression to the mean.

A trial for a condition defined by a number enrols people whose number is high. Blood pressure above a line, a symptom score above a line, a biomarker above a line: somebody is screened, the reading clears the cut, and the person is in. The same reading is then written down as their baseline, and at the end of the trial they are measured again.

The people who were given nothing improve. That is not news — the essay on regression to the mean measured the effect and gave its size, (1ρ)(1-\rho) times the selected group’s distance from the mean. What it did not ask, because it had only one reading to work with, is whether the effect belongs to the people or to the reading. It belongs to the reading, and the difference decides what a trial protocol can do about it. A reading used to select cannot also be the baseline. A second reading, taken after enrolment, removes the whole artefact. Averaging several readings at screening, which sounds like the same repair done more carefully, only shrinks it — and at the rate Spearman–Brown sets, slowly.

A cut on one reading

The model is the classical one for measurement. A person has a true score and every reading is that true score plus fresh error, standardised so a single reading has variance one; the test–retest correlation between two readings is ρ\rho, which is also the share of a reading’s variance that belongs to the person. Here ρ=0.6\rho = 0.6. The cut sits at the ninetieth percentile of one reading, 1.282 standard deviations, so it enrols the top tenth of whoever is screened. Nobody is treated.

On one cohort of two thousand people, 215 clear the cut. Their mean screening reading is 1.744 and their mean follow-up reading is 0.982: a fall of 0.762 ± 0.055 with nothing done to anyone.

A cohort screened once, the top tenth enrolled, and followed up with nothing given — correlation 0.62000 people read once at screening and once at follow-up, with a test–retest correlation of 0.6 and no treatment. The 215 above a cut at the top ten per cent of one reading (1.282 standard deviations) are enrolled. Their mean screening reading is 1.744 and their mean follow-up reading 0.982, a fall of 0.762 ± 0.055 with nothing done to anyone. The closed form for the fall is (1 − ρ) times the truncated-normal mean, (1 − 0.6) × 1.755 = 0.702.-3-2-10123-3-2-10123screening readingfollow-up reading, nothing givenenrolment cut215 enrolledmean screening1.744mean follow-up0.982fall, nothing given0.762 ± 0.055closed form0.7022000 people, correlation 0.6, top 10% of one reading enrolledthe untreated improve
Fig. 1 Two thousand people read at screening and again at follow-up, with nothing given. Those right of the dashed cut are enrolled. The horizontal line is their mean follow-up reading, well below every one of their screening readings; the slider sets the reading’s test–retest correlation.

The closed form for that fall needs only the mean of a normal cut off below — the same truncated normal that selecting on a sum of two causes is built from, applied here to one reading rather than to a pair. A standard normal truncated at cc has mean m(c)=φ(c)/(1Φ(c))m(c) = \varphi(c)/(1 - \Phi(c)), the inverse Mills ratio, which at the ninetieth percentile is 1.755. The enrolled group’s true scores are shrunk towards zero by ρ\rho — a reading is partly the person and partly the occasion, and only the person comes back — so the group’s expected follow-up is ρm(c)\rho\, m(c) = 1.053, and the fall is

(1ρ)m(c)=0.4×1.755=0.702.(1-\rho)\, m(c) = 0.4 \times 1.755 = 0.702 .

The cohort of two thousand read 0.762, about one of its standard errors above. On a cohort of 400,000 the same protocol reads 0.7054 ± 0.0041 from 39,782 enrolled people. The count never forms a truncated moment and the closed form never draws a person, which is the discipline two independent routes to one number enforce, and they agree.

The size is set by the cut, before anybody is screened

Nothing in (1ρ)m(c)(1-\rho)\,m(c) refers to the condition, the trial or the people. It is a property of how reliable the reading is and how selective the cut is, and it can be computed before the first screening visit.

At ρ=0.6\rho = 0.6, enrolling the top half gives a fall of 0.319 standard deviations; the top fifth 0.560; the top tenth 0.702; the top twentieth 0.825; the top hundredth 1.066. At a correlation of 0.4 and the top hundredth it is 1.599, and at 0.8 it is 0.533.

The more selective the cut, the more the untreated improve. The fall an untreated enrolled group shows at follow-up, (1 − ρ) times the mean of a normal truncated at the cut, against the share of the screened who are enrolled, at test–retest correlations of 0.4, 0.6, 0.8. At a correlation of 0.6 it is 0.319 standard deviations when half are enrolled, 0.702 at a tenth and 1.066 at one in a hundred; at 0.4 and one in a hundred it is 1.599, and at 0.8 it is 0.533. The point is counted on 400,000 simulated people at a tenth: 0.7054 ± 0.0041.
Fig. 2 The untreated fall against the share of the screened who are enrolled, at three test–retest correlations. The point is counted on 400,000 simulated people at a tenth and sits on its curve.

Two directions, both against intuition, and both already visible in the formula. The more selective the trial, the larger the phantom improvement — a trial restricted to the most severe cases, which is the trial most likely to look worthwhile, is the one whose untreated arm moves most. And the noisier the reading, the larger the improvement: a measurement that barely reproduces itself enrols a group that is mostly luck, and luck is what leaves.

A fall of 0.7 standard deviations is larger than most effects anybody runs a trial to find. A before-and-after comparison on a group enrolled this way, with no control arm, reports that effect for any treatment at all, including one that does nothing — which is the same mechanism as the winner’s curse, selecting people on a noisy number rather than studies on a noisy estimate.

The enrolled are not who the screening said

Before asking how to remove the artefact, it is worth saying what it does to the trial’s population, because that is where it does damage even when a control arm cancels it out of the comparison.

The enrolled group’s mean true score is 1.053. The cut was 1.282. The average person in the trial does not truly meet its entry criterion. Counted over the whole distribution rather than at the mean, the share of the enrolled whose true score is above the cut is 33.1% in closed form, and 33.2% on the 39,782 people of the large cohort. Two in three people enrolled for exceeding a threshold do not exceed it; they had a high day.

A second reading shows the same thing from the outside. Read everybody who was enrolled once more, after enrolment, and 60.98% of them fall under the cut. On a cohort of four thousand, 241 of the 394 enrolled — 61.2% — do.

Six in ten of the enrolled are under the cut at their next reading. The 394 people of 4000 whose screening reading passed the cut at 1.282, each read once more after enrolment, at a test–retest correlation of 0.6. 241 of them, 61.2%, read under the cut the second time. The closed form, ∫ φ(x)Φ((c − ρx)/√(1 − ρ²)) dx over the enrolled divided by the share enrolled, is 60.98%. The enrolled group's true mean is 1.053, well short of the 1.755 its screening readings average.
Fig. 3 Each enrolled person’s screening reading against a second reading taken after enrolment. The dashed line is the cut; everybody below it at the second reading would not have been enrolled on that day.

The share depends on the cut exactly the way the fall does. Enrol the top half, and 29.5% of the enrolled read under the line at a second visit and 78.2% truly exceed it; enrol the top hundredth, and 81.2% read under it and only 8.4% truly exceed it. A trial that screens hard for severity gets a population most of whose severity was the screening visit.

The integral behind those numbers is short. A person read at xx has true score distributed as N(ρx,ρ(1ρ))N(\rho x, \rho(1-\rho)), so the chance the true score clears the cut is Q((cρx)/ρ(1ρ))Q\big((c - \rho x)/\sqrt{\rho(1-\rho)}\big), and averaging that over the enrolled readings — a normal density from cc upwards, divided by the share enrolled — gives the share. The second-reading version replaces the true score’s spread with a reading’s. Neither involves a simulation; both are checked against one.

The entry cut is a screening test, and 33.1% is its predictive value

Read the cut as a diagnostic test and that 33.1% stops being a curiosity. The condition is having a true score above the line. The test is one reading. A positive result is a reading above the line. The share of positives who truly have the condition is what the essay on a positive screening result called the positive predictive value, and it obeys the same arithmetic.

The pieces are all computable. With a true-score variance of 0.6, the share of the population whose true score clears 1.282 is Q(1.282/0.6)Q(1.282/\sqrt{0.6}) = 4.90% — the prevalence. One reading catches 67.5% of those people, the test’s sensitivity, and correctly leaves out 93.0% of the rest, its specificity. A test that is two-thirds sensitive and ninety-three per cent specific, applied to a condition one person in twenty has, returns positives of whom a third are real. Nothing about the trial is unusual; the cut is simply a moderately good test for a fairly rare condition, and the base rate does what a base rate always does.

That reading also explains why a more selective trial is worse off. Moving the cut to the top hundredth makes the “condition” rarer — its prevalence falls to about one in seven hundred — and the predictive value falls with it, to the 8.4% above. The most selective trials are screening hardest for the rarest condition, which is where every screening test is least believable.

A fresh reading removes it; averaging only shrinks it

The mechanism says what repairs it. The screening reading carries the person’s true score and one occasion’s luck, and selection kept the people whose luck was good. Any later reading of the same person carries the true score and new luck, which selection never saw. So measured from a reading taken after enrolment, the untreated group’s follow-up differs by nothing: both readings are the same true score plus independent error, and the selection is already in the past of both.

That is exactly zero in closed form. On the cohort of 400,000 the fall measured from a fresh reading reads 0.0109 ± 0.0045, which is 2.4 of its own standard errors from zero — about as far as one count in sixty lands by chance, and the closed form here is not an approximation that could be slightly off.

Averaging several screening readings sounds like the same idea. It is not, because every reading that goes into the average is a reading the selection acted on. Put the same 400,000 people through five protocols, each enrolling on the mean of kk screening readings against the same clinical cut, and measure each group’s fall from that mean:

  • one screening reading: 0.705 against a closed form of 0.702
  • the mean of 2: 0.423 against 0.421
  • the mean of 3: 0.305 against 0.301
  • the mean of 5: 0.196 against 0.193
  • the mean of 10: 0.100 against 0.101
Measured from what enrolled them, the untreated fall; measured from a fresh reading, they do not. One simulated cohort of 400,000 people at a test–retest correlation of 0.6, a clinical cut at 1.282, and the same people put through six protocols. Measured from the screening reading that enrolled them, the untreated enrolled fall by 0.705, 0.423, 0.305, 0.196, 0.100 standard deviations when screening averages 1, 2, 3, 5, 10 readings, against closed forms of 0.702, 0.421, 0.301, 0.193, 0.101. Measured from a fresh reading taken after enrolment they fall by 0.0109 ± 0.0045, against a closed form of exactly zero — 2.4 standard errors away on this count.
Fig. 4 One cohort put through six ways of taking its baseline. Measured from the readings that enrolled them, the untreated fall by an amount that shrinks as more readings are averaged. Measured from one fresh reading taken after enrolment, they fall by nothing within the count’s noise.

The averaged reading is more reliable, and the formula already says by how much. The mean of kk readings has reliability

ρk=kρ1+(k1)ρ,\rho_k = \frac{k\rho}{1 + (k-1)\rho},

the Spearman–Brown formula, which is 0.7500 at two readings, 0.8824 at five and 0.9375 at ten. A more reliable reading enrols a group with less luck in it and the fall is (1ρk)(1-\rho_k) times a truncated mean, so it shrinks. But ρk\rho_k never reaches one, and the artefact is never removed.

How many readings averaging would need

Swept out to fifty readings the fall is 0.0520 at twenty readings and 0.0212 at fifty, where the reliability of the average is 0.9868. It first goes under 0.05 standard deviations at 21 readings, and under a tenth of the single-reading artefact at 15.

Averaging the screening readings shrinks the artefact and never removes it. The untreated enrolled group's fall when the screening reading is the mean of k readings, at a test–retest correlation of 0.6 and a fixed cut of 1.282. The average's reliability is Spearman–Brown's kρ/(1 + (k − 1)ρ), which is 0.7500 at two readings, 0.8824 at five and 0.9375 at ten, and the fall is 0.702, 0.421, 0.193 and 0.101. It first drops below 0.05 at 21 readings and is still 0.0212 at fifty. The points are counted on 400,000 people. A single fresh reading after enrolment puts it at zero.
Fig. 5 The untreated fall as the number of averaged screening readings grows, with the counted protocols as points. It falls fast and then slowly, and at no number of readings does it reach the zero a single fresh reading gives.

The decay is slow because 1ρk=(1ρ)/(1+(k1)ρ)1 - \rho_k = (1-\rho)/(1 + (k-1)\rho) falls like 1/k1/k, and the truncated mean does not fall at all — a more reliable screening reading is still a reading of people selected for being high. Twenty readings to buy what one extra reading, taken at a different moment, buys outright.

The averaging protocols also enrol fewer people at the same clinical cut: a tenth of those screened at one reading, 7.60% at two, 6.73% at three, 5.46% at ten and 5.01% at fifty. The people who drop out are the ones whose single high reading was a high day. That is the averaging doing its job, and it is worth knowing when the averaged protocol is costed, because the screening has to reach nearly twice as many people to fill the same trial.

So the two repairs are not the same repair at different strengths. Averaging screening readings changes who is enrolled; a fresh baseline changes what the enrolled are compared against. Only the second makes the untreated arm’s change a measurement of the treatment rather than of the screening, and protocols that do both — average readings to decide eligibility, then take a separate reading as the baseline — are using each for what it does.

What makes a reading fresh

The repair rests on a single property of the second reading, and it is not the number of readings or their precision. It is that the second reading’s error is independent of the error the selection acted on. In the model here that holds by construction, because every reading draws new error. In a clinic it is a claim about the protocol.

A second reading taken five minutes after the first, in the same room, by the same person, at the same point in the day, shares much of whatever made the first one high: the walk to the clinic, the anxiety of the visit, the week’s diet, the season. To that extent it is not a fresh reading but a second look at the same luck, and it removes only the part of the artefact that belonged to what the two readings do not share. The quantity that governs the repair is therefore the correlation between readings taken the way the baseline and the follow-up are actually taken — different visits, weeks apart — and not the agreement of two readings taken back to back, which is always higher.

The same point bears on averaging. Readings averaged within one visit shrink only the within-visit part of the error; the part that is shared across a visit is carried into the average undiminished, and the effective Spearman–Brown gain is smaller than the count of readings suggests. None of this is measured here — the model has one kind of error, not two — so it is stated as the consequence of the mechanism rather than as a number, and the number it would need is a variance-components estimate of how much of a measurement’s noise is the visit.

Correcting afterwards, and the correlation that misleads

Suppose the trial is already over, the baseline was the screening reading, and there is no control arm. The closed form says the artefact is (1ρ)(1-\rho) times the enrolled group’s distance from the population mean, so it can be subtracted — if ρ\rho is known. The data at hand contain two readings of every enrolled person, and the obvious estimate of ρ\rho is their correlation.

That estimate is wrong, and by a large amount. Inside the enrolled group the correlation between screening and follow-up is 0.295 in closed form and 0.2902 counted, against a population value of 0.6. Plugged into the correction it removes (10.295)×1.755=(1 - 0.295) \times 1.755 = 1.238 standard deviations, which is 176% of the artefact — the correction over-subtracts by 0.536, and a treatment that did nothing is reported as harmful.

Inside the enrolled group the slope is still ρ and the correlation is not. The test–retest correlation and the regression slope of follow-up on screening, read only among the enrolled, against how selective the cut is, at a population correlation of 0.6. The slope is 0.6 at every cut, because the cut truncates the screening reading and E[follow-up | screening] = ρ × screening holds at every screening value. The correlation falls — 0.412 when half are enrolled, 0.295 at a tenth, 0.227 at one in a hundred — because the screening reading's variance inside the group is 0.169 at a tenth rather than one. Counted on the 39,782 enrolled of 400,000: correlation 0.2902, slope 0.5918.
Fig. 6 Read inside the enrolled group, the slope of follow-up on screening stays at the population’s 0.6 at every cut while the correlation falls with every tightening of it. Points are counted at a tenth.

The reason is the cut. Selecting on the screening reading truncates it, and a truncated variable has less variance: inside the enrolled tenth the screening reading’s variance is 0.169 rather than one. A correlation is a slope rescaled by the two variables’ spreads, and one spread has collapsed, so the correlation collapses with it — to 0.412 when half are enrolled and 0.227 at a hundredth. The slope does not. The expected follow-up given the screening reading is ρ\rho times the reading at every reading, selected or not, so the regression slope inside the enrolled group is still 0.6 — 0.5918 counted — and a correction built on the slope is right.

This is range restriction, and it is the same fact complete-case regression rests on: selecting on the regressor leaves the regression of the outcome on it untouched, while every summary that involves the regressor’s spread moves. The practical sentence is short. Read the slope inside a selected group, not the correlation.

What a control arm fixes, and what it does not

A randomised control arm enrolled by the same cut falls by the same 0.702, and the difference between the arms cancels it. That is the standard defence and it is correct, with two limits worth stating.

It cancels the artefact from the comparison and leaves it in each arm’s change, so any report of how much the treated arm improved is still mostly screening. And it does nothing for the population: the trial still enrolled a group two thirds of whom do not meet its entry criterion, and the effect it estimates is an effect in that group.

That second limit reaches back into the trial’s size. A sample size is computed for an effect in the population the trial means to enrol, and if a treatment’s benefit grows with the severity it treats, the severity that matters is the enrolled group’s true one. On paper the enrolled average 1.755 standard deviations; in truth they average 1.053, a ratio of exactly ρ\rho60%. A benefit proportional to severity is then 60% of the benefit the calculation assumed, and since the number of subjects a power calculation asks for scales as one over the square of the effect, the trial needs 2.78 times the people it planned for to have the power it reports. Proportional benefit is an assumption, and it is stated as one: a treatment whose benefit does not depend on severity loses nothing here. But the direction of the error does not depend on the assumption — a severity-driven benefit can only be smaller in the enrolled group than the screening readings promise, never larger.

It also changes what adjusting for baseline buys. Under randomisation adjusting for a baseline is only a question of precision, and the precision bought is 1r21 - r^2 of the unadjusted variance, where rr is the covariate’s correlation with the outcome — inside the enrolled group. For the screening reading that correlation has collapsed to 0.295, so adjusting for it leaves 0.913 of the variance. A fresh reading taken after enrolment keeps more of its correlation with the follow-up, 0.429 in closed form and 0.4263 counted, because it was not truncated; adjusting for it leaves 0.816. The reading that is right for the comparison is also the better covariate, and for the same reason: selection did not touch it.

What the closed forms assume

Normal true scores and normal error. That is what makes the enrolled group’s expected true score ρ\rho times its reading, a straight line. For a population whose true scores are skewed or heavy-tailed the shrinkage is not proportional, and the artefact at a given cut is not (1ρ)m(c)(1-\rho)\,m(c) — the next question below.

Readings exchangeable over time. A condition that genuinely improves after the moment people seek care, or a seasonal quantity screened in its bad season, adds a real change to the artefact and a fresh baseline does not remove that.

One fixed cut. Protocols that re-screen people who narrowly fail, or that enrol on a combination of readings, change the selection and so change the closed form; the principle — only a reading the selection did not see can be a clean baseline — does not change.

A large cohort. The counts use 400,000 simulated people so that the closed forms can be checked to three decimals. A real trial’s group of fifty or a hundred has the same expected artefact and a standard error around it of a tenth of a standard deviation, which is the same order as the effects it is looking for.

What is proved is the truncated-normal algebra, Spearman–Brown, and the untouched slope; what is measured is that simulated cohorts land on them. What is only argued is that real outcome measures behave like true score plus error — for most of them it is the working assumption rather than something established.

Where this goes next

Every closed form in this essay leans on one line: the enrolled group’s expected true score is ρ\rho times its reading. That line is exactly what a normal population of true scores gives, and nothing guarantees a population is normal. A population of true scores with a heavy tail — a few people genuinely far out, many near the middle — is the ordinary shape of severity, income, output and risk, and there a high reading is more likely to be a genuinely high person than the correlation suggests. The next question is how far the regression of a selected group depends on the shape of the population rather than on its correlation: a lead that a heavy tail keeps holds the correlation fixed at exactly 0.6, changes only the population, and finds that the top one per cent keeps anywhere from 44% to 78% of its lead.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCovariate adjustmentThe inverse Mills ratioMeasurement errorMonte CarloRange restrictionRegression to the meanScreeningSelection biasSpearman–BrownTest–retest reliabilityThresholdTruncated normal