Reversals that are not errors

Regression to the mean

Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.

Measure something twice. Take the group that did worst the first time and look at how they did the second time. They improved.

They will improve whatever is done to them in between, including nothing.

Two measurements of the same thing, correlated 0.60Pick the worst 15% on the first measurement and their average rises by 0.55 on the second. Pick the best and theirs falls by 0.59. No treatment was given to anybody.first measurementsecond+0.55-0.59predicted 0.64 from the correlation alonenobody was treated
Fig. 1 Five hundred pairs of measurements of the same underlying thing. The marked points are the worst and best fifteen per cent on the first measurement. Nobody was treated.

The measurement

With two measurements correlated at 0.6, selecting the worst 15% on the first and looking at the second gives an average improvement of about 0.72 standard deviations.

The closed form says the shift should be (1 − ρ) times the group’s distance from the mean, which for this selection predicts about 0.69. The figure computes both and the assertion behind it requires them to agree, so a frame where the mechanism did not account for the movement would fail the build.

The best group moves the other way by a similar amount. Nothing was done to either group.

12,000 studies of a real effect of 0.3, n = 16Power is 20%. The studies that reached significance report a mean effect of 0.613 — 2.04 times the truth. Every one of them is honest; the selection did the inflating.02505007501e+3-0.50000.5001effect the study reportsstudiesthe truth, 0.3what gets published, 0.61shaded: reached p below 0.05power 20%, inflation 2.04×
Fig. 2 The same mechanism selecting studies rather than individuals.

Why it happens

Any measurement is partly signal and partly noise. Selecting the extreme low group selects for both — people who are genuinely low, and people who had a bad measurement.

The genuine part persists to the second measurement. The noise part does not: a bad draw the first time says nothing about the second. So the selected group’s second measurement is closer to the mean, by exactly the fraction of their extremity that was noise.

That fraction is 1 − ρ, which is why the size of the effect is predictable in advance with no data about the intervention at all.

Move the correlation slider and the two limits are visible. At ρ = 0.95 the measurements are nearly the same thing and there is almost no movement. At ρ = 0 the first measurement carries no information about the second, and the selected group returns all the way to the population mean.

Two groups, one treatment, and both readings of the same numbersThe treatment wins in group A (93.0% against 87.0%) and in group B (73.0% against 69.0%), and loses overall (78.1% against 82.6%). Nothing here is a trick; the allocation differs between the groups.treatedcontrolverdictgroup A93.0%87.0%treatmentgroup B73.0%69.0%treatmentboth together78.1%82.6%controlgroup A: 88 treated of 351 · group B: 263 of 351wins in both groups, loses overallthe rates never changeonly the allocation does
Fig. 3 A neighbouring trap: correct numbers producing a conclusion that reverses on reweighting.

What it explains

The list is long, and its length is the reason this matters more than most statistical curiosities.

Interventions targeted at the worst cases. Remedial teaching for the lowest-scoring pupils, safety measures at the most dangerous junctions, management attention to the worst-performing branch. All of these select on an extreme measurement and all of them will show improvement whether or not they work. Evaluating them without a control group selected the same way is guaranteed to produce a positive result.

Punishment appearing to work better than praise. Kahneman’s flight instructor example: praise a pilot after an exceptionally good manoeuvre and the next is worse; criticise after a bad one and the next is better. Both are regression, and the instructor draws the opposite conclusion.

The sophomore slump, the Sports Illustrated cover jinx, the second album — the same mechanism wherever selection is on an exceptional performance.

Medical improvement after treatment begins. Patients typically seek treatment when symptoms are at their worst, which is a selection on an extreme measurement. Improvement follows regardless, which is a large part of what an uncontrolled trial measures and a substantial component of what gets attributed to placebo.

What a positive test means, sensitivity 90%, specificity 95%At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in ten, 33 are. The test has not changed.00.2500.5000.7501how common the condition ischance a positive result is true1 in 10,0001 in 1,0001 in 1001 in 101 in 11.8%66.7%one test, every prevalencethe base rate outweighs the test
Fig. 4 A third case where the direction of a conditional probability decides the answer.

Its relation to the winner’s curse

These are the same mechanism seen from two ends, which is worth making explicit.

The winner’s curse selects studies on an extreme estimate and finds the estimates inflated. Regression to the mean selects individuals on an extreme measurement and finds the next measurement less extreme.

In both cases: selection on a noisy quantity, followed by observation of the underlying quantity, produces a systematic gap in a direction set by which end was selected.

Recognising the shared structure is useful, because the fix is shared too. Do not evaluate on the measurement that was selected on. Select on one measurement and evaluate on an independent one, or select at random.

Twenty 95% intervals for a proportion that really is 0.350 of the twenty miss the true value. The 95% is a property of the procedure across repetitions — no single interval has a 95% chance of anything, because it either contains 0.35 or it does not.the truth, 0.350.00.20.40.60.80 of 20 missedthe 95% belongs to the procedure, not to one interval
Fig. 5 What a properly repeated procedure looks like, which is the standard a single before-and-after comparison is failing to meet.
What a one-sample t test at n = 20 can detectAt an effect of 0.5 standard deviations the test finds it 56% of the time. Below that, a non-significant result is the expected outcome of a real effect — which is why "no significant difference" is not evidence of no difference.00.2500.5000.750100.5001true effect, in standard deviationsprobability of rejecting80%, the usual target56% at d = 0.54,000 simulated studies per pointthe second number a p-value needs
Fig. 6 And the size of effect a study of a given size could detect, against the 0.7 standard deviations this mechanism supplies for free.

What a control group actually does

This clarifies something usually taught as a rule without a reason.

A control group is not there to make the study look rigorous. It is there because the selected group would have moved anyway, and the only way to know how much of the observed movement is the intervention is to watch a group that was selected the same way and not treated.

That last clause matters and is often missed. A control group drawn from the general population does not control for regression to the mean, because it was not selected on an extreme measurement and so has nothing to regress from. The control has to share the selection, not merely the treatment’s absence.

Get that wrong and the study measures regression to the mean and reports it as an effect — which is, measurably, worth about 0.7 standard deviations at a correlation of 0.6. That is larger than most real effects anyone is looking for.

Galton, and where the name comes from

The phenomenon was named by Francis Galton in the 1880s, studying the heights of parents and their adult children. Tall parents had children shorter than themselves on average; short parents had taller ones. He called it “regression towards mediocrity” and the modern sense of the word regression descends from that paper.

The part worth knowing is that Galton initially read it as a physical tendency — a force pulling the population back toward its mean, which would need explaining. It is not. It is a consequence of imperfect correlation, present whenever two measurements of anything are correlated at less than one, and requiring no mechanism at all.

That misreading is still the default. An intervention appears to work; the appearance requires no cause; and looking for the cause is the error.

The symmetry that gives it away

The diagnostic is simple and rarely applied: the effect runs both ways.

If the worst group improves because of regression, the best group must decline by the same mechanism. The figure shows exactly that — the marked groups at each end move toward the centre, and the shifts are comparable in size.

So a study reporting that the worst performers improved after an intervention has a free control available in its own data. Did the best performers decline over the same period? If they did, the mechanism is regression and the intervention has not been shown to do anything. If they held steady or improved, something else is happening.

That check costs nothing and is almost never done, because the best performers are not the group anyone was interested in.

Where the correlation comes from

The size of the effect is (1 − ρ), so everything depends on what determines ρ, and it is worth separating the two contributions.

Measurement error reduces ρ directly. A noisy instrument produces low correlation between repeat measurements and therefore large regression, regardless of how stable the underlying quantity is. Improving the instrument reduces the effect.

Genuine variation over time also reduces ρ, and improving the instrument does nothing about it. A quantity that really does fluctuate — pain, blood pressure, monthly sales — will show regression however precisely each measurement is taken.

Distinguishing them matters for what to do. If the correlation is low because of measurement noise, averaging several measurements at each time point raises ρ and shrinks the effect. If it is low because the quantity genuinely moves, averaging over a longer window does the same thing for a different reason. Either way, more measurement per occasion is the practical defence, and it is cheaper than a control group.

The version that shows up in league tables

The most consequential application, and the one where the arithmetic is routinely ignored.

Rank hospitals, schools or regions by an outcome measured over one period. The bottom performers are partly genuinely worse and partly unlucky. Measure again the next period and the bottom group improves — whatever was done to them, including nothing.

Two failures follow. Interventions targeted at the bottom appear effective, and their apparent effect is the regression. And the ranking itself is unstable, because much of what determined it was noise, so the following year’s ranking looks different and the difference gets reported as change.

The fix is the same as everywhere else in this essay: select on one period and evaluate on another, or account for the measurement error explicitly. Neither is difficult and both are rarer than the league table.

The unifying statement

Three essays on this site describe the same mechanism.

The winner’s curse selects studies on an extreme estimate. Regression to the mean selects individuals on an extreme measurement. The forking paths select an analysis on an extreme p-value.

In every case: select on a noisy quantity, then observe something that shares only the signal, and the observation moves toward the mean by the fraction that was noise.

Stated that way it stops being three phenomena and becomes one, with three names because it was discovered separately in three literatures.

Two measurements are not always available

The recommendation — select on one measurement, evaluate on another — assumes two measurements exist. Often they do not, and it is worth saying what is available then.

Split the single measurement. Where the measure is an aggregate of many small observations, half can select and the other half evaluate. This is the same idea at a smaller scale and it works.

Model the measurement error explicitly. If the reliability of the measure is known from elsewhere, the expected regression can be computed and subtracted. That is what the (1 − ρ) formula is for, and it is why the figure asserts the prediction rather than merely showing the movement.

Or select at random. An intervention delivered to a randomly chosen group has no selection to regress from, and the comparison is clean. That is less satisfying because it does not target the people who appear to need it most — but targeting on a noisy measurement is what created the problem.

The last option is usually rejected on ethical grounds, which is a real argument and worth weighing against the fact that an untargeted trial is the only one that will say whether the intervention does anything.

The version in performance management

A specific application, because it is where the mechanism causes the most avoidable damage.

Rank staff, teams or branches on a noisy performance measure. Intervene on the bottom. Measure again. They improve, and the intervention is credited.

Now intervene on the top — reward them, promote them, hold them up as an example. Measure again. They decline, and the intervention is blamed.

The manager who ran both experiments has learned that criticism works and praise does not, from data in which neither did anything at all. Kahneman’s account of the flight instructors is exactly this, and it is the clearest case in the literature of a firmly held belief produced by a statistical artefact.

The defence is the one this essay keeps arriving at: do not evaluate on the measure that selected. Rank on one period and evaluate on the next, or use a measure aggregated over long enough that the noise component is small.

What separates this from a real effect

The obvious question — how would anyone tell — has a clean answer.

A real effect moves the selected group beyond what regression predicts. The prediction is (1 − ρ) times the group’s distance from the mean, and ρ can be estimated from the data itself by correlating two measurements in an untreated group.

So the null hypothesis is not “no change”. It is “the change predicted by the correlation alone”, and that is a specific number rather than zero.

Testing against zero is what produces the false positives, and it is what almost every before-and-after evaluation does. Testing against the regression prediction requires knowing ρ and is otherwise no harder.

That is the constructive version of the whole essay: the artefact is not merely a warning, it is a quantitative prediction, and predictions can be subtracted.