Regression to the mean
Worth reading first: Simpson's reversal is a region, not a table.
Measure something twice. Take the group that did worst the first time and look at how they did the second time. They improved.
They will improve whatever is done to them in between, including nothing.
The measurement
With two measurements correlated at 0.6, selecting the worst 15% on the first and looking at the second gives an average improvement of about 0.72 standard deviations.
The closed form says the shift should be (1 − ρ) times the group’s distance from the mean, which for this selection predicts about 0.69. The figure computes both and the assertion behind it requires them to agree, so a frame where the mechanism did not account for the movement would fail the build.
The best group moves the other way by a similar amount. Nothing was done to either group.
Why it happens
Any measurement is partly signal and partly noise. Selecting the extreme low group selects for both — people who are genuinely low, and people who had a bad measurement.
The genuine part persists to the second measurement. The noise part does not: a bad draw the first time says nothing about the second. So the selected group’s second measurement is closer to the mean, by exactly the fraction of their extremity that was noise.
That fraction is 1 − ρ, which is why the size of the effect is predictable in advance with no data about the intervention at all.
Move the correlation slider and the two limits are visible. At ρ = 0.95 the measurements are nearly the same thing and there is almost no movement. At ρ = 0 the first measurement carries no information about the second, and the selected group returns all the way to the population mean.
The shift is a measurement of the correlation
The closed form has a use beyond predicting the movement, and it is the one that matters for reading somebody else’s study.
The expected return is times the selected group’s distance from the mean. So dividing an observed “improvement” by that distance recovers — which means a single-arm before-and-after study on a selected group measures the test–retest correlation of its outcome and reports it as a treatment effect.
Written as a table, the share of a selected group’s distance from the mean that comes back for nothing:
- ρ = 0.3 — 70% of the distance returns
- ρ = 0.6 — 40%
- ρ = 0.9 — 10%
So the same intervention, applied to the same people, produces an apparent effect that varies by a factor of seven depending on nothing but how reproducible the measurement is.
What a real effect has to beat
That gives the number a study needs before any of its improvement belongs to it.
A group selected a standard deviation and a half below the mean, on an outcome with a test–retest correlation of 0.6, returns 0.6 standard deviations on its own. A treatment producing a 0.5 standard-deviation improvement in such a group has produced less than nothing detectable — the group would have moved further without it, and the study would still report success.
The threshold is not zero and it is not small. It is times however far out the selection went, it is computable before the study runs from a pilot’s own repeat measurements, and any study that selects on a baseline and has no control arm should report it beside its result.
What it explains
The list is long, and its length is the reason this matters more than most statistical curiosities.
Interventions targeted at the worst cases. Remedial teaching for the lowest-scoring pupils, safety measures at the most dangerous junctions, management attention to the worst-performing branch. All of these select on an extreme measurement and all of them will show improvement whether or not they work. Evaluating them without a control group selected the same way is guaranteed to produce a positive result.
Punishment appearing to work better than praise. Kahneman’s flight instructor example: praise a pilot after an exceptionally good manoeuvre and the next is worse; criticise after a bad one and the next is better. Both are regression, and the instructor draws the opposite conclusion.
The sophomore slump, the Sports Illustrated cover jinx, the second album — the same mechanism wherever selection is on an exceptional performance.
Medical improvement after treatment begins. Patients typically seek treatment when symptoms are at their worst, which is a selection on an extreme measurement. Improvement follows regardless, which is a large part of what an uncontrolled trial measures and a substantial component of what gets attributed to placebo.
Its relation to the winner’s curse
These are the same mechanism seen from two ends, which is worth making explicit.
The winner’s curse selects studies on an extreme estimate and finds the estimates inflated. Regression to the mean selects individuals on an extreme measurement and finds the next measurement less extreme.
In both cases: selection on a noisy quantity, followed by observation of the underlying quantity, produces a systematic gap in a direction set by which end was selected.
Recognising the shared structure is useful, because the fix is shared too. Do not evaluate on the measurement that was selected on. Select on one measurement and evaluate on an independent one, or select at random.
What a control group actually does
This clarifies something usually taught as a rule without a reason.
A control group is not there to make the study look rigorous. It is there because the selected group would have moved anyway, and the only way to know how much of the observed movement is the intervention is to watch a group that was selected the same way and not treated.
That last clause matters and is often missed. A control group drawn from the general population does not control for regression to the mean, because it was not selected on an extreme measurement and so has nothing to regress from. The control has to share the selection, not merely the treatment’s absence.
Get that wrong and the study measures regression to the mean and reports it as an effect — which is, measurably, worth about 0.7 standard deviations at a correlation of 0.6. That is larger than most real effects anyone is looking for.
Galton, and where the name comes from
The phenomenon was named by Francis Galton in the 1880s, studying the heights of parents and their adult children. Tall parents had children shorter than themselves on average; short parents had taller ones. He called it “regression towards mediocrity” and the modern sense of the word regression descends from that paper.
The part worth knowing is that Galton initially read it as a physical tendency — a force pulling the population back toward its mean, which would need explaining. It is not. It is a consequence of imperfect correlation, present whenever two measurements of anything are correlated at less than one, and requiring no mechanism at all.
That misreading is still the default. An intervention appears to work; the appearance requires no cause; and looking for the cause is the error.
The symmetry that gives it away
The diagnostic is simple and rarely applied: the effect runs both ways.
If the worst group improves because of regression, the best group must decline by the same mechanism. The figure shows exactly that — the marked groups at each end move toward the centre, and the shifts are comparable in size.
So a study reporting that the worst performers improved after an intervention has a free control available in its own data. Did the best performers decline over the same period? If they did, the mechanism is regression and the intervention has not been shown to do anything. If they held steady or improved, something else is happening.
That check costs nothing and is almost never done, because the best performers are not the group anyone was interested in.
Where the correlation comes from
The size of the effect is (1 − ρ), so everything depends on what determines ρ, and it is worth separating the two contributions.
Measurement error reduces ρ directly. A noisy instrument produces low correlation between repeat measurements and therefore large regression, regardless of how stable the underlying quantity is. Improving the instrument reduces the effect.
Genuine variation over time also reduces ρ, and improving the instrument does nothing about it. A quantity that really does fluctuate — pain, blood pressure, monthly sales — will show regression however precisely each measurement is taken.
Distinguishing them matters for what to do. If the correlation is low because of measurement noise, averaging several measurements at each time point raises ρ and shrinks the effect. If it is low because the quantity genuinely moves, averaging over a longer window does the same thing for a different reason. Either way, more measurement per occasion is the practical defence, and it is cheaper than a control group.
The version that shows up in league tables
The most consequential application, and the one where the arithmetic is routinely ignored.
Rank hospitals, schools or regions by an outcome measured over one period. The bottom performers are partly genuinely worse and partly unlucky. Measure again the next period and the bottom group improves — whatever was done to them, including nothing.
Two failures follow. Interventions targeted at the bottom appear effective, and their apparent effect is the regression. And the ranking itself is unstable, because much of what determined it was noise, so the following year’s ranking looks different and the difference gets reported as change.
The fix is the same as everywhere else in this essay: select on one period and evaluate on another, or account for the measurement error explicitly. Neither is difficult and both are rarer than the league table.
The unifying statement
Three essays on this site describe the same mechanism.
The winner’s curse selects studies on an extreme estimate. Regression to the mean selects individuals on an extreme measurement. The forking paths select an analysis on an extreme p-value.
In every case: select on a noisy quantity, then observe something that shares only the signal, and the observation moves toward the mean by the fraction that was noise.
Stated that way it stops being three phenomena and becomes one, with three names because it was discovered separately in three literatures.
Two measurements are not always available
The recommendation — select on one measurement, evaluate on another — assumes two measurements exist. Often they do not, and it is worth saying what is available then.
Split the single measurement. Where the measure is an aggregate of many small observations, half can select and the other half evaluate. This is the same idea at a smaller scale and it works.
Model the measurement error explicitly. If the reliability of the measure is known from elsewhere, the expected regression can be computed and subtracted. That is what the (1 − ρ) formula is for, and it is why the figure asserts the prediction rather than merely showing the movement.
Or select at random. An intervention delivered to a randomly chosen group has no selection to regress from, and the comparison is clean. That is less satisfying because it does not target the people who appear to need it most — but targeting on a noisy measurement is what created the problem.
The last option is usually rejected on ethical grounds, which is a real argument and worth weighing against the fact that an untargeted trial is the only one that will say whether the intervention does anything.
The version in performance management
A specific application, because it is where the mechanism causes the most avoidable damage.
Rank staff, teams or branches on a noisy performance measure. Intervene on the bottom. Measure again. They improve, and the intervention is credited.
Now intervene on the top — reward them, promote them, hold them up as an example. Measure again. They decline, and the intervention is blamed.
The manager who ran both experiments has learned that criticism works and praise does not, from data in which neither did anything at all. Kahneman’s account of the flight instructors is exactly this, and it is the clearest case in the literature of a firmly held belief produced by a statistical artefact.
The defence is the one this essay keeps arriving at: do not evaluate on the measure that selected. Rank on one period and evaluate on the next, or use a measure aggregated over long enough that the noise component is small.
What separates this from a real effect
The obvious question — how would anyone tell — has a clean answer.
A real effect moves the selected group beyond what regression predicts. The prediction is (1 − ρ) times the group’s distance from the mean, and ρ can be estimated from the data itself by correlating two measurements in an untreated group.
So the null hypothesis is not “no change”. It is “the change predicted by the correlation alone”, and that is a specific number rather than zero.
Testing against zero is what produces the false positives, and it is what almost every before-and-after evaluation does. Testing against the regression prediction requires knowing ρ and is otherwise no harder.
That is the constructive version of the whole essay: the artefact is not merely a warning, it is a quantitative prediction, and predictions can be subtracted.
The size of the effect, in closed form
The measurement above is a simulation, and the whole quantity is available in closed form, which is worth doing because it turns a demonstration into a prediction.
For a pair of standardised measurements with correlation ρ, selecting on the first and looking at the second gives
That is the entire phenomenon. The second measurement is pulled toward zero by exactly the factor ρ, so the apparent improvement is (1 − ρ) times however extreme the selected group was on the first measurement.
Both pieces are computable. Selecting the bottom 20% of a standard normal gives a group whose mean is −1.400; at ρ = 0.6 their second measurement averages −0.840, an apparent improvement of 0.56 standard deviations from nothing at all. Selecting the bottom 10% gives a group at −1.755 and a follow-up at −1.053. Selecting the bottom half gives −0.798 and −0.479.
Two things follow that are hard to see from a single simulation.
The effect grows as the selection gets more extreme. Picking the worst decile produces a larger phantom improvement than picking the worst half, and picking the single worst case produces the largest of all. Interventions are targeted at the worst cases, which is exactly where the artefact is biggest.
The effect grows as the measurement gets noisier. Sweeping the correlation and predicting the shift before drawing anything:
- ρ = 0.2 predicts a shift of 1.12 standard deviations
- ρ = 0.4 predicts 0.84
- ρ = 0.6 predicts 0.56
- ρ = 0.8 predicts 0.28
- ρ = 0.95 predicts 0.07
and the simulation lands on each of them. So the phantom improvement is largest precisely where the measurement is least reliable, which inverts the intuition that a noisy measurement produces a small apparent effect. A noisy measurement produces a large apparent effect, reliably, in the direction of whatever was selected.
Why the prediction is the useful form
Having the closed form changes what can be done with the phenomenon, and it is the difference between a caution and a test.
A caution says that some of an observed improvement may be regression to the mean. That is unfalsifiable and easy to wave away, and it is how the effect is usually discussed — the same shape as a table that cannot answer the question it raises.
A prediction says: given the correlation between the two measurements and how the group was selected, the improvement attributable to regression alone is this number. Now the observed improvement can be compared with it. If a programme selecting the worst decile on a measurement with ρ = 0.6 reports an improvement of 0.7 standard deviations, the phantom accounts for 0.70 of it and there is nothing left to explain. If it reports 1.5, there is 0.8 to account for and the programme may be doing something.
That comparison needs one number nobody usually reports: the test–retest correlation of the outcome measure. It is cheap to estimate — it requires measuring the same units twice with no intervention — and its absence is what makes most before-and-after claims uninterpretable rather than merely uncertain.
The strongest version of the argument is therefore not that regression to the mean might explain a result. It is that without ρ nobody can say how much of the result it explains, and a design that failed to measure ρ has chosen not to be able to answer the first question its own numbers raise.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The slope of a density nobody can see — both name measurement error, regression to the mean, selection bias
- The estimate after the choice — both name regression to the mean, selection bias
Named objects
A flat tag is an object no other essay names yet.
CorrelationMeasurement errorRegression to the meanSelection biasStandard deviation