Regression to the mean
Measure something twice. Take the group that did worst the first time and look at how they did the second time. They improved.
They will improve whatever is done to them in between, including nothing.
The measurement
With two measurements correlated at 0.6, selecting the worst 15% on the first and looking at the second gives an average improvement of about 0.72 standard deviations.
The closed form says the shift should be (1 − ρ) times the group’s distance from the mean, which for this selection predicts about 0.69. The figure computes both and the assertion behind it requires them to agree, so a frame where the mechanism did not account for the movement would fail the build.
The best group moves the other way by a similar amount. Nothing was done to either group.
Why it happens
Any measurement is partly signal and partly noise. Selecting the extreme low group selects for both — people who are genuinely low, and people who had a bad measurement.
The genuine part persists to the second measurement. The noise part does not: a bad draw the first time says nothing about the second. So the selected group’s second measurement is closer to the mean, by exactly the fraction of their extremity that was noise.
That fraction is 1 − ρ, which is why the size of the effect is predictable in advance with no data about the intervention at all.
Move the correlation slider and the two limits are visible. At ρ = 0.95 the measurements are nearly the same thing and there is almost no movement. At ρ = 0 the first measurement carries no information about the second, and the selected group returns all the way to the population mean.
What it explains
The list is long, and its length is the reason this matters more than most statistical curiosities.
Interventions targeted at the worst cases. Remedial teaching for the lowest-scoring pupils, safety measures at the most dangerous junctions, management attention to the worst-performing branch. All of these select on an extreme measurement and all of them will show improvement whether or not they work. Evaluating them without a control group selected the same way is guaranteed to produce a positive result.
Punishment appearing to work better than praise. Kahneman’s flight instructor example: praise a pilot after an exceptionally good manoeuvre and the next is worse; criticise after a bad one and the next is better. Both are regression, and the instructor draws the opposite conclusion.
The sophomore slump, the Sports Illustrated cover jinx, the second album — the same mechanism wherever selection is on an exceptional performance.
Medical improvement after treatment begins. Patients typically seek treatment when symptoms are at their worst, which is a selection on an extreme measurement. Improvement follows regardless, which is a large part of what an uncontrolled trial measures and a substantial component of what gets attributed to placebo.
Its relation to the winner’s curse
These are the same mechanism seen from two ends, which is worth making explicit.
The winner’s curse selects studies on an extreme estimate and finds the estimates inflated. Regression to the mean selects individuals on an extreme measurement and finds the next measurement less extreme.
In both cases: selection on a noisy quantity, followed by observation of the underlying quantity, produces a systematic gap in a direction set by which end was selected.
Recognising the shared structure is useful, because the fix is shared too. Do not evaluate on the measurement that was selected on. Select on one measurement and evaluate on an independent one, or select at random.
What a control group actually does
This clarifies something usually taught as a rule without a reason.
A control group is not there to make the study look rigorous. It is there because the selected group would have moved anyway, and the only way to know how much of the observed movement is the intervention is to watch a group that was selected the same way and not treated.
That last clause matters and is often missed. A control group drawn from the general population does not control for regression to the mean, because it was not selected on an extreme measurement and so has nothing to regress from. The control has to share the selection, not merely the treatment’s absence.
Get that wrong and the study measures regression to the mean and reports it as an effect — which is, measurably, worth about 0.7 standard deviations at a correlation of 0.6. That is larger than most real effects anyone is looking for.
Galton, and where the name comes from
The phenomenon was named by Francis Galton in the 1880s, studying the heights of parents and their adult children. Tall parents had children shorter than themselves on average; short parents had taller ones. He called it “regression towards mediocrity” and the modern sense of the word regression descends from that paper.
The part worth knowing is that Galton initially read it as a physical tendency — a force pulling the population back toward its mean, which would need explaining. It is not. It is a consequence of imperfect correlation, present whenever two measurements of anything are correlated at less than one, and requiring no mechanism at all.
That misreading is still the default. An intervention appears to work; the appearance requires no cause; and looking for the cause is the error.
The symmetry that gives it away
The diagnostic is simple and rarely applied: the effect runs both ways.
If the worst group improves because of regression, the best group must decline by the same mechanism. The figure shows exactly that — the marked groups at each end move toward the centre, and the shifts are comparable in size.
So a study reporting that the worst performers improved after an intervention has a free control available in its own data. Did the best performers decline over the same period? If they did, the mechanism is regression and the intervention has not been shown to do anything. If they held steady or improved, something else is happening.
That check costs nothing and is almost never done, because the best performers are not the group anyone was interested in.
Where the correlation comes from
The size of the effect is (1 − ρ), so everything depends on what determines ρ, and it is worth separating the two contributions.
Measurement error reduces ρ directly. A noisy instrument produces low correlation between repeat measurements and therefore large regression, regardless of how stable the underlying quantity is. Improving the instrument reduces the effect.
Genuine variation over time also reduces ρ, and improving the instrument does nothing about it. A quantity that really does fluctuate — pain, blood pressure, monthly sales — will show regression however precisely each measurement is taken.
Distinguishing them matters for what to do. If the correlation is low because of measurement noise, averaging several measurements at each time point raises ρ and shrinks the effect. If it is low because the quantity genuinely moves, averaging over a longer window does the same thing for a different reason. Either way, more measurement per occasion is the practical defence, and it is cheaper than a control group.
The version that shows up in league tables
The most consequential application, and the one where the arithmetic is routinely ignored.
Rank hospitals, schools or regions by an outcome measured over one period. The bottom performers are partly genuinely worse and partly unlucky. Measure again the next period and the bottom group improves — whatever was done to them, including nothing.
Two failures follow. Interventions targeted at the bottom appear effective, and their apparent effect is the regression. And the ranking itself is unstable, because much of what determined it was noise, so the following year’s ranking looks different and the difference gets reported as change.
The fix is the same as everywhere else in this essay: select on one period and evaluate on another, or account for the measurement error explicitly. Neither is difficult and both are rarer than the league table.
The unifying statement
Three essays on this site describe the same mechanism.
The winner’s curse selects studies on an extreme estimate. Regression to the mean selects individuals on an extreme measurement. The forking paths select an analysis on an extreme p-value.
In every case: select on a noisy quantity, then observe something that shares only the signal, and the observation moves toward the mean by the fraction that was noise.
Stated that way it stops being three phenomena and becomes one, with three names because it was discovered separately in three literatures.
Two measurements are not always available
The recommendation — select on one measurement, evaluate on another — assumes two measurements exist. Often they do not, and it is worth saying what is available then.
Split the single measurement. Where the measure is an aggregate of many small observations, half can select and the other half evaluate. This is the same idea at a smaller scale and it works.
Model the measurement error explicitly. If the reliability of the measure is known from elsewhere, the expected regression can be computed and subtracted. That is what the (1 − ρ) formula is for, and it is why the figure asserts the prediction rather than merely showing the movement.
Or select at random. An intervention delivered to a randomly chosen group has no selection to regress from, and the comparison is clean. That is less satisfying because it does not target the people who appear to need it most — but targeting on a noisy measurement is what created the problem.
The last option is usually rejected on ethical grounds, which is a real argument and worth weighing against the fact that an untargeted trial is the only one that will say whether the intervention does anything.
The version in performance management
A specific application, because it is where the mechanism causes the most avoidable damage.
Rank staff, teams or branches on a noisy performance measure. Intervene on the bottom. Measure again. They improve, and the intervention is credited.
Now intervene on the top — reward them, promote them, hold them up as an example. Measure again. They decline, and the intervention is blamed.
The manager who ran both experiments has learned that criticism works and praise does not, from data in which neither did anything at all. Kahneman’s account of the flight instructors is exactly this, and it is the clearest case in the literature of a firmly held belief produced by a statistical artefact.
The defence is the one this essay keeps arriving at: do not evaluate on the measure that selected. Rank on one period and evaluate on the next, or use a measure aggregated over long enough that the noise component is small.
What separates this from a real effect
The obvious question — how would anyone tell — has a clean answer.
A real effect moves the selected group beyond what regression predicts. The prediction is (1 − ρ) times the group’s distance from the mean, and ρ can be estimated from the data itself by correlating two measurements in an untreated group.
So the null hypothesis is not “no change”. It is “the change predicted by the correlation alone”, and that is a specific number rather than zero.
Testing against zero is what produces the false positives, and it is what almost every before-and-after evaluation does. Testing against the regression prediction requires knowing ρ and is otherwise no harder.
That is the constructive version of the whole essay: the artefact is not merely a warning, it is a quantitative prediction, and predictions can be subtracted.