Concept

Regression to the mean — where it appears

That an extreme measurement tends to be followed by a less extreme one, because part of what made it extreme was noise. It is arithmetic rather than a force, and mistaking it for one is how an intervention applied to the worst cases gets credit for their improvement.

Named by 7 essays across 3 fields — each of them below, with the objects they name alongside it.

Two measurements of the same thing, correlated 0.60. Pick the worst 15% on the first measurement and their average rises by 0.78 on the second. Pick the best and theirs falls by 0.48. No treatment was given to anybody.

Regression to the mean

Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.

paradox · Rtm
How many of a league table's top ten are small groups, ranked four ways. A hundred groups with sizes from 4 to 400, of which 36% have twenty units or fewer. Small groups make up 36.3% of the true top ten, 62.1% of the top ten by raw means, 13.4% by posterior means and 22.9% by the posterior chance of being in the top ten. The three rankings recover 4.43, 5.38 and 5.47 of the true top ten.

A league table of a hundred

A hundred groups with sizes from 4 to 400, and a top ten to publish. Ranked by their own means, small groups fill 62.1% of the top ten against their 36.3% share of the true top ten. Ranked by posterior means they fill 13.4%. The ranking built from each group's chance of being in the top ten recovers 5.47 of the true ten, the best of three and barely half; and the group ranked first could hold any rank from 1 to 31.

borrowed · Shrinkage
Three estimates of the same effect, and the honest one is the worst. 8,000 trials in which every arm has an effect of exactly 0. The arm carried forward is reported by its first stage at +0.187 above the truth — it was chosen for being ahead — and by its second stage at -0.0032, which is unbiased by construction because the selection could not see it. The combination that gets published is at 0.123, exactly the share of the first stage's bias the first stage contributes. Root mean squared error: 0.240, 0.261, 0.182 — the unbiased estimate is the least accurate of the three.

The estimate after the choice

An arm chosen for being ahead is ahead by more than it should be, and the trial then publishes the average of the stage that chose it and the stage that did not. The unbiased estimate is the one built from a third of the data — and it is the least accurate of the three.

adaptive · Curse
Two groups, a baseline and a follow-up, and nothing happening in between — baseline reliability 0.6. 600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.075 and −0.032, so the change-score analysis reports a group difference of 0.043. The regression of follow-up on baseline and group reports 0.409, against a closed form of (1 − λ) × 1.00 = 0.400: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 0.614, the baseline's reliability.

Two analyses of one baseline

Two groups read at baseline and again at follow-up, with no change for anybody. Subtracting the baseline reports a group difference of −0.0014 and adjusting for it reports 0.4008 — and each analysis is exactly right about one reason the groups started apart and wrong by 0.40 about the other.

paradox · Rtm
A cohort screened once, the top tenth enrolled, and followed up with nothing given — correlation 0.6. 2000 people read once at screening and once at follow-up, with a test–retest correlation of 0.6 and no treatment. The 215 above a cut at the top ten per cent of one reading (1.282 standard deviations) are enrolled. Their mean screening reading is 1.744 and their mean follow-up reading 0.982, a fall of 0.762 ± 0.055 with nothing done to anyone. The closed form for the fall is (1 − ρ) times the truncated-normal mean, (1 − 0.6) × 1.755 = 0.702.

The measurement that got them enrolled

Enrol the top tenth of one screening reading and give them nothing, and they fall by 0.702 standard deviations at follow-up. Measured from a fresh reading taken after enrolment they fall by nothing. Averaging ten screening readings still leaves 0.101, and it takes twenty-one to get under 0.05.

paradox · Rtm
The top of a heavy-tailed population keeps its lead; the top of a light-tailed one gives it back. Select the top share on the first reading and read the group again: the share of its mean lead the second reading keeps, by integration over the true score (lines) and counted on 400,000 draws a parent in 20 batches (points, with two standard errors). The normal keeps exactly 0.6 at every selection. At the top half the Laplace keeps 0.541, the t 0.535 and the uniform 0.648 — the heavy tails keep LESS than the correlation. By the top one per cent the order has reversed: 0.761, 0.784 and 0.443. At one in ten thousand the t keeps 0.977 and the uniform 0.346.

A lead that a heavy tail keeps

Four populations whose readings all correlate at exactly 0.6, and whose least-squares slopes all read 0.6. Select the top one per cent on one reading and measure them again: they keep 60% of their lead if the true scores are normal, 76.1% if they are Laplace, 78.4% if they are a t on four degrees of freedom — and 44.3% if they are uniform. The correlation predicts the regression of the extremes for one shape of population only.

paradox · Rtm
Twenty studies of a thousand readings estimate the regression of the extremes: the t, four degrees parent. Each thin line is one study of 1000 readings from the t, four degrees parent: Tweedie's formula with the log-density's slope estimated by a degree-5 log-spline, drawn up to that study's largest reading. The thick line is the share kept by integration over the true score, and the dashed line the correlation, 0.6. The top ten readings of the median study begin at 2.45. Over 400 studies the corrected share kept by the top one per cent averages 0.8190, with a spread of 0.0806, against 0.7842 by integration; a Gaussian kernel averages 0.7914 with a spread of 0.0853.

The slope of a density nobody can see

Tweedie's formula corrects a reading by the slope of the readings' own log-density, and a study has its readings. Estimated from a thousand of them, the correction for the top one per cent beats the correlation's linear rule on 84.0% to 98.0% of studies from heavy-tailed populations and on 75.0% to 81.5% from a bounded one — and costs an error of 0.09 to 0.12 where the population is normal and the rule was already exact. At 250 readings the log-spline loses to the rule it replaces, and at 16,000 the same log-spline gets worse on a power tail.

paradox · Rtm

Named alongside it

The objects these essays reach for when they reach for this one.

Measurement errorSelection biasClosed formMonte CarloTest–retest reliabilityShrinkageCorrelationCovariate adjustmentHeavy tailPosterior meanTweedie's formulaAdaptive design

All concepts