Two analyses of one baseline
Worth reading first: Regression to the mean.
In 1967 Frederic Lord described a university dining hall and two statisticians. Students were weighed in September and again in June, boys and girls separately, and for both groups the distribution of weight at the end of the year was the same as at the start. The first statistician compared the mean gain in each group, found it zero in both, and concluded that the food had affected the two groups no differently. The second fitted a regression of June weight on September weight with a term for group, found that a boy gained reliably more than a girl of the same starting weight, and concluded the opposite.
Both analyses were done correctly. They were applied to one dataset and they disagree about it, and the disagreement does not shrink with more students. It has been called a paradox ever since, and the name has outlived the confusion, because what separates the two statisticians is arithmetic that can be written down.
That arithmetic is the one the essay on regression to the mean measured: a reading is partly signal and partly noise, and a second reading gives back the noise. What that essay did not need to ask is back towards what. With one population there is only one mean to return to. With two groups whose true means differ there are two, and a unit’s second reading returns towards its own group’s mean rather than the mean of everybody. Lord’s paradox is that sentence and nothing else.
A dataset in which nothing happens
Build the dining hall with no food in it. Two groups whose true means are one standard deviation apart; each unit’s true score fixed for the year; a baseline reading and a follow-up reading, each the true score plus independent error. The within-group variance of the true score is 0.6 and each reading adds error of variance 0.4, so either reading has a reliability of 0.6 inside a group — the share of its variance that is the unit rather than the occasion. Nobody changes. Whatever an analysis reports as a difference between the groups, the truth is zero.
On one draw of six hundred units, the two groups’ mean changes are −0.075 and −0.032. Their difference — the change-score analysis — is 0.043, which is noise of the size six hundred units should produce. The regression of follow-up on baseline and group reports a group coefficient of 0.409, with a pooled within-group slope of 0.614. The first analysis says there is nothing, and it is right. The second says there is something, roughly six of its own standard errors from zero, and nothing happened.
The picture shows where the second number comes from. Each cloud is centred on the diagonal, because on average nobody moved. But each cloud’s regression line is flatter than the diagonal, because the slope of follow-up on baseline is the baseline’s reliability rather than one, and a flattened line through a cloud centred at a different point crosses the diagonal at a different place. Stand at one baseline reading and look up: the two lines are not at the same height. That vertical gap is exactly what an analysis of covariance estimates as the effect of group.
Two lines with one slope and two crossings
The gap has a closed form, and it needs one fact about normal readings. If a true score has mean in group and the baseline reading is with reliability inside the group, then
A reading is shrunk towards its group’s mean by the share of its deviation that was error. Because the follow-up reading is the same true score plus fresh error, the expected follow-up given the baseline is this same line. Two groups give two lines of slope , one crossing the diagonal at each group’s mean, and the vertical distance between them is
at every baseline reading. With a reliability of 0.6 and a gap of one standard deviation that is 0.4, and the single draw above read 0.409.
Follow one reading. A unit in the first group read at 1.5 sits two standard deviations above its group’s mean of −0.5, and its second reading is expected to give back 40% of that, landing at 0.7. A unit in the second group read at the same 1.5 sits only one standard deviation above its group’s mean of 0.5, gives back 40% of that, and lands at 1.1. The two units had identical baselines, received identical treatment — none — and are expected to differ by 0.4 at follow-up, because the same reading means something different about a unit depending on which population it was drawn from.
That is the first statistician’s answer and the second’s in one picture. Averaged over each group, every unit’s regression is cancelled by another unit’s regression the other way, and the mean change is zero. Held at a fixed baseline, every unit in the second group is regressing towards a higher mean than every unit in the first, and the conditional difference is 0.4. Neither is an error of computation. They are answers to different questions.
Which question each number answers
The change score answers: did the groups change by different amounts, on average? In the dining hall, no.
The adjusted coefficient answers: among units that read the same at baseline, does group membership predict the follow-up? In the dining hall, yes. A boy and a girl who weigh the same in September are not in the same position: the boy is below the boys’ mean and the girl above the girls’, each September reading is partly that day’s fluctuation, and in June each is expected to move back towards the mean of their own group — the boy up and the girl down.
What neither number answers by itself is the question Lord’s statisticians were asked, which was causal: did the food affect one group differently from the other? To get from either number to that question needs a statement about what would have happened without the food, and the two analyses carry two different statements. Holland and Rubin’s 1983 reading of the paradox put it that way: the data are compatible with both analyses, and the choice between them is a choice of untestable assumption rather than of method. It is the point three causal structures fitted to one covariance matrix makes about arrows, arriving here as a point about baselines.
That reading is correct and it is easy to leave abstract. It can be made concrete by building the two assumptions as two worlds and counting both analyses in each.
Two reasons for the same gap
In the first world the groups are pre-existing populations — sexes, schools, wards — and their true means differ by one standard deviation. That is the dining hall.
In the second world there is one population, and the group a unit ends up in is chosen by its baseline reading: a unit reading high is more likely to be put in the second group, through a probit in the reading whose steepness, 0.8041, is solved so that the baseline gap between the groups is again exactly one standard deviation. Nothing but the measurement decides. That is a clinic that refers its high readers, a school that streams on a test, a programme that enrols on a score.
Both worlds have the same baseline gap and the same absence of any change. Counted over 2,000 datasets of 200 units each at a reliability of 0.6:
- Pre-existing groups. The change score reports −0.0014 ± 0.0028. ANCOVA reports 0.4008 ± 0.0028.
- Groups chosen by the reading. ANCOVA reports −0.0031 ± 0.0029. The change score reports −0.4035 ± 0.0028.
Each analysis is right in exactly one world and wrong by the same amount, 0.4, in the other. The change score is wrong in the second world for the reason it is right in the first: it treats the baseline gap as a permanent fact about the groups. When the gap was produced by selecting on noisy readings, the second group’s high readings were partly luck, the luck does not repeat at follow-up, and the second group appears to fall towards the population mean by of its lead — a change score of −0.4, reported as a group difference, from a world where every unit stayed where it was. That is plain regression to the mean, measured by an analysis that did not account for it.
ANCOVA is right in the second world because the group was decided by the reading and by nothing else, so once the reading is held fixed the group carries no further information about the follow-up. It is wrong in the first world because there the group was decided by the true score, the reading is only a noisy proxy for the true score, and holding a noisy proxy fixed does not hold the true score fixed.
The two worlds are one construction each, and they are not claimed to exhaust what a real gap can come from. A real comparison can be any mixture — groups that differ in truth and were partly sorted on a measurement — and then neither analysis is right and the truth lies between them at a place fixed by the mixture. The two worlds are also not literally indistinguishable: selecting on a probit bends the within-group distribution of baselines slightly away from normal, and a large sample could see that. What no sample can see is the general question, which is why the groups differ, and a dataset carries no column for it.
The disagreement is the unreliability of the baseline
The size of every wrong answer in that table is times the gap, and is the baseline’s reliability. So the paradox should vanish when the baseline is measured without error, and it does.
Swept across baseline reliabilities of 0.2, 0.4, 0.6, 0.8 and 1, ANCOVA on pre-existing groups reports 0.803, 0.596, 0.401, 0.197 and 0.002. The change score on groups chosen by the reading reports −0.793, −0.594, −0.403, −0.205 and 0.000. The other two cells sit at zero throughout. Across all thirty counted cells — three analyses, two worlds, five reliabilities — the worst disagreement between a count and its closed form is 2.34 standard errors, which is about what thirty honest comparisons produce.
Read ANCOVA’s error in the first world as what it is. The analysis is trying to adjust for the true score, which is the real difference between the groups. It is given a reading of the true score with reliability , and a regression on a mismeasured covariate estimates an attenuated slope — rather than one — and so removes only the share of the confounding the covariate carries. The share is left in the group coefficient and reported as an effect. This is the ordinary result about adjusting for a confounder measured with error, and it is the same result as adjusting for a covariate that does not mean what the regression assumes it means: the arithmetic of adjustment is one operation, and whether it removes the problem depends on facts about the covariate the arithmetic cannot check.
With a perfect baseline the reading is the true score, the within-group slope is one, both groups’ lines lie on the diagonal, and subtracting the baseline and adjusting for it become the same operation. There is nothing left to regress. The paradox, stated precisely, is the unreliability of the baseline, multiplied by the gap between the groups.
“Unreliability” has to be read broadly for that sentence to be true of the dining hall. A scale can be precise to the gram and a September weighing still not be a student’s weight for the year, because weight moves from week to week. What enters is everything about the baseline that does not persist to the follow-up — instrument error and genuine occasion-to-occasion fluctuation alike — which is the same split the essay on regression to the mean drew between a noisy instrument and a quantity that really moves. A precise instrument on a fluctuating quantity leaves the paradox exactly where it was.
When the groups were randomised
Take the reason for the gap away and the paradox goes with it. With the groups randomised the true means are equal, , every term of the form times a gap is zero, and all three analyses — ignore the baseline, subtract it, adjust for it — estimate the same effect without bias. What is left to choose between them is precision, and that has a closed form too.
Standardise both readings to variance one with test–retest correlation , and put units in each arm. Then
The last factor is exact rather than asymptotic: it is what estimating the slope costs, from . Relative to ignoring the baseline, the change score costs and ANCOVA .
Counted over 4,000 randomised trials of a hundred units at each of seven correlations, at the change score’s variance is 0.773 of ignoring the baseline against a closed form of 0.800, and ANCOVA’s is 0.642 against 0.647. At they are 0.199 and 0.192 — almost the same, because at high correlation the fitted slope is nearly one and adjustment nearly is subtraction.
The change score has a break-even that is worth knowing, because the intuition runs the other way. Below a correlation of one half, subtracting the baseline makes a randomised comparison less precise than ignoring it. At its variance is 1.183 of the unadjusted comparison’s, and at it is 2.093, against a closed form of exactly two: subtracting an unrelated reading adds its whole variance and removes nothing. ANCOVA cannot do that, because it estimates how much of the baseline to subtract rather than subtracting all of it, and its only cost is the charge for that estimate — at it reads 1.009 against a closed form of 1.010, one hundred units paying a single degree of freedom.
This is the same bargain as arranging units in pairs before any outcome exists: information about the units, used in the design or in the analysis, is paid for with a degree of freedom and repaid with the variance it explains. Under randomisation the choice is only about that repayment, and adjusting is the better buy at every correlation. It is outside randomisation that the choice stops being about precision and starts being about which world the data came from, and there no amount of precision helps.
What a real effect looks like through each analysis
The two worlds above contain no effect, which made the errors visible against a truth of zero. Put an effect in, and the errors do not go away; they are added to it.
Give the second group a real change of 0.4 at follow-up. Among pre-existing groups the change score reports the 0.4 and ANCOVA reports 0.800 — the effect and the paradox, indistinguishable. Among groups chosen by the reading ANCOVA reports the 0.4 and the change score reports 0.000: a real effect exactly cancelled by regression to the mean, and a programme that worked declared to have done nothing. The size of the effect was picked to make that cancellation visible, and nothing about it is rare. An effect smaller in size than times the gap, pointing the other way from the wrong analysis’s error, comes out with its sign reversed. At test–retest correlations between 0.5 and 0.8, a range commonly reported for outcomes that fluctuate from one occasion to the next, that threshold is between a fifth and a half of whatever gap separated the groups at the start.
That is why the question of which analysis to trust cannot be put off until the result is in. The two analyses were designed for two assumptions about why the groups differ, and a comparison whose groups were not randomised has to say which assumption it is making before the numbers exist — for the same reason a table that reverses on aggregation cannot be read until somebody says which weighting the question calls for.
What the two worlds assume, and where the arithmetic stops
The true scores are normal inside each group. That is what makes a straight line with slope . For any other distribution of true scores the conditional mean is curved — a heavy-tailed population keeps more of an extreme reading’s lead than the correlation predicts — and then ANCOVA’s fitted line is only the best straight approximation to two curves. The sign of the paradox survives; its closed form does not.
The two groups share one reliability. A baseline measured more precisely in one group than the other gives the two lines different slopes, the gap between them stops being constant, and “the” adjusted group difference depends on where it is read.
Nobody changes. The dining hall was built so that the truth is zero and every non-zero answer is an artefact. A world in which units genuinely grow, and grow by amounts related to where they started, adds a second regression — a real one — on top of the measurement one, and separating the two needs more than two readings.
The groups are large. Nothing here is a small-sample effect. Every number is a limit that 200 units approach and 600 units approach more closely, which is the uncomfortable part: a wrong analysis on a large comparison is wrong with a small standard error, and a precise answer to a question nobody asked looks exactly like a precise answer to the one that was.
What is proved here is the algebra of the two lines and the three variances; what is measured is that 2,000 datasets a cell land on that algebra. What is not settled by either is which world a particular comparison came from, and that is the part of Lord’s paradox that is not arithmetic at all.
Where this goes next
The second world is the one worth following. There the group was chosen by the baseline reading, and the damage came from using that same reading as the baseline: the change score measured the selection’s luck wearing off. The obvious repair is to stop using the reading that selected as the reading that measures, and it raises a question this essay did not need to ask — whether taking a second baseline after the selection removes the artefact completely, or whether averaging several readings at selection does as well. The two sound equivalent and are not. The next question on regression to the mean is enrolment on a threshold, where the selection is explicit, the untreated arm improves by a closed-form amount, and one fresh reading turns out to do what no number of averaged screening readings can.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Adjusting for a shadow — both name closed form, covariate adjustment, measurement error, monte carlo, test–retest reliability
- The slope of a density nobody can see — both name closed form, measurement error, monte carlo, regression to the mean, test–retest reliability
- A basis is a subspace — both name assignment mechanism, monte carlo, variance reduction
- A count that has to be estimated — both name closed form, monte carlo, randomisation
- A covariate with no levels — both name closed form, monte carlo, randomisation
- A rate times a size — both name closed form, monte carlo, variance reduction
Named objects
A flat tag is an object no other essay names yet.
Assignment mechanismAttenuationChange scoreClosed formConfoundingCovariate adjustmentLord's paradoxMeasurement errorMonte CarloPrecisionRandomisationRegression to the meanTest–retest reliabilityVariance reduction