Reversals that are not errors

Two analyses of one baseline

Two groups read at baseline and again at follow-up, with no change for anybody. Subtracting the baseline reports a group difference of −0.0014 and adjusting for it reports 0.4008 — and each analysis is exactly right about one reason the groups started apart and wrong by 0.40 about the other.

Worth reading first: Regression to the mean.

In 1967 Frederic Lord described a university dining hall and two statisticians. Students were weighed in September and again in June, boys and girls separately, and for both groups the distribution of weight at the end of the year was the same as at the start. The first statistician compared the mean gain in each group, found it zero in both, and concluded that the food had affected the two groups no differently. The second fitted a regression of June weight on September weight with a term for group, found that a boy gained reliably more than a girl of the same starting weight, and concluded the opposite.

Both analyses were done correctly. They were applied to one dataset and they disagree about it, and the disagreement does not shrink with more students. It has been called a paradox ever since, and the name has outlived the confusion, because what separates the two statisticians is arithmetic that can be written down.

That arithmetic is the one the essay on regression to the mean measured: a reading is partly signal and partly noise, and a second reading gives back the noise. What that essay did not need to ask is back towards what. With one population there is only one mean to return to. With two groups whose true means differ there are two, and a unit’s second reading returns towards its own group’s mean rather than the mean of everybody. Lord’s paradox is that sentence and nothing else.

A dataset in which nothing happens

Build the dining hall with no food in it. Two groups whose true means are one standard deviation apart; each unit’s true score fixed for the year; a baseline reading and a follow-up reading, each the true score plus independent error. The within-group variance of the true score is 0.6 and each reading adds error of variance 0.4, so either reading has a reliability of 0.6 inside a group — the share of its variance that is the unit rather than the occasion. Nobody changes. Whatever an analysis reports as a difference between the groups, the truth is zero.

On one draw of six hundred units, the two groups’ mean changes are −0.075 and −0.032. Their difference — the change-score analysis — is 0.043, which is noise of the size six hundred units should produce. The regression of follow-up on baseline and group reports a group coefficient of 0.409, with a pooled within-group slope of 0.614. The first analysis says there is nothing, and it is right. The second says there is something, roughly six of its own standard errors from zero, and nothing happened.

Two groups, a baseline and a follow-up, and nothing happening in between — baseline reliability 0.6600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.075 and −0.032, so the change-score analysis reports a group difference of 0.043. The regression of follow-up on baseline and group reports 0.409, against a closed form of (1 − λ) × 1.00 = 0.400: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 0.614, the baseline's reliability.-3-2-10123-3-2-10123baseline readingfollow-up readinggroup 1mean change −0.07group 2mean change −0.03at one baseline readingthe lines are 0.40 apartchange scoresays 0.04ANCOVAsays 0.41600 units, true group means 1.00 apart, no changeone dataset, two answers
Fig. 1 Six hundred units in two groups whose true means are one standard deviation apart, each read at baseline and at follow-up with nothing changing. The diagonal is “no change”. Each group’s regression line is flatter than the diagonal and crosses it at that group’s own mean, so at any single baseline reading the two lines sit apart. The slider sets the baseline’s reliability.

The picture shows where the second number comes from. Each cloud is centred on the diagonal, because on average nobody moved. But each cloud’s regression line is flatter than the diagonal, because the slope of follow-up on baseline is the baseline’s reliability rather than one, and a flattened line through a cloud centred at a different point crosses the diagonal at a different place. Stand at one baseline reading and look up: the two lines are not at the same height. That vertical gap is exactly what an analysis of covariance estimates as the effect of group.

Two lines with one slope and two crossings

The gap has a closed form, and it needs one fact about normal readings. If a true score TT has mean μg\mu_g in group gg and the baseline reading is X=T+eX = T + e with reliability λ\lambda inside the group, then

E[TX=x, g]=(1λ)μg+λx.E[\,T \mid X = x,\ g\,] = (1 - \lambda)\,\mu_g + \lambda\, x .

A reading is shrunk towards its group’s mean by the share of its deviation that was error. Because the follow-up reading is the same true score plus fresh error, the expected follow-up given the baseline is this same line. Two groups give two lines of slope λ\lambda, one crossing the diagonal at each group’s mean, and the vertical distance between them is

(1λ)(μ2μ1)(1 - \lambda)(\mu_2 - \mu_1)

at every baseline reading. With a reliability of 0.6 and a gap of one standard deviation that is 0.4, and the single draw above read 0.409.

One baseline reading, followed towards two different means. The expected follow-up reading given the baseline, for each of two groups whose true means are −0.50 and 0.50, at a baseline reliability of 0.6. Both lines have slope 0.6; they cross the diagonal at their own group's mean. A unit read at 1.50 in the first group is expected back at 0.700, a fall of 0.800, which is 40% of its distance from −0.50. The same reading in the second group is expected back at 1.100, a fall of 0.400, 40% of its distance from 0.50. Nobody changed; the 0.400 between the two is what an analysis that holds the baseline fixed reports as a group difference.
Fig. 2 The two expected-follow-up lines alone. A unit read at 1.5 in the first group is expected back at 0.7, having given up 40% of its distance from that group’s mean of −0.5; the same reading in the second group is expected back at 1.1, having given up 40% of its distance from 0.5. The 0.4 between them is a difference in where each unit is regressing to.

Follow one reading. A unit in the first group read at 1.5 sits two standard deviations above its group’s mean of −0.5, and its second reading is expected to give back 40% of that, landing at 0.7. A unit in the second group read at the same 1.5 sits only one standard deviation above its group’s mean of 0.5, gives back 40% of that, and lands at 1.1. The two units had identical baselines, received identical treatment — none — and are expected to differ by 0.4 at follow-up, because the same reading means something different about a unit depending on which population it was drawn from.

That is the first statistician’s answer and the second’s in one picture. Averaged over each group, every unit’s regression is cancelled by another unit’s regression the other way, and the mean change is zero. Held at a fixed baseline, every unit in the second group is regressing towards a higher mean than every unit in the first, and the conditional difference is 0.4. Neither is an error of computation. They are answers to different questions.

Which question each number answers

The change score answers: did the groups change by different amounts, on average? In the dining hall, no.

The adjusted coefficient answers: among units that read the same at baseline, does group membership predict the follow-up? In the dining hall, yes. A boy and a girl who weigh the same in September are not in the same position: the boy is below the boys’ mean and the girl above the girls’, each September reading is partly that day’s fluctuation, and in June each is expected to move back towards the mean of their own group — the boy up and the girl down.

What neither number answers by itself is the question Lord’s statisticians were asked, which was causal: did the food affect one group differently from the other? To get from either number to that question needs a statement about what would have happened without the food, and the two analyses carry two different statements. Holland and Rubin’s 1983 reading of the paradox put it that way: the data are compatible with both analyses, and the choice between them is a choice of untestable assumption rather than of method. It is the point three causal structures fitted to one covariance matrix makes about arrows, arriving here as a point about baselines.

That reading is correct and it is easy to leave abstract. It can be made concrete by building the two assumptions as two worlds and counting both analyses in each.

Two reasons for the same gap

In the first world the groups are pre-existing populations — sexes, schools, wards — and their true means differ by one standard deviation. That is the dining hall.

In the second world there is one population, and the group a unit ends up in is chosen by its baseline reading: a unit reading high is more likely to be put in the second group, through a probit in the reading whose steepness, 0.8041, is solved so that the baseline gap between the groups is again exactly one standard deviation. Nothing but the measurement decides. That is a clinic that refers its high readers, a school that streams on a test, a programme that enrols on a score.

Both worlds have the same baseline gap and the same absence of any change. Counted over 2,000 datasets of 200 units each at a reliability of 0.6:

  • Pre-existing groups. The change score reports −0.0014 ± 0.0028. ANCOVA reports 0.4008 ± 0.0028.
  • Groups chosen by the reading. ANCOVA reports −0.0031 ± 0.0029. The change score reports −0.4035 ± 0.0028.
Each analysis is right about one reason for the gap, and wrong about the other. Both analyses counted over 2,000 datasets of 200 units in each of two worlds with a baseline gap of 1.00 and nothing changing, at a baseline reliability of 0.6. Where the groups are pre-existing populations the change score reports −0.0014 ± 0.0028 and ANCOVA 0.4008 ± 0.0028. Where the baseline reading itself decided who went into which group, ANCOVA reports −0.0031 ± 0.0029 and the change score −0.4035 ± 0.0028. The wrong analysis is off by (1 − λ) × gap = 0.400 in each world, in opposite directions.
Fig. 3 Both analyses counted in both worlds, where nothing changes and the right answer is zero. The change score is right about pre-existing groups and ANCOVA about groups the reading chose; each is off by 0.4 in the other world, in opposite directions.

Each analysis is right in exactly one world and wrong by the same amount, 0.4, in the other. The change score is wrong in the second world for the reason it is right in the first: it treats the baseline gap as a permanent fact about the groups. When the gap was produced by selecting on noisy readings, the second group’s high readings were partly luck, the luck does not repeat at follow-up, and the second group appears to fall towards the population mean by (1λ)(1-\lambda) of its lead — a change score of −0.4, reported as a group difference, from a world where every unit stayed where it was. That is plain regression to the mean, measured by an analysis that did not account for it.

ANCOVA is right in the second world because the group was decided by the reading and by nothing else, so once the reading is held fixed the group carries no further information about the follow-up. It is wrong in the first world because there the group was decided by the true score, the reading is only a noisy proxy for the true score, and holding a noisy proxy fixed does not hold the true score fixed.

The two worlds are one construction each, and they are not claimed to exhaust what a real gap can come from. A real comparison can be any mixture — groups that differ in truth and were partly sorted on a measurement — and then neither analysis is right and the truth lies between them at a place fixed by the mixture. The two worlds are also not literally indistinguishable: selecting on a probit bends the within-group distribution of baselines slightly away from normal, and a large sample could see that. What no sample can see is the general question, which is why the groups differ, and a dataset carries no column for it.

The disagreement is the unreliability of the baseline

The size of every wrong answer in that table is (1λ)(1-\lambda) times the gap, and λ\lambda is the baseline’s reliability. So the paradox should vanish when the baseline is measured without error, and it does.

Swept across baseline reliabilities of 0.2, 0.4, 0.6, 0.8 and 1, ANCOVA on pre-existing groups reports 0.803, 0.596, 0.401, 0.197 and 0.002. The change score on groups chosen by the reading reports −0.793, −0.594, −0.403, −0.205 and 0.000. The other two cells sit at zero throughout. Across all thirty counted cells — three analyses, two worlds, five reliabilities — the worst disagreement between a count and its closed form is 2.34 standard errors, which is about what thirty honest comparisons produce.

The disagreement is the unreliability of the baseline. Both analyses in both worlds at baseline reliabilities of 0.2, 0.4, 0.6, 0.8, 1, counted over 2,000 datasets a point and drawn against their closed forms. ANCOVA on pre-existing groups reports (1 − λ) × gap: 0.803, 0.596, 0.401, 0.197, 0.002. The change score on groups chosen by the reading reports −(1 − λ) × gap: −0.793, −0.594, −0.403, −0.205, 0.000. The other two cells sit at zero at every reliability. At a reliability of 1 — a baseline read without error — all four agree, because a unit's reading is then its true score and there is nothing to regress. The worst disagreement between count and closed form anywhere on the figure is 2.34 standard errors.
Fig. 4 The two wrong analyses against the baseline’s reliability, with the two right ones at zero. The lines are ±(1 − λ) times the gap and the points are counted. At a reliability of 1 all four agree.

Read ANCOVA’s error in the first world as what it is. The analysis is trying to adjust for the true score, which is the real difference between the groups. It is given a reading of the true score with reliability λ\lambda, and a regression on a mismeasured covariate estimates an attenuated slope — λ\lambda rather than one — and so removes only the share λ\lambda of the confounding the covariate carries. The share 1λ1-\lambda is left in the group coefficient and reported as an effect. This is the ordinary result about adjusting for a confounder measured with error, and it is the same result as adjusting for a covariate that does not mean what the regression assumes it means: the arithmetic of adjustment is one operation, and whether it removes the problem depends on facts about the covariate the arithmetic cannot check.

Two groups, a baseline and a follow-up, and nothing happening in between — baseline reliability 1. 600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.020 and −0.022, so the change-score analysis reports a group difference of −0.003. The regression of follow-up on baseline and group reports −0.038, against a closed form of (1 − λ) × 1.00 = 0.000: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 1.035, the baseline's reliability.
Fig. 5 The same two groups with the baseline read without error. Each group’s regression line now has slope one and lies on the diagonal, the two lines coincide, and the change score and the adjusted difference are both zero within their noise.

With a perfect baseline the reading is the true score, the within-group slope is one, both groups’ lines lie on the diagonal, and subtracting the baseline and adjusting for it become the same operation. There is nothing left to regress. The paradox, stated precisely, is the unreliability of the baseline, multiplied by the gap between the groups.

“Unreliability” has to be read broadly for that sentence to be true of the dining hall. A scale can be precise to the gram and a September weighing still not be a student’s weight for the year, because weight moves from week to week. What enters λ\lambda is everything about the baseline that does not persist to the follow-up — instrument error and genuine occasion-to-occasion fluctuation alike — which is the same split the essay on regression to the mean drew between a noisy instrument and a quantity that really moves. A precise instrument on a fluctuating quantity leaves the paradox exactly where it was.

When the groups were randomised

Take the reason for the gap away and the paradox goes with it. With the groups randomised the true means are equal, μ2μ1=0\mu_2 - \mu_1 = 0, every term of the form (1λ)(1-\lambda) times a gap is zero, and all three analyses — ignore the baseline, subtract it, adjust for it — estimate the same effect without bias. What is left to choose between them is precision, and that has a closed form too.

Standardise both readings to variance one with test–retest correlation ρ\rho, and put n/2n/2 units in each arm. Then

var(post)=4n,var(change)=4n2(1ρ),var(ANCOVA)=4n(1ρ2)n3n4.\operatorname{var}(\text{post}) = \frac{4}{n},\qquad \operatorname{var}(\text{change}) = \frac{4}{n}\cdot 2(1-\rho),\qquad \operatorname{var}(\text{ANCOVA}) = \frac{4}{n}\,(1-\rho^2)\,\frac{n-3}{n-4}.

The last factor is exact rather than asymptotic: it is what estimating the slope costs, from E[1/χn22]=1/(n4)E[1/\chi^2_{n-2}] = 1/(n-4). Relative to ignoring the baseline, the change score costs 2(1ρ)2(1-\rho) and ANCOVA (1ρ2)(n3)/(n4)(1-\rho^2)(n-3)/(n-4).

Counted over 4,000 randomised trials of a hundred units at each of seven correlations, at ρ=0.6\rho = 0.6 the change score’s variance is 0.773 of ignoring the baseline against a closed form of 0.800, and ANCOVA’s is 0.642 against 0.647. At ρ=0.9\rho = 0.9 they are 0.199 and 0.192 — almost the same, because at high correlation the fitted slope is nearly one and adjustment nearly is subtraction.

What each analysis costs when the groups were randomised. With the groups randomised all three analyses are unbiased and differ only in variance. Relative to ignoring the baseline, the change score's variance is 2(1 − ρ) and ANCOVA's is (1 − ρ²)(n − 3)/(n − 4) at n = 100; the lines are those closed forms and the points are counted over 4,000 trials each. At ρ = 0.6 the change score reads 0.773 against 0.800 and ANCOVA 0.642 against 0.647. The change score is worse than ignoring the baseline below a correlation of one half and better above it; ANCOVA is at or below both everywhere except for its own slope charge at ρ = 0, where it reads 1.009.
Fig. 6 Variance of the change score and of ANCOVA, as a multiple of ignoring the baseline, when the groups were randomised. Lines are the closed forms and points are counted. The change score is worse than ignoring the baseline below a correlation of one half; ANCOVA is never worse than either, except by its slope charge at zero.

The change score has a break-even that is worth knowing, because the intuition runs the other way. Below a correlation of one half, subtracting the baseline makes a randomised comparison less precise than ignoring it. At ρ=0.4\rho = 0.4 its variance is 1.183 of the unadjusted comparison’s, and at ρ=0\rho = 0 it is 2.093, against a closed form of exactly two: subtracting an unrelated reading adds its whole variance and removes nothing. ANCOVA cannot do that, because it estimates how much of the baseline to subtract rather than subtracting all of it, and its only cost is the charge for that estimate — at ρ=0\rho = 0 it reads 1.009 against a closed form of 1.010, one hundred units paying a single degree of freedom.

This is the same bargain as arranging units in pairs before any outcome exists: information about the units, used in the design or in the analysis, is paid for with a degree of freedom and repaid with the variance it explains. Under randomisation the choice is only about that repayment, and adjusting is the better buy at every correlation. It is outside randomisation that the choice stops being about precision and starts being about which world the data came from, and there no amount of precision helps.

What a real effect looks like through each analysis

The two worlds above contain no effect, which made the errors visible against a truth of zero. Put an effect in, and the errors do not go away; they are added to it.

Give the second group a real change of 0.4 at follow-up. Among pre-existing groups the change score reports the 0.4 and ANCOVA reports 0.800 — the effect and the paradox, indistinguishable. Among groups chosen by the reading ANCOVA reports the 0.4 and the change score reports 0.000: a real effect exactly cancelled by regression to the mean, and a programme that worked declared to have done nothing. The size of the effect was picked to make that cancellation visible, and nothing about it is rare. An effect smaller in size than (1λ)(1-\lambda) times the gap, pointing the other way from the wrong analysis’s error, comes out with its sign reversed. At test–retest correlations between 0.5 and 0.8, a range commonly reported for outcomes that fluctuate from one occasion to the next, that threshold is between a fifth and a half of whatever gap separated the groups at the start.

That is why the question of which analysis to trust cannot be put off until the result is in. The two analyses were designed for two assumptions about why the groups differ, and a comparison whose groups were not randomised has to say which assumption it is making before the numbers exist — for the same reason a table that reverses on aggregation cannot be read until somebody says which weighting the question calls for.

What the two worlds assume, and where the arithmetic stops

The true scores are normal inside each group. That is what makes E[TX]E[T \mid X] a straight line with slope λ\lambda. For any other distribution of true scores the conditional mean is curved — a heavy-tailed population keeps more of an extreme reading’s lead than the correlation predicts — and then ANCOVA’s fitted line is only the best straight approximation to two curves. The sign of the paradox survives; its closed form does not.

The two groups share one reliability. A baseline measured more precisely in one group than the other gives the two lines different slopes, the gap between them stops being constant, and “the” adjusted group difference depends on where it is read.

Nobody changes. The dining hall was built so that the truth is zero and every non-zero answer is an artefact. A world in which units genuinely grow, and grow by amounts related to where they started, adds a second regression — a real one — on top of the measurement one, and separating the two needs more than two readings.

The groups are large. Nothing here is a small-sample effect. Every number is a limit that 200 units approach and 600 units approach more closely, which is the uncomfortable part: a wrong analysis on a large comparison is wrong with a small standard error, and a precise answer to a question nobody asked looks exactly like a precise answer to the one that was.

What is proved here is the algebra of the two lines and the three variances; what is measured is that 2,000 datasets a cell land on that algebra. What is not settled by either is which world a particular comparison came from, and that is the part of Lord’s paradox that is not arithmetic at all.

Where this goes next

The second world is the one worth following. There the group was chosen by the baseline reading, and the damage came from using that same reading as the baseline: the change score measured the selection’s luck wearing off. The obvious repair is to stop using the reading that selected as the reading that measures, and it raises a question this essay did not need to ask — whether taking a second baseline after the selection removes the artefact completely, or whether averaging several readings at selection does as well. The two sound equivalent and are not. The next question on regression to the mean is enrolment on a threshold, where the selection is explicit, the untreated arm improves by a closed-form amount, and one fresh reading turns out to do what no number of averaged screening readings can.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Assignment mechanismAttenuationChange scoreClosed formConfoundingCovariate adjustmentLord's paradoxMeasurement errorMonte CarloPrecisionRandomisationRegression to the meanTest–retest reliabilityVariance reduction