Reversals that are not errors

A lead that a heavy tail keeps

Four populations whose readings all correlate at exactly 0.6, and whose least-squares slopes all read 0.6. Select the top one per cent on one reading and measure them again: they keep 60% of their lead if the true scores are normal, 76.1% if they are Laplace, 78.4% if they are a t on four degrees of freedom — and 44.3% if they are uniform. The correlation predicts the regression of the extremes for one shape of population only.

Worth reading first: Regression to the mean.

Every measurement of regression to the mean so far has used one sentence as its engine: a reading is partly the person and partly the occasion, and a second reading keeps the share of the lead that was the person. The essay that introduced the effect sized it from the correlation alone, two analyses of one baseline built Lord’s paradox from it, and the measurement that got them enrolled computed a trial’s phantom improvement as (1ρ)(1-\rho) times a truncated mean.

That sentence is a theorem about normal populations. It says that the expected true score behind a reading xx is ρx\rho x, a straight line through the origin with slope equal to the correlation, and it is true when the true scores are normal and false otherwise. The question this essay counts is how false, and where.

The answer is that it holds in the middle of every population and fails at the top of all of them except the normal, in a direction set by the shape of the population’s tail. At the top one per cent, a group that the correlation says keeps 60% of its lead keeps anywhere from 44.3% to 78.4%, and the correlation is exactly 0.6 in all four cases.

The top of a heavy-tailed population keeps its lead; the top of a light-tailed one gives it back. Select the top share on the first reading and read the group again: the share of its mean lead the second reading keeps, by integration over the true score (lines) and counted on 400,000 draws a parent in 20 batches (points, with two standard errors). The normal keeps exactly 0.6 at every selection. At the top half the Laplace keeps 0.541, the t 0.535 and the uniform 0.648 — the heavy tails keep LESS than the correlation. By the top one per cent the order has reversed: 0.761, 0.784 and 0.443. At one in ten thousand the t keeps 0.977 and the uniform 0.346.
Fig. 1 Select the top share of a population on one reading and read the group again. The share of the group’s mean lead the second reading keeps, for four populations whose readings correlate at exactly 0.6: lines by integration over the true score, points counted on 400,000 draws a population with two standard errors.

Four populations that one correlation cannot tell apart

The construction holds everything fixed except the shape of the true scores. In each of four worlds, a true score TT has variance 0.6 and every reading adds normal error of variance 0.4, so a single reading has variance one and any two readings of the same person correlate at exactly 0.6. The four worlds differ only in the distribution TT is drawn from: a normal; a Laplace, whose tails fall exponentially rather than like a normal’s; a t on four degrees of freedom, whose tails fall like a power; and a uniform, which has no tail at all and stops at ±1.342.

Four parents with one variance. The four distributions of the true score: normal, Laplace, t with four degrees of freedom scaled, and uniform on ±1.342, each with variance 0.6 — measured on a grid as 0.600, 0.600, 0.600, 0.601. At the centre the Laplace is tallest at 0.913 and the uniform flattest at 0.373. What separates them for a reading far from the centre is in their tails, which at this scale are nearly invisible: the density at a true score of 3 is normal 2.85e-4, Laplace 3.82e-3, t, four degrees 3.25e-3, uniform 0.00e+0.
Fig. 2 The four distributions of the true score, each with variance 0.6. They differ visibly at the centre and almost invisibly in the tails, which is where the selected group lives.

At the centre they look different and not dramatically so: the Laplace is tallest, at a density of 0.913, the uniform flattest at 0.373. Out at a true score of three they differ by orders of magnitude that the plot cannot show — the normal’s density there is 2.85×1042.85 \times 10^{-4}, the Laplace’s and the t’s more than ten times that, the uniform’s exactly zero. A top one per cent is selected from exactly that region.

Drawn 400,000 times a world, two readings each, the least-squares slope of the second reading on the first is 0.5996, 0.6002, 0.6001 and 0.5972, each with a standard error of 0.0013, and the correlations are 0.5995, 0.6003, 0.5990 and 0.5977. A regression table, a reliability study and a test–retest coefficient report one number for all four worlds, and it is the number they were built to share.

The regression of a reading is a slope of its density

The expected true score behind a reading has an exact expression that makes the role of the population’s shape visible, and it is due to Tweedie. With normal error of variance σ2\sigma^2,

E[TX=x]  =  x+σ2ddxlogfX(x),E[T \mid X = x] \;=\; x + \sigma^2\,\frac{d}{dx}\log f_X(x),

where fXf_X is the density of the readings. The lead a second reading gives back is the error variance times how steeply the log-density of readings is falling at that reading. For a normal population the readings are normal, their log-density is a parabola, its slope is x-x in these units, and the give-back is 0.4x0.4x — a fixed share, 1ρ1-\rho, at every reading. That is the whole of the linear rule.

For any other population the log-density of readings is not a parabola, and the give-back is not a fixed share.

The regression is the error variance times the slope of the readings' log-density. The density of the readings on a logarithmic scale, for the four parents. Tweedie's formula says the expected true score is the reading plus the error variance, 0.4, times the slope of the natural log of this density, so the steeper the curve falls at a reading, the more of the reading a second one gives back. At a reading of 2.5 the lead given back is 1.000 for the normal, 0.728 for the Laplace, 0.789 for the t, four degrees, 1.407 for the uniform; at a reading of 4 it is 1.600, 0.730, 0.536, 2.795. The normal's log-density is a parabola whose slope grows with the reading, so its give-back is a fixed share, 0.4; the Laplace's straightens to a constant slope, so its give-back stops growing; the t's flattens towards zero slope, more slowly; the uniform's steepens past its edge.
Fig. 3 The density of readings on a logarithmic scale, for the four populations. The regression at a reading is 0.4 times the slope of this curve there; the numbers at the right are the lead given back at a reading of 2.5.

At a reading of 2.5 the normal gives back exactly 1.000 of it. The Laplace gives back 0.728 and the t 0.789, because their log-densities have already begun to straighten; the uniform gives back 1.407, because a reading that high can only have come from a true score near the edge at 1.342 plus a large error, so most of it is error. At a reading of four the normal gives back 1.600, the Laplace 0.730, the t 0.536 and the uniform 2.795.

The Laplace’s number barely moved between 2.5 and four, and that is a closed form rather than a coincidence. Far out, the Laplace’s log-density becomes a straight line with slope 1/b-1/b, where b=0.3b = \sqrt{0.3} is its scale, so the give-back settles at σ2/b=0.4/0.3\sigma^2/b = 0.4/\sqrt{0.3}, which is 0.730. A Laplace population gives back a fixed amount at the extremes, not a fixed share, so the share it keeps rises towards one as the reading grows. The t’s log-density flattens further still, towards a slope of zero, and its give-back falls towards nothing. The uniform’s steepens without limit past its edge.

Three routes to one curve

Every curve here is computed three ways that share no arithmetic, because a formula with a derivative of a log-density in it is exactly the kind of number one route is not enough for.

The first route integrates: the posterior mean of the true score is a ratio of two integrals over the parent density times the normal error density, computed by Simpson’s rule on pieces split at the parent’s non-smooth points. The second applies Tweedie’s formula, differentiating the logarithm of the first route’s denominator numerically. The third uses closed forms where they exist — the normal’s straight line, the uniform’s posterior as a normal truncated to ±1.342, and the Laplace’s as a mixture of two truncated normals. Across eighty-one readings from −4 to 4, the integral and the closed forms agree to 9.04×10109.04 \times 10^{-10} for the three parents that have them, and Tweedie’s formula agrees with the integral to 7.38×1087.38 \times 10^{-8} for all four.

One correlation, four regressions: how much of a reading's lead survives a second reading. The expected second reading divided by the first, for readings from 0.5 to 4 standard deviations, where the true scores come from a normal, a Laplace, a t with four degrees of freedom or a uniform parent, every one with variance 0.6, and every reading adds normal error of variance 0.4. All four put the two readings at a correlation of exactly 0.6. For the normal parent the share kept is 0.6 at every reading. At a reading of 2 the Laplace parent keeps 0.644, the t parent 0.604 and the uniform 0.507; at a reading of 4 they keep 0.817, 0.866 and 0.301. The curves are posterior means by integration, checked against Tweedie's formula to 7e-8 and against closed forms for three of the parents to 9e-10.
Fig. 4 The expected second reading as a share of the first, for readings from half a standard deviation to four. The normal’s line is flat at the correlation; the others cross it.

Read as the share of a reading a second reading keeps: at a reading of one standard deviation, the Laplace keeps 0.493, the t 0.492 and the uniform 0.691 — the heavy tails keep less than the correlation, and the light tail more. At a reading of two they keep 0.644, 0.604 and 0.507; at four, 0.817, 0.866 and 0.301. Every non-normal curve crosses the correlation somewhere between one and two standard deviations, and the ordering at the top is the reverse of the ordering in the middle.

Why a slope of 0.6 is consistent with all of this

It is natural to suspect the least-squares slope should have noticed. It did not, and it could not, because a straight line fitted to the whole population is an average of the regression over every reading, weighted by how many readings there are. The middle holds nearly all of them.

One fitted line, and four different curves around it. Draws of 400,000 true scores from each parent, two readings each. The least-squares slope of the second reading on the first is 0.5996 (normal), 0.6002 (Laplace), 0.6001 (t, four degrees), 0.5972 (uniform) — the correlation, for all four. The figure plots what the line leaves: the mean second reading in bins of the first, minus 0.6 times the first, as points, with the posterior-mean curve drawn through them. The normal's departures are flat at zero. The heavy-tailed parents' curves are S-shaped — below the line in the middle of the positive half and above it at the extremes — and the uniform's is the mirror image. A single fitted line averages over exactly these departures, which is why its slope is the same for all four.
Fig. 5 What the fitted line leaves: the mean second reading in bins of the first, minus 0.6 times the first, with the posterior-mean curve through the points. Flat for the normal; S-shaped the other ways for the rest.

For the two heavy-tailed populations the departures from the line are S-shaped — the expected second reading sits below the line in the middle of the positive half and above it at the extremes — and for the uniform the S is reflected. Integrated against the density of readings, each S cancels to zero, which is why every population’s slope is 0.6. The departures that matter to anyone selecting on a reading are the ones at the ends of the S, where there are almost no readings to pull the fitted line towards them.

This is a general property of summaries, not a special weakness of regression: a correlation describes a joint distribution by one number, and one number cannot distinguish shapes that share it. What is particular here is that the summary is correct about the middle and wrong about exactly the part a selection reads, the same asymmetry by which a tail converges last while the middle of a distribution is already right.

What the top of each population keeps

Selecting a group is not reading one value; it is averaging over every reading above a cut, and the share of the group’s lead that survives is the ratio of two integrals over the true score. The opening figure gives it at seven selections, and the numbers carry the argument.

The normal keeps 0.6000 at every selection, as it must, and the count reads 0.5986 at the top fifth, 0.5935 at the top one per cent and 0.5912 at the top thousandth, each within its standard error.

The Laplace keeps 0.5411 at the top half, 0.6254 at the top tenth, 0.7612 at the top one per cent, 0.8308 at the top thousandth and 0.8691 at the top ten-thousandth. Counted, the top one per cent keeps 0.7587 ± 0.0044 and the top thousandth 0.8228 ± 0.0091.

The t keeps 0.5355, 0.6016, 0.7842, 0.9262 and 0.9768 at the same selections, and counts at 0.7934 ± 0.0047 and 0.9305 ± 0.0073 for the top one per cent and thousandth — the first about two of its standard errors above the integral, which is where the t’s rare very large draws make any finite count lumpy.

The uniform keeps 0.6482, 0.5478, 0.4432, 0.3845 and 0.3458, and counts at 0.4371 ± 0.0053 and 0.3805 ± 0.0090.

Two things in that table run against intuition, and the first is easy to miss. At broad selections the heavy-tailed populations regress more than the correlation says, not less. A top half, or a top fifth, is mostly readings between one and two standard deviations, which is where the heavy-tailed curves sit below 0.6. Only once the selection reaches into the region where the log-density has straightened does the reversal happen, between the top tenth and the top twentieth. The second is how far the t goes: its top thousandth reads 5.502 at selection and 5.096 on a second reading, a group that keeps more than nine-tenths of an extreme lead that a normal population would have given back to 3.301.

What a correction built on the correlation removes

The practical form of the rule is a correction: predict a selected group’s second reading as ρ\rho times its first, and treat anything beyond that as a real change. On a normal population that is exact. On the other three it is wrong by amounts that grow with the selection.

What a correction built on the correlation gets wrong at the top. The error of predicting a selected group's second reading as 0.6 times its first — the correction the correlation supplies — at the top one per cent and the top one in a thousand, in standard deviations of one reading. For the normal parent it is exact. For the Laplace it predicts too little by 0.492 and 0.996, for the t by 0.588 and 1.795: it strips out lead the group really has. For the uniform it predicts too much by 0.383 and 0.639, leaving in lead that was noise. The top one in a thousand of the t parent reads 5.502 at selection and 5.096 on a second reading.
Fig. 6 The error of predicting a selected group’s second reading as 0.6 times its first, at the top one per cent and the top thousandth. Negative bars are lead the correction strips out although it was real; positive bars are lead it leaves in although it was noise.

For the Laplace population the correction predicts too low by 0.4925 standard deviations at the top one per cent and 0.9964 at the top thousandth; for the t by 0.5881 and 1.7950. In both it removes lead the group really has, and a real effect in that group is under-reported by the same amount. For the uniform it predicts too high by 0.3832 and 0.6392, leaving in lead that was the occasion rather than the person: a group with nothing done to it would be credited with an improvement of that size that a correct model would have predicted.

The practical settings divide by the shape of what is measured. Quantities with a long upper tail — income, output, citations, the size of an effect across a literature, the severity of a rare condition — are the Laplace-and-t side: the top of a league table in them is more likely to stay near the top than the test–retest correlation says, and a correction built on that correlation over-corrects the extremes. Quantities with a hard ceiling — a score on a test with a maximum, a proportion near one, a rating on a bounded scale — are the uniform side, where the top regresses more than the correlation says and the correction under-corrects. The winner of a contest scored on a bounded scale owes more of the win to the day than the same contest’s reliability suggests, which is the winner’s curse in the population where it bites hardest.

What an enrolment artefact becomes

The enrolment case, measured most carefully of all, is the one most exposed to the shape of the population. An untreated group enrolled on a single reading improves by (1ρ)(1-\rho) times its truncated mean, 0.702 standard deviations at the top tenth, and that number was derived from a normal population of true scores. The same selection in the other three worlds, with the same correlation, gives different phantom improvements, and the table above already contains them as the gap between a group’s reading at selection and its expected second reading.

At the top tenth the four populations barely differ: the untreated fall is 0.702 for the normal, 0.668 for the Laplace, 0.701 for the t and 0.775 for the uniform. A top tenth is still mostly readings from the middle of the distribution, where the four regressions are close. At the top one per cent they separate — 1.066, 0.730, 0.689 and 1.36 — and at the top thousandth they are 1.347, 0.730, 0.406 and 1.825. A trial that enrols the most severe thousandth of a population whose severity has a power-law tail should expect its untreated arm to improve by less than a third of what the normal model predicts; the same trial on a bounded score should expect a third more.

Two of that essay’s conclusions survive the change of shape and one does not, and which is which follows from the mechanism. A fresh baseline still removes the artefact completely, in every world: the follow-up and a reading taken after enrolment are both the person’s true score plus error the selection never saw, so they have the same expected value whatever the population’s shape, and the fall between them is zero. A randomised control arm still cancels the artefact from the comparison, since both arms are selected the same way. What does not survive is the arithmetic of averaging screening readings. Spearman–Brown gives the reliability of an average, and that reliability set the size of the remaining artefact only because a normal population makes the regression a fixed share of the reading. In a heavy-tailed population, averaging still helps, and how much it helps at a given cut is no longer a formula in the reliability alone.

The shrinkage of an estimate is the same arithmetic

Tweedie’s formula is not special to readings of people. Replace the true score with a true effect, the reading with an estimate, and the population with the distribution of effects across groups, and it becomes the rule for how far an estimate should be shrunk. The weight a hierarchical model uses, B=se2/(se2+τ2)B = \mathrm{se}^2/(\mathrm{se}^2 + \tau^2), is its normal special case: a fixed share of every estimate’s distance from the mean, whatever the distance.

That is exactly the rule this essay found wrong at the extremes, and it explains a failure measured elsewhere in the collection. When borrowing goes wrong, a group far from the others is pulled towards them by the same share as every other group, and estimated six times worse than by its own mean. Under a normal distribution of effects that pull is correct, because a far estimate is mostly noise. Under a heavy-tailed distribution of effects, a far estimate is more likely a far effect, and the correct pull is a fixed amount rather than a fixed share — so the correct estimate keeps most of its distance, and normal shrinkage takes it away. The misfit group in that essay is one point from a population of effects that the normal model assumed had no tail.

What is proved here and what is only computed for these four

Proved. Tweedie’s formula is exact whenever the error is normal, whatever the population of true scores; it is a line of calculus, and the three routes here check the arithmetic rather than the formula. That the normal population keeps exactly ρ\rho of its lead at every reading and every selection follows from it. That the Laplace’s give-back tends to σ2/b\sigma^2/b follows from the Laplace log-density becoming linear, and that a bounded population’s expected true score cannot exceed its bound is immediate.

Computed for these four populations. Every share kept, every correction error and every crossing point belongs to the four parents chosen, all at variance 0.6 and all with error variance 0.4. They were chosen to span the three kinds of tail — exponential, power and none — rather than to model a particular measurement. The direction of each departure is what should carry over: a population whose readings’ log-density flattens in the upper tail keeps more of an extreme lead than its correlation, and one whose log-density steepens keeps less.

Assumed and not examined. The error is normal throughout, and that assumption is doing real work. If the error has the heavy tail instead — occasional wild readings of an ordinary person — the conclusion reverses: an extreme reading is then more likely to be a wild occasion, and it gives back more than the correlation says. Tweedie’s formula in the form used here does not apply to that case, and nothing here measures it. Neither does anything here decide, for a particular quantity, which shape its true scores have; a quantile plot of a sample of forty is famously poor at telling a normal from its neighbours, and the difference that matters lives beyond where most samples have any points.

Where this goes next: estimating the slope of a density nobody can see

Tweedie’s formula has one property this essay has not used and the next measurement should. It needs the density of the readings, not of the true scores — and the readings are observed. In principle a study with enough readings can estimate fXf_X, differentiate its logarithm, and compute the right regression for its own population without ever knowing the parent, which is the idea behind the empirical-Bayes corrections that a prior estimated from the data introduced for a single variance.

The catch is precisely where the correction is needed. The slope of a log-density is hardest to estimate in the tail, where there are fewest readings, and the top one per cent of a thousand readings is ten points. The measurement that would settle whether the idea survives contact with a real sample size is the share kept by the top one per cent, estimated from a kernel or a fitted log-spline of a thousand readings in each of these four worlds, against the integrals above — and whether the estimated correction beats the linear one, loses to it, or merely replaces a known bias with an unknown variance. It is distinct from this essay because it asks what a single dataset can recover of the shape, where this one assumed the shape was known.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCorrelationHeavy tailMeasurement errorMonte CarloNormalityPosterior meanRegression to the meanSelection biasShrinkageTest–retest reliabilityTweedie's formulaUniform distribution