A lead that a heavy tail keeps
Worth reading first: Regression to the mean.
Every measurement of regression to the mean so far has used one sentence as its engine: a reading is partly the person and partly the occasion, and a second reading keeps the share of the lead that was the person. The essay that introduced the effect sized it from the correlation alone, two analyses of one baseline built Lord’s paradox from it, and the measurement that got them enrolled computed a trial’s phantom improvement as times a truncated mean.
That sentence is a theorem about normal populations. It says that the expected true score behind a reading is , a straight line through the origin with slope equal to the correlation, and it is true when the true scores are normal and false otherwise. The question this essay counts is how false, and where.
The answer is that it holds in the middle of every population and fails at the top of all of them except the normal, in a direction set by the shape of the population’s tail. At the top one per cent, a group that the correlation says keeps 60% of its lead keeps anywhere from 44.3% to 78.4%, and the correlation is exactly 0.6 in all four cases.
Four populations that one correlation cannot tell apart
The construction holds everything fixed except the shape of the true scores. In each of four worlds, a true score has variance 0.6 and every reading adds normal error of variance 0.4, so a single reading has variance one and any two readings of the same person correlate at exactly 0.6. The four worlds differ only in the distribution is drawn from: a normal; a Laplace, whose tails fall exponentially rather than like a normal’s; a t on four degrees of freedom, whose tails fall like a power; and a uniform, which has no tail at all and stops at ±1.342.
At the centre they look different and not dramatically so: the Laplace is tallest, at a density of 0.913, the uniform flattest at 0.373. Out at a true score of three they differ by orders of magnitude that the plot cannot show — the normal’s density there is , the Laplace’s and the t’s more than ten times that, the uniform’s exactly zero. A top one per cent is selected from exactly that region.
Drawn 400,000 times a world, two readings each, the least-squares slope of the second reading on the first is 0.5996, 0.6002, 0.6001 and 0.5972, each with a standard error of 0.0013, and the correlations are 0.5995, 0.6003, 0.5990 and 0.5977. A regression table, a reliability study and a test–retest coefficient report one number for all four worlds, and it is the number they were built to share.
The regression of a reading is a slope of its density
The expected true score behind a reading has an exact expression that makes the role of the population’s shape visible, and it is due to Tweedie. With normal error of variance ,
where is the density of the readings. The lead a second reading gives back is the error variance times how steeply the log-density of readings is falling at that reading. For a normal population the readings are normal, their log-density is a parabola, its slope is in these units, and the give-back is — a fixed share, , at every reading. That is the whole of the linear rule.
For any other population the log-density of readings is not a parabola, and the give-back is not a fixed share.
At a reading of 2.5 the normal gives back exactly 1.000 of it. The Laplace gives back 0.728 and the t 0.789, because their log-densities have already begun to straighten; the uniform gives back 1.407, because a reading that high can only have come from a true score near the edge at 1.342 plus a large error, so most of it is error. At a reading of four the normal gives back 1.600, the Laplace 0.730, the t 0.536 and the uniform 2.795.
The Laplace’s number barely moved between 2.5 and four, and that is a closed form rather than a coincidence. Far out, the Laplace’s log-density becomes a straight line with slope , where is its scale, so the give-back settles at , which is 0.730. A Laplace population gives back a fixed amount at the extremes, not a fixed share, so the share it keeps rises towards one as the reading grows. The t’s log-density flattens further still, towards a slope of zero, and its give-back falls towards nothing. The uniform’s steepens without limit past its edge.
Three routes to one curve
Every curve here is computed three ways that share no arithmetic, because a formula with a derivative of a log-density in it is exactly the kind of number one route is not enough for.
The first route integrates: the posterior mean of the true score is a ratio of two integrals over the parent density times the normal error density, computed by Simpson’s rule on pieces split at the parent’s non-smooth points. The second applies Tweedie’s formula, differentiating the logarithm of the first route’s denominator numerically. The third uses closed forms where they exist — the normal’s straight line, the uniform’s posterior as a normal truncated to ±1.342, and the Laplace’s as a mixture of two truncated normals. Across eighty-one readings from −4 to 4, the integral and the closed forms agree to for the three parents that have them, and Tweedie’s formula agrees with the integral to for all four.
Read as the share of a reading a second reading keeps: at a reading of one standard deviation, the Laplace keeps 0.493, the t 0.492 and the uniform 0.691 — the heavy tails keep less than the correlation, and the light tail more. At a reading of two they keep 0.644, 0.604 and 0.507; at four, 0.817, 0.866 and 0.301. Every non-normal curve crosses the correlation somewhere between one and two standard deviations, and the ordering at the top is the reverse of the ordering in the middle.
Why a slope of 0.6 is consistent with all of this
It is natural to suspect the least-squares slope should have noticed. It did not, and it could not, because a straight line fitted to the whole population is an average of the regression over every reading, weighted by how many readings there are. The middle holds nearly all of them.
For the two heavy-tailed populations the departures from the line are S-shaped — the expected second reading sits below the line in the middle of the positive half and above it at the extremes — and for the uniform the S is reflected. Integrated against the density of readings, each S cancels to zero, which is why every population’s slope is 0.6. The departures that matter to anyone selecting on a reading are the ones at the ends of the S, where there are almost no readings to pull the fitted line towards them.
This is a general property of summaries, not a special weakness of regression: a correlation describes a joint distribution by one number, and one number cannot distinguish shapes that share it. What is particular here is that the summary is correct about the middle and wrong about exactly the part a selection reads, the same asymmetry by which a tail converges last while the middle of a distribution is already right.
What the top of each population keeps
Selecting a group is not reading one value; it is averaging over every reading above a cut, and the share of the group’s lead that survives is the ratio of two integrals over the true score. The opening figure gives it at seven selections, and the numbers carry the argument.
The normal keeps 0.6000 at every selection, as it must, and the count reads 0.5986 at the top fifth, 0.5935 at the top one per cent and 0.5912 at the top thousandth, each within its standard error.
The Laplace keeps 0.5411 at the top half, 0.6254 at the top tenth, 0.7612 at the top one per cent, 0.8308 at the top thousandth and 0.8691 at the top ten-thousandth. Counted, the top one per cent keeps 0.7587 ± 0.0044 and the top thousandth 0.8228 ± 0.0091.
The t keeps 0.5355, 0.6016, 0.7842, 0.9262 and 0.9768 at the same selections, and counts at 0.7934 ± 0.0047 and 0.9305 ± 0.0073 for the top one per cent and thousandth — the first about two of its standard errors above the integral, which is where the t’s rare very large draws make any finite count lumpy.
The uniform keeps 0.6482, 0.5478, 0.4432, 0.3845 and 0.3458, and counts at 0.4371 ± 0.0053 and 0.3805 ± 0.0090.
Two things in that table run against intuition, and the first is easy to miss. At broad selections the heavy-tailed populations regress more than the correlation says, not less. A top half, or a top fifth, is mostly readings between one and two standard deviations, which is where the heavy-tailed curves sit below 0.6. Only once the selection reaches into the region where the log-density has straightened does the reversal happen, between the top tenth and the top twentieth. The second is how far the t goes: its top thousandth reads 5.502 at selection and 5.096 on a second reading, a group that keeps more than nine-tenths of an extreme lead that a normal population would have given back to 3.301.
What a correction built on the correlation removes
The practical form of the rule is a correction: predict a selected group’s second reading as times its first, and treat anything beyond that as a real change. On a normal population that is exact. On the other three it is wrong by amounts that grow with the selection.
For the Laplace population the correction predicts too low by 0.4925 standard deviations at the top one per cent and 0.9964 at the top thousandth; for the t by 0.5881 and 1.7950. In both it removes lead the group really has, and a real effect in that group is under-reported by the same amount. For the uniform it predicts too high by 0.3832 and 0.6392, leaving in lead that was the occasion rather than the person: a group with nothing done to it would be credited with an improvement of that size that a correct model would have predicted.
The practical settings divide by the shape of what is measured. Quantities with a long upper tail — income, output, citations, the size of an effect across a literature, the severity of a rare condition — are the Laplace-and-t side: the top of a league table in them is more likely to stay near the top than the test–retest correlation says, and a correction built on that correlation over-corrects the extremes. Quantities with a hard ceiling — a score on a test with a maximum, a proportion near one, a rating on a bounded scale — are the uniform side, where the top regresses more than the correlation says and the correction under-corrects. The winner of a contest scored on a bounded scale owes more of the win to the day than the same contest’s reliability suggests, which is the winner’s curse in the population where it bites hardest.
What an enrolment artefact becomes
The enrolment case, measured most carefully of all, is the one most exposed to the shape of the population. An untreated group enrolled on a single reading improves by times its truncated mean, 0.702 standard deviations at the top tenth, and that number was derived from a normal population of true scores. The same selection in the other three worlds, with the same correlation, gives different phantom improvements, and the table above already contains them as the gap between a group’s reading at selection and its expected second reading.
At the top tenth the four populations barely differ: the untreated fall is 0.702 for the normal, 0.668 for the Laplace, 0.701 for the t and 0.775 for the uniform. A top tenth is still mostly readings from the middle of the distribution, where the four regressions are close. At the top one per cent they separate — 1.066, 0.730, 0.689 and 1.36 — and at the top thousandth they are 1.347, 0.730, 0.406 and 1.825. A trial that enrols the most severe thousandth of a population whose severity has a power-law tail should expect its untreated arm to improve by less than a third of what the normal model predicts; the same trial on a bounded score should expect a third more.
Two of that essay’s conclusions survive the change of shape and one does not, and which is which follows from the mechanism. A fresh baseline still removes the artefact completely, in every world: the follow-up and a reading taken after enrolment are both the person’s true score plus error the selection never saw, so they have the same expected value whatever the population’s shape, and the fall between them is zero. A randomised control arm still cancels the artefact from the comparison, since both arms are selected the same way. What does not survive is the arithmetic of averaging screening readings. Spearman–Brown gives the reliability of an average, and that reliability set the size of the remaining artefact only because a normal population makes the regression a fixed share of the reading. In a heavy-tailed population, averaging still helps, and how much it helps at a given cut is no longer a formula in the reliability alone.
The shrinkage of an estimate is the same arithmetic
Tweedie’s formula is not special to readings of people. Replace the true score with a true effect, the reading with an estimate, and the population with the distribution of effects across groups, and it becomes the rule for how far an estimate should be shrunk. The weight a hierarchical model uses, , is its normal special case: a fixed share of every estimate’s distance from the mean, whatever the distance.
That is exactly the rule this essay found wrong at the extremes, and it explains a failure measured elsewhere in the collection. When borrowing goes wrong, a group far from the others is pulled towards them by the same share as every other group, and estimated six times worse than by its own mean. Under a normal distribution of effects that pull is correct, because a far estimate is mostly noise. Under a heavy-tailed distribution of effects, a far estimate is more likely a far effect, and the correct pull is a fixed amount rather than a fixed share — so the correct estimate keeps most of its distance, and normal shrinkage takes it away. The misfit group in that essay is one point from a population of effects that the normal model assumed had no tail.
What is proved here and what is only computed for these four
Proved. Tweedie’s formula is exact whenever the error is normal, whatever the population of true scores; it is a line of calculus, and the three routes here check the arithmetic rather than the formula. That the normal population keeps exactly of its lead at every reading and every selection follows from it. That the Laplace’s give-back tends to follows from the Laplace log-density becoming linear, and that a bounded population’s expected true score cannot exceed its bound is immediate.
Computed for these four populations. Every share kept, every correction error and every crossing point belongs to the four parents chosen, all at variance 0.6 and all with error variance 0.4. They were chosen to span the three kinds of tail — exponential, power and none — rather than to model a particular measurement. The direction of each departure is what should carry over: a population whose readings’ log-density flattens in the upper tail keeps more of an extreme lead than its correlation, and one whose log-density steepens keeps less.
Assumed and not examined. The error is normal throughout, and that assumption is doing real work. If the error has the heavy tail instead — occasional wild readings of an ordinary person — the conclusion reverses: an extreme reading is then more likely to be a wild occasion, and it gives back more than the correlation says. Tweedie’s formula in the form used here does not apply to that case, and nothing here measures it. Neither does anything here decide, for a particular quantity, which shape its true scores have; a quantile plot of a sample of forty is famously poor at telling a normal from its neighbours, and the difference that matters lives beyond where most samples have any points.
Where this goes next: estimating the slope of a density nobody can see
Tweedie’s formula has one property this essay has not used and the next measurement should. It needs the density of the readings, not of the true scores — and the readings are observed. In principle a study with enough readings can estimate , differentiate its logarithm, and compute the right regression for its own population without ever knowing the parent, which is the idea behind the empirical-Bayes corrections that a prior estimated from the data introduced for a single variance.
The catch is precisely where the correction is needed. The slope of a log-density is hardest to estimate in the tail, where there are fewest readings, and the top one per cent of a thousand readings is ten points. The measurement that would settle whether the idea survives contact with a real sample size is the share kept by the top one per cent, estimated from a kernel or a fitted log-spline of a thousand readings in each of these four worlds, against the integrals above — and whether the estimated correction beats the linear one, loses to it, or merely replaces a known bias with an unknown variance. It is distinct from this essay because it asks what a single dataset can recover of the shape, where this one assumed the shape was known.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Adjusting for a shadow — both name closed form, measurement error, monte carlo, test–retest reliability
- The sample is a condition — both name closed form, correlation, monte carlo, selection bias
- A lag the sample has less of — both name closed form, monte carlo, shrinkage
- A league table of a hundred — both name posterior mean, regression to the mean, shrinkage
- A simulation that stops when it looks settled — both name closed form, monte carlo, selection bias
- Leaving each row out of its own first stage — both name closed form, heavy tail, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Closed formCorrelationHeavy tailMeasurement errorMonte CarloNormalityPosterior meanRegression to the meanSelection biasShrinkageTest–retest reliabilityTweedie's formulaUniform distribution