Reversals that are not errors

A correction built on a retest

Tweedie's correction for regression to the mean needs the error variance, and a study gets it from a few people measured twice. Twenty of them are enough: with normal errors the estimated correction then beats the correlation's rule on 73% to 77% of studies of a thousand readings from a heavy-tailed population, though its variance is uncertain by a third. What twenty cannot reveal is the error's shape. Laplace errors of the same variance make the correction overstate the lead a top group keeps by about 0.17, the plain rule then wins on most studies, and a study needs about two hundred retest pairs before a check on its own differences notices.

Worth reading first: Regression to the mean.

The slope of a density nobody can see corrected a selected group’s readings for regression to the mean by Tweedie’s formula — the reading plus the error variance times the slope of the readings’ log-density — with the slope estimated from the study’s own readings. At a thousand readings the estimated correction beat the linear rule, which shrinks every reading by the correlation between two readings, on most studies from every population that was not normal. It took the error variance as known: 0.4, from a test–retest study somewhere else.

A study usually has to estimate that too, and the obvious source is a retest subsample: some of the people measured twice, so that the differences between their two readings are differences of two errors. Half the mean squared difference estimates the error variance. The same pairs give the correlation between two readings, which is the reliability the linear rule shrinks by. So both corrections can be built from what the study itself has, and the question is how many people it has to measure twice for the more flexible one to stay ahead.

Both corrections from the same retest

The design is the earlier essay’s. Studies of a thousand people, each with a reading that is a true score plus an error of variance 0.4, from populations of true scores that are Laplace, t with four degrees of freedom and uniform, each scaled so that two readings correlate at 0.6. Each study selects its top one per cent — ten people — and predicts the share of their lead over the population mean that a second reading would keep. The target is that share, worked out by integrating over the population, and every estimate is scored by how far from it it lands.

Now the first m people in each study are measured twice. From their pairs the study takes the error variance as half the mean squared difference, and the reliability as the correlation of the two readings. Tweedie’s correction uses the estimated variance with a kernel estimate of the log-density’s slope, the better of the two estimators at this size in the earlier essay. The linear rule uses the estimated reliability. Neither is given a number the study could not have computed.

Twenty people are enough

How many people a study has to retest before its estimated correction is worth having, normal errorsRoot mean squared error, over 400 studies of a thousand readings at each retest size, of the share of the top one per cent's lead a second reading keeps: Tweedie's correction with a kernel log-slope and the error variance estimated from the retested people (solid), and the linear rule at the reliability those people estimate (dashed). Laplace parent: 0.121, 0.108, 0.084, 0.086, 0.084, 0.078 against 0.343, 0.259, 0.201, 0.183, 0.171, 0.164. T, four degrees parent: 0.133, 0.103, 0.086, 0.084, 0.080, 0.078 against 0.358, 0.278, 0.235, 0.212, 0.205, 0.187. Uniform parent: 0.247, 0.192, 0.156, 0.140, 0.118, 0.117 against 0.260, 0.208, 0.185, 0.166, 0.160, 0.157 — at 10, 20, 50, 100, 200, 1000 retested.00.1000.2000.3000.400people retested (log scale)error of the share kept (root mean square)1020501002001000Laplace parentt, four degrees parentuniform parentsolid: Tweedie with the variance estimated · dashed: the linear rule at the estimated reliability400 studies of 1,000 readings at each size; top one per centerrors normal, variance 0.4
Fig. 1 The error of each correction, as a root mean square over 400 studies of a thousand readings, against the number of people retested, with normal errors. Solid: Tweedie’s correction with the error variance estimated. Dashed: the linear rule at the estimated reliability. For the two heavy-tailed populations the estimated correction is better at every retest size.

With the variance known, the correction’s error in the t population is 0.080. With the variance estimated from a thousand retested people it is 0.078, from a hundred 0.084, from twenty 0.103, and from ten 0.133. The linear rule, at the reliability the same people estimate, has an error of 0.278 with twenty retested and 0.187 with a thousand. In the Laplace population the pattern is the same: the estimated correction’s error is 0.108 at twenty retested and the rule’s 0.259.

What makes this work is how little of the correction depends on the variance. The variance estimated from m pairs of normal errors has a relative error of about 2/m\sqrt{2/m}: at twenty people its standard deviation is 31.5% of its value, at ten 42.9%. But the correction multiplies the variance by the log-density’s slope and adds the product to the reading, so a variance off by a third moves the corrected share by a third of the correction itself. In the t population the correction takes the share from one to about 0.8, so a third of it is about 0.07 — well under half of the 0.18 the rule is wrong by when it assumes the population is normal. The slope estimate, which the retest does not touch, contributes most of the remaining error at every retest size from fifty up, which is why the curves flatten there.

The rule does not get better at the same rate, because its error is not mostly noise. In the t population the share kept is 0.784, and a perfectly estimated reliability is 0.6: the rule is wrong by 0.18 however many people are retested, and estimating the reliability only adds noise on top. More retested people bring the estimated correction down to its floor and leave the rule at its own.

How often the estimate wins

How often the estimated correction beats the linear rule, by the number retested and the error law. Share of 400 studies on which Tweedie's estimated correction is closer to the share kept than the linear rule at the estimated reliability. Normal errors (solid): Laplace 67%, 73%, 81%, 85%, 88%, 94%; t, four degrees 72%, 77%, 86%, 89%, 93%, 96%; uniform 59%, 62%, 71%, 70%, 78%, 81%. Laplace errors (dashed): Laplace 42%, 31%, 24%, 15%, 12%, 6%; t, four degrees 43%, 37%, 30%, 22%, 18%, 13%; uniform 33%, 20%, 18%, 18%, 13%, 13% — at 10, 20, 50, 100, 200, 1000 retested.
Fig. 2 The share of 400 studies on which the estimated correction lands closer to the truth than the rule, against the number retested. Solid lines are normal errors and dashed are Laplace errors of the same variance; above one half, estimating pays.

With normal errors the estimate wins on most studies at every retest size. In the t population it wins on 72% of studies with ten people retested, 77% with twenty, 89% with a hundred and 96% with a thousand. In the Laplace population 67%, 73%, 85% and 94%. The uniform population is the hard case, as it was before: its readings have a boundary that a kernel smooths over, so the correction carries a bias of its own, and the estimate wins on 59% of studies at ten retested, 71% at fifty and 81% at a thousand.

So the earlier essay’s conclusion survives the retest. A study of a thousand readings from a population with heavy tails does better correcting its top group by an estimated log-density and a variance from twenty retested people than by the correlation those same twenty people give it. That is a cheap condition, and it is met by most studies that have a reliability figure at all.

An error with a heavy tail

The dashed lines in the same figure tell a different story, and it is the one the earlier essay’s last section feared. Tweedie’s formula is exact only when the error is normal. It says the expected true score is the reading plus the error variance times the log-density’s slope, and that identity comes from differentiating a normal error’s density. With any other error law, the formula is still computable — the study still has readings, a slope and a variance — but it is no longer the regression of the true score on the reading.

What a heavy-tailed error does to Tweedie's correction: the share kept, the correction's estimate of it and the linear rule's. For each parent and error law, at a thousand readings with the whole sample retested: the share of the top one per cent's lead a second reading keeps, Tweedie's correction with the true error variance averaged over 400 studies, and the linear rule's estimated reliability. Laplace parent, normal errors: kept 0.761, Tweedie 0.782, rule 0.599; Laplace parent, Laplace errors: kept 0.643, Tweedie 0.807, rule 0.600; t, four degrees parent, normal errors: kept 0.784, Tweedie 0.802, rule 0.600; t, four degrees parent, Laplace errors: kept 0.666, Tweedie 0.830, rule 0.599; uniform parent, normal errors: kept 0.443, Tweedie 0.526, rule 0.599; uniform parent, Laplace errors: kept 0.328, Tweedie 0.702, rule 0.599.
Fig. 3 For each population, under normal errors and under Laplace errors of the same variance: the share of the top group’s lead a second reading keeps, Tweedie’s correction with the true error variance, and the linear rule. Under Laplace errors the correction lands far above the truth and the rule lands closer.

The error law changes the target, and the correction does not notice. With Laplace errors of variance 0.4, the top one per cent of the t population keeps 0.666 of its lead, counted on two million people, against 0.784 with normal errors. Tweedie’s correction, given the exact error variance, says 0.830 — too high by 0.164, which is a larger error than the rule’s under normal errors. The linear rule says 0.599, too low by 0.067. In the Laplace population the share kept is 0.643, the correction says 0.807 and the rule 0.600. In the uniform population the share kept falls to 0.328, and the correction overshoots it by 0.374.

The direction is the informative part. A heavy-tailed error produces extreme readings by itself: some of the people in the top one per cent are there because their error was unusually large, not because their true score was. A normal error makes that rare, so a normal-error formula attributes most of an extreme reading to the person; a Laplace error makes it common, so more of the lead is noise and less of it survives a second reading. Tweedie’s formula computed as though the error were normal assigns the noise to the person and over-predicts what the second reading keeps — precisely the error the correction was built to remove, reintroduced by its own assumption.

Where a top group’s lead comes from

The direction can be counted rather than argued, because a simulation knows each person’s true score. On the two million people behind each share, the top one per cent of the t population under normal errors reads 3.192 on average, and 0.689 of that is error; a fifth of them, 20.5%, have an error larger than two of its standard deviations. Under Laplace errors the top one per cent reads 3.330 — the selection reaches a little further, because the readings’ tail is heavier — and 1.111 of it is error, with 40.8% of the group carrying an error past two standard deviations. The second reading draws a fresh error, so whatever part of the lead was error does not come back, and the share kept falls from 0.783 to 0.666.

The same count checks the integral it replaces. Under normal errors the counted share kept in the t population is 0.7834 and the integral is 0.7842, which is two routes to one number that share nothing but the model; under Laplace errors only the count is available, and the agreement under normal errors is the reason to trust it there.

The uniform population shows the mechanism at its starkest. Its true scores stop at 1.342, well short of the reading of 2.187 at which its top one per cent begins, so every selected reading is mostly error even with normal errors: the top group reads 2.444 on average and 1.360 of it is error. Under Laplace errors the group reads 2.739, of which 1.836 is error, and 84.0% of its members owe their place to an error beyond two standard deviations. A second reading keeps a third of the lead. Tweedie’s formula under the normal assumption, which cannot see that the true scores have run out, predicts it keeps seven tenths.

And the comparison reverses. Under Laplace errors the estimated correction beats the rule on 13% of studies in the t population with a thousand retested, and on 6% in the Laplace population. The rule wins not because it is right — it assumes normal true scores and normal errors — but because its two wrong assumptions push in opposite directions here, and Tweedie’s correction keeps one of them. A heavy-tailed population makes a top group keep more than the rule says, and a heavy-tailed error makes it keep less; the rule, which knows about neither, lands between them. That is the cancellation when borrowing goes wrong warned about from the other side, where an assumption that is wrong in a known direction can still beat an estimate that is right about one thing and wrong about another, and it is the reason the plain account of regression to the mean — shrink by the reliability — survives in practice better than its assumptions deserve.

A larger retest does not help the correction

Under Laplace errors the retest size works in the rule’s favour and not the correction’s. In the t population the estimated correction’s error is 0.211 with ten people retested and 0.178 with a thousand — and with the variance known exactly it is also 0.178, so the whole of what remains is the formula’s bias, which no amount of retesting touches. The rule’s error falls from 0.299 with ten retested to 0.076 with a thousand, because its only problem at small retests was the noise in its estimated reliability, and the reliability it converges on, 0.6, happens to lie close to the 0.666 the top group keeps.

So the two corrections cross as the retest grows. With ten retested the estimated correction still lands closer on 43% of studies, because the rule’s noise is then as large as the correction’s bias; by twenty it is 37%, and from fifty retested on the rule wins on most studies. A study under Laplace errors that invests in a large retest makes its simple correction better and its sophisticated one no better at all, which is the reverse of the case under normal errors and the reason the error’s shape, and not its size, is what a retest is most needed to settle.

What the retest pairs can see

The study’s retest pairs carry the information that would warn it. Each pair’s difference is the difference of two errors, so its distribution is the error law’s shape, convolved with itself: normal for normal errors, and with an excess kurtosis of 3⁄2 for Laplace errors. A study can compute the kurtosis of its own differences and ask whether it is larger than normal errors would produce.

Whether a study's own retest pairs can show that its error is not normal. The share of studies whose retest differences have an excess kurtosis above its 95th percentile under normal errors, when the errors are Laplace: 11.8% at 10, 19.8% at 20, 34.9% at 50, 57.4% at 100, 80.3% at 200, 100.0% at 1000, over 4,000 studies at each size.
Fig. 4 The chance that the excess kurtosis of a study’s retest differences exceeds its 95th percentile under normal errors, when the errors are Laplace, against the number of people retested. Twenty pairs almost never notice; two hundred notice four times in five.

It needs many more pairs than the variance does. Ten retested people flag a Laplace error 11.8% of the time and twenty 19.8% — barely above the 5% the check fires with normal errors. A hundred flag it 57.4% of the time, two hundred 80.3%, and a thousand every time. Fourth moments are estimated far less precisely than second ones, and the heavy tail the check is looking for is exactly the part of the distribution a few pairs are least likely to sample.

So the two things a study needs from its retest have very different prices. The error variance is cheap: twenty pairs estimate it well enough that the correction keeps nearly all its advantage. The error law is expensive: about two hundred pairs before the study can tell a Laplace error from a normal one four times in five. A study that retests twenty people learns the number it needs and cannot learn whether the formula it puts the number into is the right one.

What a corrected report should state

How many people were retested. Below about twenty, the estimated variance is uncertain by more than a third and the correction’s advantage starts to erode; above it, the retest is not what limits the correction.

Whether the retest differences look normal, and how many there were. With fewer than about two hundred pairs a clean kurtosis check is weak evidence, and the report should say so rather than read silence as normality. With the earlier essay’s emphasis on how far into the tail the selection reached, this is the second condition a correction depends on that a reader cannot see from the corrected numbers.

Both corrections, again. Under normal errors, the gap between Tweedie’s share and the rule’s measures how far the population is from normal. Under heavy-tailed errors it measures, in part, how wrong the correction’s own assumption is, and the two cannot be told apart without the error’s shape.

Counted, integrated and assumed

The shares kept under normal errors are integrals over each population, as in a lead that a heavy tail keeps; under Laplace errors there is no integral as simple, and each is counted on two million people, with a standard error under a thousandth. Every error, win rate and correction is counted over 400 studies of a thousand readings at each retest size, with the retested people taken as the first m of each study. The kurtosis check’s critical value is counted under normal errors at each size, on 4,000 studies, and its power on 4,000 more. Only the kernel estimator is used, only the top one per cent is selected, and only one heavy-tailed error law is tried; an error that is skewed rather than heavy-tailed, or one whose variance depends on the true score, would break the formula in other ways that are not measured here.

Still open: a correction that does not assume the error

The repair the measurements point at is to stop assuming the error is normal. Tweedie’s identity has a general form — the expected true score given a reading is an integral of the true-score density against the error density — and with the retest pairs a study can estimate the error density directly, as the deconvolution of the differences’ distribution, and the true-score density as the deconvolution of the readings’ distribution by the estimated error law. Deconvolution is notoriously slow to converge, and with heavy-tailed errors it is somewhat easier than with normal ones, which is an unusual case where the hard error law helps.

Whether a deconvolution estimate built from a thousand readings and two hundred retest pairs beats both Tweedie’s normal-error form and the linear rule under Laplace errors — and what it costs under normal errors, where the normal form is exact — is the measurement this leaves. The prior the data estimates faced a milder version of the same choice, between a parametric shape and an estimated one, and the winner’s curse is the reason any of it matters: the people who need correcting most are the ones selected because their readings were extreme.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Empirical BayesHeavy tailKernel density estimateKurtosisLaplace distributionMeasurement errorMonte CarloRegression to the meanReliabilitySelection effectTest retestTweedie's formula