A correction built on a retest
Worth reading first: Regression to the mean.
The slope of a density nobody can see corrected a selected group’s readings for regression to the mean by Tweedie’s formula — the reading plus the error variance times the slope of the readings’ log-density — with the slope estimated from the study’s own readings. At a thousand readings the estimated correction beat the linear rule, which shrinks every reading by the correlation between two readings, on most studies from every population that was not normal. It took the error variance as known: 0.4, from a test–retest study somewhere else.
A study usually has to estimate that too, and the obvious source is a retest subsample: some of the people measured twice, so that the differences between their two readings are differences of two errors. Half the mean squared difference estimates the error variance. The same pairs give the correlation between two readings, which is the reliability the linear rule shrinks by. So both corrections can be built from what the study itself has, and the question is how many people it has to measure twice for the more flexible one to stay ahead.
Both corrections from the same retest
The design is the earlier essay’s. Studies of a thousand people, each with a reading that is a true score plus an error of variance 0.4, from populations of true scores that are Laplace, t with four degrees of freedom and uniform, each scaled so that two readings correlate at 0.6. Each study selects its top one per cent — ten people — and predicts the share of their lead over the population mean that a second reading would keep. The target is that share, worked out by integrating over the population, and every estimate is scored by how far from it it lands.
Now the first m people in each study are measured twice. From their pairs the study takes the error variance as half the mean squared difference, and the reliability as the correlation of the two readings. Tweedie’s correction uses the estimated variance with a kernel estimate of the log-density’s slope, the better of the two estimators at this size in the earlier essay. The linear rule uses the estimated reliability. Neither is given a number the study could not have computed.
Twenty people are enough
With the variance known, the correction’s error in the t population is 0.080. With the variance estimated from a thousand retested people it is 0.078, from a hundred 0.084, from twenty 0.103, and from ten 0.133. The linear rule, at the reliability the same people estimate, has an error of 0.278 with twenty retested and 0.187 with a thousand. In the Laplace population the pattern is the same: the estimated correction’s error is 0.108 at twenty retested and the rule’s 0.259.
What makes this work is how little of the correction depends on the variance. The variance estimated from m pairs of normal errors has a relative error of about : at twenty people its standard deviation is 31.5% of its value, at ten 42.9%. But the correction multiplies the variance by the log-density’s slope and adds the product to the reading, so a variance off by a third moves the corrected share by a third of the correction itself. In the t population the correction takes the share from one to about 0.8, so a third of it is about 0.07 — well under half of the 0.18 the rule is wrong by when it assumes the population is normal. The slope estimate, which the retest does not touch, contributes most of the remaining error at every retest size from fifty up, which is why the curves flatten there.
The rule does not get better at the same rate, because its error is not mostly noise. In the t population the share kept is 0.784, and a perfectly estimated reliability is 0.6: the rule is wrong by 0.18 however many people are retested, and estimating the reliability only adds noise on top. More retested people bring the estimated correction down to its floor and leave the rule at its own.
How often the estimate wins
With normal errors the estimate wins on most studies at every retest size. In the t population it wins on 72% of studies with ten people retested, 77% with twenty, 89% with a hundred and 96% with a thousand. In the Laplace population 67%, 73%, 85% and 94%. The uniform population is the hard case, as it was before: its readings have a boundary that a kernel smooths over, so the correction carries a bias of its own, and the estimate wins on 59% of studies at ten retested, 71% at fifty and 81% at a thousand.
So the earlier essay’s conclusion survives the retest. A study of a thousand readings from a population with heavy tails does better correcting its top group by an estimated log-density and a variance from twenty retested people than by the correlation those same twenty people give it. That is a cheap condition, and it is met by most studies that have a reliability figure at all.
An error with a heavy tail
The dashed lines in the same figure tell a different story, and it is the one the earlier essay’s last section feared. Tweedie’s formula is exact only when the error is normal. It says the expected true score is the reading plus the error variance times the log-density’s slope, and that identity comes from differentiating a normal error’s density. With any other error law, the formula is still computable — the study still has readings, a slope and a variance — but it is no longer the regression of the true score on the reading.
The error law changes the target, and the correction does not notice. With Laplace errors of variance 0.4, the top one per cent of the t population keeps 0.666 of its lead, counted on two million people, against 0.784 with normal errors. Tweedie’s correction, given the exact error variance, says 0.830 — too high by 0.164, which is a larger error than the rule’s under normal errors. The linear rule says 0.599, too low by 0.067. In the Laplace population the share kept is 0.643, the correction says 0.807 and the rule 0.600. In the uniform population the share kept falls to 0.328, and the correction overshoots it by 0.374.
The direction is the informative part. A heavy-tailed error produces extreme readings by itself: some of the people in the top one per cent are there because their error was unusually large, not because their true score was. A normal error makes that rare, so a normal-error formula attributes most of an extreme reading to the person; a Laplace error makes it common, so more of the lead is noise and less of it survives a second reading. Tweedie’s formula computed as though the error were normal assigns the noise to the person and over-predicts what the second reading keeps — precisely the error the correction was built to remove, reintroduced by its own assumption.
Where a top group’s lead comes from
The direction can be counted rather than argued, because a simulation knows each person’s true score. On the two million people behind each share, the top one per cent of the t population under normal errors reads 3.192 on average, and 0.689 of that is error; a fifth of them, 20.5%, have an error larger than two of its standard deviations. Under Laplace errors the top one per cent reads 3.330 — the selection reaches a little further, because the readings’ tail is heavier — and 1.111 of it is error, with 40.8% of the group carrying an error past two standard deviations. The second reading draws a fresh error, so whatever part of the lead was error does not come back, and the share kept falls from 0.783 to 0.666.
The same count checks the integral it replaces. Under normal errors the counted share kept in the t population is 0.7834 and the integral is 0.7842, which is two routes to one number that share nothing but the model; under Laplace errors only the count is available, and the agreement under normal errors is the reason to trust it there.
The uniform population shows the mechanism at its starkest. Its true scores stop at 1.342, well short of the reading of 2.187 at which its top one per cent begins, so every selected reading is mostly error even with normal errors: the top group reads 2.444 on average and 1.360 of it is error. Under Laplace errors the group reads 2.739, of which 1.836 is error, and 84.0% of its members owe their place to an error beyond two standard deviations. A second reading keeps a third of the lead. Tweedie’s formula under the normal assumption, which cannot see that the true scores have run out, predicts it keeps seven tenths.
And the comparison reverses. Under Laplace errors the estimated correction beats the rule on 13% of studies in the t population with a thousand retested, and on 6% in the Laplace population. The rule wins not because it is right — it assumes normal true scores and normal errors — but because its two wrong assumptions push in opposite directions here, and Tweedie’s correction keeps one of them. A heavy-tailed population makes a top group keep more than the rule says, and a heavy-tailed error makes it keep less; the rule, which knows about neither, lands between them. That is the cancellation when borrowing goes wrong warned about from the other side, where an assumption that is wrong in a known direction can still beat an estimate that is right about one thing and wrong about another, and it is the reason the plain account of regression to the mean — shrink by the reliability — survives in practice better than its assumptions deserve.
A larger retest does not help the correction
Under Laplace errors the retest size works in the rule’s favour and not the correction’s. In the t population the estimated correction’s error is 0.211 with ten people retested and 0.178 with a thousand — and with the variance known exactly it is also 0.178, so the whole of what remains is the formula’s bias, which no amount of retesting touches. The rule’s error falls from 0.299 with ten retested to 0.076 with a thousand, because its only problem at small retests was the noise in its estimated reliability, and the reliability it converges on, 0.6, happens to lie close to the 0.666 the top group keeps.
So the two corrections cross as the retest grows. With ten retested the estimated correction still lands closer on 43% of studies, because the rule’s noise is then as large as the correction’s bias; by twenty it is 37%, and from fifty retested on the rule wins on most studies. A study under Laplace errors that invests in a large retest makes its simple correction better and its sophisticated one no better at all, which is the reverse of the case under normal errors and the reason the error’s shape, and not its size, is what a retest is most needed to settle.
What the retest pairs can see
The study’s retest pairs carry the information that would warn it. Each pair’s difference is the difference of two errors, so its distribution is the error law’s shape, convolved with itself: normal for normal errors, and with an excess kurtosis of 3⁄2 for Laplace errors. A study can compute the kurtosis of its own differences and ask whether it is larger than normal errors would produce.
It needs many more pairs than the variance does. Ten retested people flag a Laplace error 11.8% of the time and twenty 19.8% — barely above the 5% the check fires with normal errors. A hundred flag it 57.4% of the time, two hundred 80.3%, and a thousand every time. Fourth moments are estimated far less precisely than second ones, and the heavy tail the check is looking for is exactly the part of the distribution a few pairs are least likely to sample.
So the two things a study needs from its retest have very different prices. The error variance is cheap: twenty pairs estimate it well enough that the correction keeps nearly all its advantage. The error law is expensive: about two hundred pairs before the study can tell a Laplace error from a normal one four times in five. A study that retests twenty people learns the number it needs and cannot learn whether the formula it puts the number into is the right one.
What a corrected report should state
How many people were retested. Below about twenty, the estimated variance is uncertain by more than a third and the correction’s advantage starts to erode; above it, the retest is not what limits the correction.
Whether the retest differences look normal, and how many there were. With fewer than about two hundred pairs a clean kurtosis check is weak evidence, and the report should say so rather than read silence as normality. With the earlier essay’s emphasis on how far into the tail the selection reached, this is the second condition a correction depends on that a reader cannot see from the corrected numbers.
Both corrections, again. Under normal errors, the gap between Tweedie’s share and the rule’s measures how far the population is from normal. Under heavy-tailed errors it measures, in part, how wrong the correction’s own assumption is, and the two cannot be told apart without the error’s shape.
Counted, integrated and assumed
The shares kept under normal errors are integrals over each population, as in a lead that a heavy tail keeps; under Laplace errors there is no integral as simple, and each is counted on two million people, with a standard error under a thousandth. Every error, win rate and correction is counted over 400 studies of a thousand readings at each retest size, with the retested people taken as the first m of each study. The kurtosis check’s critical value is counted under normal errors at each size, on 4,000 studies, and its power on 4,000 more. Only the kernel estimator is used, only the top one per cent is selected, and only one heavy-tailed error law is tried; an error that is skewed rather than heavy-tailed, or one whose variance depends on the true score, would break the formula in other ways that are not measured here.
Still open: a correction that does not assume the error
The repair the measurements point at is to stop assuming the error is normal. Tweedie’s identity has a general form — the expected true score given a reading is an integral of the true-score density against the error density — and with the retest pairs a study can estimate the error density directly, as the deconvolution of the differences’ distribution, and the true-score density as the deconvolution of the readings’ distribution by the estimated error law. Deconvolution is notoriously slow to converge, and with heavy-tailed errors it is somewhat easier than with normal ones, which is an unusual case where the hard error law helps.
Whether a deconvolution estimate built from a thousand readings and two hundred retest pairs beats both Tweedie’s normal-error form and the linear rule under Laplace errors — and what it costs under normal errors, where the normal form is exact — is the measurement this leaves. The prior the data estimates faced a milder version of the same choice, between a parametric shape and an estimated one, and the winner’s curse is the reason any of it matters: the people who need correcting most are the ones selected because their readings were extreme.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The shape the estimates keep — both name empirical bayes, kurtosis, monte carlo
- Two analyses of one baseline — both name measurement error, monte carlo, regression to the mean
- A charge that depends on the rule — both name monte carlo, selection effect
- A charge that reads the draw — both name monte carlo, selection effect
- A criterion is a prediction of the hold-out — both name monte carlo, selection effect
- A list is not a rule — both name monte carlo, selection effect
Named objects
A flat tag is an object no other essay names yet.
Empirical BayesHeavy tailKernel density estimateKurtosisLaplace distributionMeasurement errorMonte CarloRegression to the meanReliabilitySelection effectTest retestTweedie's formula