The slope of a density nobody can see
Worth reading first: Regression to the mean.
A lead that a heavy tail keeps showed that the correlation between two readings decides how much of a selected group’s lead survives only when the true scores are normal. Four populations with readings correlated at exactly 0.6 kept 60%, 76.1%, 78.4% and 44.3% of their top one per cent’s lead, and the rule that predicts 60% for all of them removes real lead from the heavy-tailed ones and leaves noise in the bounded one. The exact regression is Tweedie’s formula: the expected true score behind a reading is the reading plus the error variance times the slope of the log-density of readings at that reading.
That formula has a property the essay did not use. It needs the density of the readings, and the readings are what a study has. A study with enough of them can estimate the density, differentiate its logarithm, and correct its own top group without knowing what the population of true scores looks like — the idea behind the empirical-Bayes corrections a prior estimated from the data introduced for a single variance. The catch is where it matters. The slope of a log-density is least determined in the tail, and the top one per cent of a thousand readings is ten points. This essay measures whether the idea survives a study of a realistic size, in the same four populations, with the error variance of 0.4 taken as known from a test–retest study.
Two ways to estimate a slope from readings
Two estimators are standard and both are measured. The first is a Gaussian kernel density with Silverman’s rule for its bandwidth, whose log-slope at a reading is the derivative of the estimate over the estimate. The second is Efron’s log-spline, in the form of Lindsey’s method: the readings are binned into sixty equal cells, the counts are fitted by a Poisson regression on a polynomial of degree five in the reading, and the log-density’s slope is the polynomial’s derivative. Either slope, times the error variance and added to the reading, gives the corrected expected true score, and averaging those over a study’s top ten readings gives the share of their lead the correction says a second reading keeps.
The two have different weaknesses, and a large sample shows them before any tail does. On 200,000 readings from a standard normal, where the log-density’s slope is exactly minus the reading, the log-spline reads −0.999, −1.975 and −2.931 at readings of one, two and three. The kernel reads −1.082, −1.740 and −2.557. A kernel estimate is the true density smoothed by the kernel, so its tails fall more slowly than the density’s and its log-slope is too shallow exactly where the correction is applied. A polynomial has no such smoothing bias, and pays instead in how wildly its derivative can move near the edge of the data.
Twenty studies of a thousand readings
The opening figure draws what one study’s estimate looks like, twenty times, for the t population. In the middle of the readings the twenty curves bunch around the integral. Past two standard deviations they fan out, and the top ten readings of a typical study sit well inside the fan. Every curve is an honest estimate from a thousand readings; the fan is the uncertainty in the slope of a density where there are almost no readings to fix it.
Averaged over four hundred studies, the corrected share lands near the truth for the two heavy-tailed populations. For the t population, where the integral is 0.7842 and the linear rule says 0.6, the log-spline averages 0.8190 with a spread of 0.0806 and the kernel 0.7914 with a spread of 0.0853. For the Laplace, where the integral is 0.7612, they average 0.8417 and 0.7883, with spreads of 0.0833 and 0.0735. Both lean slightly high, towards keeping more of the lead than the truth, and both are far closer to the truth than the rule’s 0.6. The log-spline’s lean is not fixed, though. In the Laplace population its average falls from 0.8417 at a thousand readings to 0.7806 at 4,000 and 0.7425 at 16,000, crossing the integral on the way, which is the first sign of the polynomial trouble a larger study makes plain.
Counted study by study, an estimate is closer to the integral than the linear rule is on 95.3% of studies for the log-spline and 98.0% for the kernel in the t population, and on 84.0% and 96.8% in the Laplace. The rule is wrong there by 0.1842 and 0.1612, which is large enough that a noisy estimate centred on the truth usually beats it.
A bounded population and a kernel that smooths its edge
The uniform population is where the kernel’s smoothing bias shows. Its true scores stop at ±1.342, so a reading far past that edge is mostly error and keeps little of its lead: the integral for the top one per cent is 0.4432. The kernel smooths the sharp shoulder of the readings’ density into a gentler slope and averages 0.5298 — more than halfway back to the rule’s 0.6. The log-spline, with no smoothing, averages 0.4607, centred on the truth but with a spread of 0.1289, the widest of any population.
So neither estimator is comfortable at the edge of a bounded population. The kernel beats the rule on 81.5% of studies despite its bias, because its spread is small; the log-spline on 75.0% despite being centred, because its spread is large. Their errors against the integral are nearly equal, 0.1187 and 0.1299, and both are smaller than the rule’s 0.1568 — by a quarter for the kernel and a sixth for the log-spline.
What an estimate costs where the rule was right
For a normal population the linear rule is exact, and an estimated correction can only add error. With a thousand readings the log-spline’s share kept is off the truth by an error of 0.1157 and the kernel’s by 0.0902. That is the price of not assuming the shape: a study that estimates its correction from its own readings when the population was normal reports its top group’s expected second reading with the same order of error a heavy-tailed study gains back by estimating.
The price is also the size of the study’s own noise. The ten people a study selects from a thousand readings have second readings of their own, and across studies the share they actually keep has a spread of 0.0989 around the 0.6 the rule predicts. An estimated correction’s error at a thousand readings is as large as the variation between the ten people one study happens to select, which is a sober measure of how much a single study can know about its own tail.
What more readings buy, and a polynomial that stops following
With more readings the kernel improves steadily in every population. For the t its error falls from 0.1360 at 250 readings to 0.0855 at a thousand, 0.0455 at 4,000 and 0.0230 at 16,000; the bandwidth narrows, the smoothing bias shrinks, and the top one per cent holds more points.
The log-spline does not. For the t its error falls from 0.2754 at 250 readings to 0.0877 at a thousand and 0.0499 at 4,000, and then rises to 0.0912 at 16,000, where its average share kept is 0.6975 against the integral’s 0.7842. A polynomial of fixed degree is fitted over the whole range of the readings, and a larger study stretches that range further into the tail. A power tail’s log-density flattens like a logarithm, which no polynomial of degree five follows over a widening range; the fit bends to match the bulk and misses the tail it was estimated for. The Laplace log-spline turns the same way at 16,000 readings, to 0.7425 below its integral. Efron’s own recommendation is a spline whose flexibility grows with the data, and the counts say why: a fixed degree is a model of the tail that more data eventually contradicts.
When an estimate beats the correlation
Size decides whether estimating is worth it, and the crossing is early for the kernel and late for the log-spline. With 250 readings — a top one per cent of three people — the log-spline beats the linear rule on only 36.3% of studies in the Laplace population, 40.0% in the t and 32.0% in the uniform, and its error there, 0.2904, 0.2754 and 0.3402, is larger than the rule’s. The kernel already beats the rule on 81.8%, 86.5% and 62.0%. By a thousand readings both beat it on most studies, and by 4,000 the log-spline beats it on 100.0%, 100.0% and 98.8%.
The ordering is the one the winner’s curse and when borrowing goes wrong set up from the other side. A rule that assumes a shape is exactly right when the shape holds and wrong by a fixed amount when it does not; an estimate that assumes nothing is never exactly right and is wrong by an amount that shrinks with the data. Which is better depends on how wrong the assumed shape is against how much data there is, and at a thousand readings, for populations as far from normal as a Laplace or a bounded uniform, the estimate has already won most of the time.
How deep a selection has to reach before estimating pays
Every number so far is for a top one per cent, where the four populations differ most. A study that selects more broadly reads further from the tail, where the regressions nearly agree, and there the comparison runs the other way. For the top tenth of a thousand readings the integral is 0.6254 for the Laplace population, 0.6016 for the t and 0.5478 for the uniform, so the linear rule is wrong by 0.0254, 0.0016 and 0.0522 — all but exact for the t.
The estimates are more precise there, since a top tenth is a hundred readings rather than ten, and the log-spline’s error falls to 0.0340 in the Laplace population and 0.0382 in the t. That precision is not enough against a rule that is nearly right. The log-spline beats the rule on 53.8% of studies in the Laplace population and 3.5% in the t; the kernel on 43.0% and 3.0%. Only the uniform, whose top tenth already keeps noticeably less than the correlation, still rewards estimating, on 76.3% of studies for the log-spline and 56.3% for the kernel. At the top twentieth the balance has mostly turned back, and the log-spline beats the rule on 91.8%, 74.0% and 85.0% of studies.
So the choice is not between two methods but between a depth of selection and a size of study. A broad selection from a moderate study is where the correlation’s rule belongs, because the population’s shape has not yet had room to matter; a narrow selection is where the shape decides the answer, and where a study has to buy the shape with readings. The same trade decides how far a hierarchical model should pull a group: a fixed share is right near the centre and wrong in the tails, and whether estimating the tail is worth its noise depends on how far out the group sits.
The error in the units a report uses
A share kept is a ratio, and a report states the second reading it expects. The top one per cent of the t population reads 3.19 standard deviations on average at selection. There the linear rule strips 0.5881 standard deviations of real lead from the group, while the kernel’s estimated correction is off by about 0.27 on a typical study of a thousand readings and by 0.02 on average — less than half as wrong on a single study, and nearly unbiased across them. In the Laplace population, reading 3.06, the rule strips 0.4925 and the kernel is off by about 0.24.
In the uniform population, reading 2.44, the rule leaves 0.3832 standard deviations of noise in the group, and the kernel’s smoothing bias alone is 0.21 of them, so an estimated correction there removes a little under half of the rule’s error and replaces the rest with a systematic error of the same sign. And where the population is normal, reading 2.67, the rule is exact and the kernel’s estimate is off by about 0.24 standard deviations on a typical study: a departure from ordinary regression to the mean of a quarter of a standard deviation, which a study estimating its own tail would report where there is none. Setting an assumed correction beside an estimated one is two routes that share nothing applied to a single study, and the count of readings behind the estimate decides whether a gap between them is a finding or the estimator’s noise.
What a study that corrects its extremes should report
The correction’s source. A share kept of 0.6 from the correlation is a claim that the population is normal. A share kept from an estimated log-density is a claim that the estimator is right in the tail, which at a thousand readings is a claim with an error of about a tenth.
The number of readings behind the tail. The top one per cent of 250 readings is three people, and at that size the more flexible estimator is worse than the rule it replaces. The enrolment that selected them sets how far into the tail the correction is applied, and the count of readings sets whether the tail can be estimated at all.
How deep the selection reached. For the t population the linear rule beat both estimates on nearly every study of a thousand readings at the top tenth and lost to both on nearly every study at the top hundredth. A correction’s worth belongs to the selection it is applied to, so a report that states one without the other has left out the condition that decides which way the comparison goes.
Both corrections, when they disagree. An estimated share kept well away from the correlation — 0.82 against 0.6 — says the population is not normal in the direction of keeping lead, and the gap between the two is the size of the decision the analysis is making for its reader.
Proved, computed and counted
Proved. Tweedie’s formula is exact with normal error whatever the population, and a correction built on any estimate of the readings’ log-density is Tweedie’s formula with that estimate substituted.
Computed without a sample. The share each population’s top one per cent keeps, by integration over the true score, which every estimate is scored against.
Counted. Every estimated share, spread, error and win rate, over four hundred studies at 250, 1,000 and 4,000 readings and a hundred at 16,000 in each population, and the large-sample check of both estimators on 200,000 normal readings.
Particular to these choices. One bandwidth rule, one polynomial degree and sixty bins, a known error variance, and a top one per cent. A quantile plot of forty points cannot tell these populations apart; four hundred studies of a thousand can tell the estimators apart, which is a different and easier question.
Still open: an error variance the study also has to estimate
Every correction here multiplies the estimated slope by an error variance taken as known. A study usually estimates that too, from a test–retest subsample, and the correction is proportional to it: an error variance overstated by a tenth overstates the lead given back by a tenth, at every reading, in every population. The two estimates compound — a slope uncertain in the tail times a variance uncertain from a small retest — and a heavy tail in the error rather than in the true scores breaks the formula outright, since Tweedie’s form assumes the error is normal. How large a test–retest sample a thousand-reading study needs before its estimated correction still beats the correlation’s, and what a heavy-tailed error does to both, is the measurement this one leaves.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two analyses of one baseline — both name closed form, measurement error, monte carlo, regression to the mean, test–retest reliability
- Adjusting for a shadow — both name closed form, measurement error, monte carlo, test–retest reliability
- A lag the sample has less of — both name closed form, monte carlo, shrinkage
- A simulation that stops when it looks settled — both name closed form, monte carlo, selection bias
- An ordering that depends on the rule — both name bias-variance, closed form, monte carlo
- Leaving each row out of its own first stage — both name closed form, heavy tail, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Bias-varianceClosed formHeavy tailMeasurement errorMonte CarloRegression to the meanSelection biasShrinkageTest–retest reliabilityTweedie's formula