Concept

The winner's curse — where it appears

That the largest of many noisy estimates overstates the quantity it is largest about. It applies to a selected model's own score as much as to a selected effect, and the overstatement is exactly what made the winner win.

Named by 12 essays across 10 fields — each of them below, with the objects they name alongside it.

What each rule gives up against an oracle that is arithmetic. Expected squared error of the candidate each rule selects, minus the expected squared error of the best candidate in the table, over 500 draws of 120 rows. Both quantities are closed forms — σ_S²(1 + q/(n − q − 1)) — so the only Monte Carlo here is over which candidate got picked. The hold-out spends half its rows measuring what the criterion computes, and pays 1.8 times as much for it. Schwarz's criterion is worst because it is answering a different question: which candidate contains the truth, rather than which one forecasts best.

A criterion is a prediction of the hold-out

A rolling hold-out spends half the sample measuring what a criterion computes from all of it. Against an oracle that is arithmetic rather than an estimate, the criterion gives up 0.01701 and the hold-out 0.03200 — and the number the hold-out reports for its own winner is optimistic by more than either.

proxy · Order-selection
Where a replication's estimate lands against a 95% interval, replication the same size. The chance that a 95% interval contains a replication's estimate is 95.00% when the original landed on the truth, 82.99% one standard error away and 48.40% two away. Averaged over where originals land it is 83.42%, and 5.00% of originals capture a replication less than half the time.

Five times in six

A 95% interval is read as a 95% chance that a replication's estimate will land inside it. With the spread known and a replication of the same size, the chance is 83.42% — five times in six — because both estimates are uncertain. An original that landed two standard errors from the truth captures a replication 48.40% of the time; a replication a tenth the size lands inside 44.54% of the time; and among significant originals from studies with 17% power, 66.94%.

alongside · Repetition
12,000 studies of a real effect of 0.3, n = 16. Power is 21%. The studies that reached significance report a mean effect of 0.621 — 2.07 times the truth. Every one of them is honest; the selection did the inflating.

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

testing · Curse
Forty O'Brien–Fleming trials at a true effect of 0.16, with the boundary written as an effect. The dashed line is the smallest effect a trial can report and still stop at each look: 0.510 at 80 observations, 0.255 at 160 observations, 0.170 at 240 observations, 0.128 at 320 observations, 0.102 at 400 observations. The true effect is 0.16, so at 3 of the five looks a trial cannot stop without reporting more than it. 29 of these forty trials stop before the last look, each marked where it stopped.

The effect a stopped trial reports

An O'Brien–Fleming trial at 88.45% power holds its error rate exactly and reports an effect 9.6% too large on average. The 11.39% of trials that stop at the second look report 1.83 times the truth, the ones that cross at the last look report 0.80 times it, and pooling every trial by its size gives the truth back to the last digit.

sequential · Stopping
What the family-wise correction does to the effect it lets through. At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one.

The correction that makes the estimate worse

Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

paths · Forking
Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

alongside · Repetition
Three estimates of the same effect, and the honest one is the worst. 8,000 trials in which every arm has an effect of exactly 0. The arm carried forward is reported by its first stage at +0.187 above the truth — it was chosen for being ahead — and by its second stage at -0.0032, which is unbiased by construction because the selection could not see it. The combination that gets published is at 0.123, exactly the share of the first stage's bias the first stage contributes. Root mean squared error: 0.240, 0.261, 0.182 — the unbiased estimate is the least accurate of the three.

The estimate after the choice

An arm chosen for being ahead is ahead by more than it should be, and the trial then publishes the average of the stage that chose it and the stage that did not. The unbiased estimate is the one built from a third of the data — and it is the least accurate of the three.

adaptive · Curse
The p-value of a study with 80% power, twenty thousand times. Twenty thousand two-sided z-tests, each on 25 observations whose true mean is 0.5603 standard deviations from the null, a noncentrality of 2.802. The bars are the counted share of p-values in bins a quarter of a power of ten wide, with the leftmost bin holding everything smaller; the line is the closed form. The middle eighty per cent of the p-values runs from 4.4×10⁻⁵ to 0.13, 3.46 orders of magnitude, the median is 0.0051, and 80.0% fall below 0.05, which is what the power means.

The p-value a replication gets

Under a true null a p-value is flat. Under a real effect its distribution is closed form and wide — a study with 80% power returns anything from 4.4×10⁻⁵ to 0.13 in eight runs of ten — and the chance that an exact replication of a p = 0.05 result is significant again is exactly one half, under both of the models people use without naming them.

testing · Uniformity
The two standardised means, and the shape of Fieller's set for their ratio at a = 3, d = 1. Each dot is one pair (zx, zy) drawn around (3, 1). Outside the horizontal band |zy| > 1.96 Fieller's set is a bounded interval, with probability 17.01%; inside the band and outside the disc of radius 1.96 it is everything outside an interval, 75.03%; inside the disc it is the whole line, 7.96%. On 40,000 counted draws Fieller covers ρ = 3.00 95.21% of the time and the delta interval 82.48%.

A ratio whose interval has to be the whole line

The delta interval for a ratio of two means covers 95.61% when the denominator is eight standard errors from zero and 1.10% at a ten-thousandth of one, and ten times as wide it still covers only 3.48%. Linearising is not the fault. Gleser and Hwang proved that every interval that is always finite fails the same way, so an interval that keeps its promise has to be the whole line some of the time.

normal · Clt
Every finding Benjamini–Hochberg made in thirty families, with its interval, at real effects of 2. 85 findings, sorted by their estimate. 12 of their ordinary 95% intervals miss the true effect, every one of them on the far side; 6 of the wider false-coverage-rate intervals miss.

Intervals for the findings

Benjamini–Hochberg's findings usually go out each with its ordinary 95% interval. With ten real effects of two standard errors among twenty tests, 11.59% of those intervals miss their effect, every miss on the far side, and the interval around the most prominent finding covers 72.36% of the time — 2.38% when the effects are one standard error. Intervals widened for the number of findings hold the share that miss under 5%.

multiplicity · Multiplicity
What a confirmation run at the chosen setting would find. The true optimum is worth 62.348. At σ = 2 the fit predicts 62.679 at the setting it recommends and the truth there is 61.821 — a gap of 0.858, which is 0.72 of the prediction's own standard error. The setting itself gives up 0.527 against the best available.

The run that confirms it

The setting a response-surface analysis recommends was chosen because the fitted surface was highest there, so the height the fit predicts at it is a maximum over a random field. At twice the noise the fit predicts 0.858 more than is there — 0.72 of the prediction's own standard error — and the gap is not noise, it is the selection.

surface · Optimum
How fast a gap has to close before a sample can see it close. The power of the test against the half-life of a disagreement, at 100, 200, 400 observations, each read against its own simulated critical value. Every pair in every reading is genuinely tied together, so a non-rejection is a miss. At 200 observations a gap that halves in 3 steps is found 99.9% of the time and one that halves in 12 steps is found 15.3% of the time — and by 35 steps the reading is 6.1%, which is the test's own size. Beyond that the curves are flat because there is nothing left to detect with.

How slow a return a sample can see

At two hundred observations the test finds a gap that halves in five steps four times in five, one that halves in eight 37.3% of the time, and one that halves in fifty 4.95% of the time — which is the rate at which it finds pairs with no mechanism at all. The boundary moves with the sample, not with its square root.

timeseries · Spurious

Named alongside it

The objects these essays reach for when they reach for this one.

Statistical powerMonte CarloSelection biasConfidence intervalCoverageEffect sizeMean squared errorMultiple comparisonsSample sizeSelection effectClosed formInterim analysis

All concepts