An interval read beside something else

Five times in six

A 95% interval is read as a 95% chance that a replication's estimate will land inside it. With the spread known and a replication of the same size, the chance is 83.42% — five times in six — because both estimates are uncertain. An original that landed two standard errors from the truth captures a replication 48.40% of the time; a replication a tenth the size lands inside 44.54% of the time; and among significant originals from studies with 17% power, 66.94%.

Worth reading first: Twenty intervals and one expected miss.

Twenty intervals and one expected miss makes the point every account of a confidence interval makes: the 95% belongs to the procedure, and it describes how often intervals built this way contain the parameter. The parameter is never seen. What is seen, eventually, is another study of the same thing, and the reading that quietly replaces the procedural one is about that: “if this is right, a replication should land inside the interval nineteen times in twenty.”

That reading is about a different quantity, and it has a different number.

Where a replication's estimate lands against a 95% interval, replication the same sizeThe chance that a 95% interval contains a replication's estimate is 95.00% when the original landed on the truth, 82.99% one standard error away and 48.40% two away. Averaged over where originals land it is 83.42%, and 5.00% of originals capture a replication less than half the time.00.2500.5000.7501-202how far the original estimate landed from the truth, in standard errorschance the replication lands insideaveraged over the originals: 83.42%95%where originals landnormal sampling, spread knownthe 95% is about the parameter, not the next estimate
Fig. 1 The chance a 95% interval contains an equal-sized replication’s estimate, as a function of how far the original estimate landed from the truth. The shaded curve shows where originals land; the dashed rule is the chance averaged over them. The slider changes the replication’s size.

Two uncertain estimates, not one

Take an estimate xˉ1\bar x_1 with a known standard error, and its 95% interval xˉ1±1.96se\bar x_1 \pm 1.96\,\mathrm{se}. A replication of the same size produces xˉ2\bar x_2 with the same standard error, independently.

The interval contains the parameter 95% of the time because xˉ1\bar x_1 is within 1.96 standard errors of it 95% of the time. It contains the replication’s estimate when xˉ2xˉ1<1.96se|\bar x_2 - \bar x_1| < 1.96\,\mathrm{se}, and the difference of two independent estimates has standard error se2\mathrm{se}\sqrt{2}. So

P(capture)  =  2Φ ⁣(1.962)1  =  83.42%P(\text{capture}) \;=\; 2\,\Phi\!\left(\frac{1.96}{\sqrt 2}\right) - 1 \;=\; 83.42\%

Five times in six, not nineteen times in twenty. The shortfall is not a defect of the interval, which does exactly what it promises. It is the second study’s own sampling error, which the parameter does not have and the replication does. A reader who expects a faithful replication to land inside a 95% interval nineteen times in twenty, and sees it land outside one time in six, is liable to conclude that something is wrong with one of the studies when both are behaving normally.

It depends on where the original landed

The average hides a wide spread, because the original estimate itself is somewhere random.

If the original landed exactly on the truth, its interval is centred on the parameter and the replication’s estimate is inside with probability 95%: the replication’s error alone decides it. If the original landed dd standard errors away, its interval is off-centre and the capture probability is

Φ(d+1.96)    Φ(d1.96)\Phi(d + 1.96) \;-\; \Phi(d - 1.96)

original’s error chance the replication lands inside
0 95.00%
1 standard error 82.99%
2 standard errors 48.40%
3 standard errors 14.92%

The hero figure is this curve, with the distribution of where originals land drawn under it. The 83.42% is its average.

One property of the curve is exact and worth pausing on. The capture probability falls below one half at an original error of 1.9599 standard errors — the same 1.96 that defines the interval, to three decimals — so the originals that capture a replication less than half the time are essentially the ones whose own intervals missed the parameter, and they are 5.00% of all originals. An interval that has missed the truth is also an interval that is more likely than not to miss the next estimate. A reader holding one cannot tell which kind it is, which is the same predicament the single interval in front of them always presents, now with a second number attached.

A smaller replication lands inside less often

Replications are frequently not the same size as the study they replicate. A pilot is replicated by a larger trial; a large survey is checked against a small follow-up.

How often a replication's estimate lands inside a 95% interval, by the replication's relative size. A replication a tenth of the size lands inside 44.54% of the time, one of the same size 83.42%, and one ten times the size 93.83%. Only an infinitely large replication reaches 95%.
Fig. 2 The average chance that a replication’s estimate lands inside the original’s 95% interval, against the replication’s size as a multiple of the original’s, on a logarithmic axis. The counted capture when both studies estimate their own spread is drawn further down.

A replication kk times the original’s size has standard error se/k\mathrm{se}/\sqrt{k}, so the difference has standard error se1+1/k\mathrm{se}\sqrt{1 + 1/k} and the capture is 2Φ(1.96/1+1/k)12\Phi\bigl(1.96/\sqrt{1 + 1/k}\bigr) - 1:

replication size captured
a tenth 44.54%
a quarter 61.93%
half 74.22%
the same 83.42%
twice 89.05%
four times 92.04%
ten times 93.83%

A replication a tenth the size of the original lands inside the original’s interval less than half the time, with nothing wrong in either study. A replication ten times the size approaches but does not reach 95%, and only an infinitely large one does, because only then is its estimate the parameter.

The direction matters when a small study is used to check a large one. The large study’s interval is narrow, the small study’s estimate is noisy, and a result “outside the original interval” is the expected outcome rather than a sign of disagreement.

A pilot, and the trial that checks it

The two directions of that asymmetry meet in the most ordinary replication there is: a pilot study followed by a trial four times its size.

Where a replication's estimate lands against a 95% interval, replication 4 times the size. The chance that a 95% interval contains a replication's estimate is 99.99% when the original landed on the truth, 97.26% one standard error away and 46.81% two away. Averaged over where originals land it is 92.04%, and 5.00% of originals capture a replication less than half the time.
Fig. 3 The same capture curve when the replication is four times the size of the original. The curve is higher and flatter at the centre and falls away as steeply once the original is two standard errors off.

When the trial is read against the pilot’s interval, the replication is the larger study and its estimate is nearly the parameter. The capture curve becomes almost an indicator of whether the pilot’s interval contained the truth: 99.99% when the pilot landed on it, 97.26% one standard error away, 46.81% two away and 1.88% three away. Averaged, 92.04%. A trial landing outside the pilot’s interval is then close to a direct statement that the pilot’s interval missed, which is what a reader instinctively takes it to mean — and here the instinct is nearly right.

Read the other way — a pilot-sized study checking the trial’s interval — the curve is low and flat: 67.29% when the trial was exactly right, 61.49% one standard error off, 46.82% two off. The small study’s own noise dominates, and whether it lands inside says little about whether the trial was right. Averaged, 61.93%.

So the same event, “the replication fell outside the original interval”, is strong evidence against the original when the replication is larger and weak evidence when it is smaller. A report of a replication that does not say which was larger cannot be read at all.

The interval a replication would need

If the question is where the replication’s estimate will land, the interval to draw is a prediction interval for it, and the calculation above gives its multiplier directly: 1.961+1/k1.96\sqrt{1 + 1/k} standard errors of the original.

For a same-sized replication the multiplier is 2.772, so the interval that captures a replication’s estimate 95% of the time is 41% wider than the confidence interval. For a replication a quarter the size it is 4.383, more than twice as wide. For one four times the size it is 2.191.

The number 2.772 turns up again in the essay on overlapping intervals, for a related reason: it is the separation at which two same-sized 95% intervals just touch, measured in standard errors of their difference. Both are the same fact about the difference of two independent estimates.

Ten replications of one study

A single replication either lands inside or it does not, and 83.42% says how often. The more informative situation is several replications of the same original — a multi-site programme, a set of registered replications — and a tally of how many fell outside.

The obvious calculation treats each replication as an independent trial with a 16.58% chance of landing outside and reads the tally against a binomial. That calculation is wrong in a way that matters, because the replications share one thing: the original’s interval. If the original landed far from the truth, every replication is likely to land outside it; if it landed close, nearly none will.

How many of 10 replications of one study land outside its 95% interval. The replications share the original, so they miss together. None of 10 lands outside with probability 31.40%, against 16.32% if they were independent; at least half land outside with probability 8.72%, against 1.51%.
Fig. 4 The distribution of how many of ten replications of one study land outside its 95% interval, with the spread known and every replication the same size. Beside each bar, the count a reader would expect if the replications were independent.

The count is a binomial mixed over where the original landed, and the mixing changes both ends of it:

  • none of ten outside: 31.40%, against 16.32% if the replications were independent;
  • at least three outside: 23.48%, against 22.23% — nearly the same in the middle;
  • at least five outside: 8.72%, against 1.51%;
  • at least eight outside: 1.56%, against about two in a hundred thousand.

Replications that share an original miss together. A programme in which eight of ten replications land outside the original interval looks, against the independent binomial, like something that essentially cannot happen by chance. It happens once in sixty-four programmes of perfectly faithful replications of a perfectly genuine effect, and when it does the explanation is not the replications. It is that the original landed badly, which is the thing the programme was set up to find out and which a 95% interval says will happen one time in twenty.

This is the same structure as errors that arrive together when tests share data: the rate per comparison is unchanged, and the distribution of how many occur at once is far wider than independence predicts. The correct reading of a replication tally is against the mixture, and a tally that is extreme against the binomial is much less extreme against it.

Twenty intervals, each with its own replication

The opposite arrangement is the one the twenty intervals already draws: twenty independent studies, each with its own interval, each replicated once.

Twenty 95% intervals for a proportion that really is 0.35. 2 of the twenty miss the true value. The 95% is a property of the procedure across repetitions — no single interval has a 95% chance of anything, because it either contains 0.35 or it does not.
Fig. 5 Twenty 95% intervals from twenty samples, with the true mean marked. Each is a separate original; a replication of each would land inside its interval with probability 83.42%, independently of the others.

Now the captures are independent, because nothing is shared, and the count of replications inside is binomial with twenty trials: 16.68 on average. At least nineteen of twenty land inside only 13.24% of the time, and all twenty 2.66% of the time. A collection of twenty genuine findings that each replicated “inside the original interval” would be a remarkable coincidence, not a vindication; the expected outcome is that three or four do not.

The two arrangements are worth holding side by side because they are so easily confused. Many studies, one replication each: a narrow binomial around 83%. One study, many replications: a mixture that is often all-inside and sometimes mostly-outside. The per-replication rate is the same in both. What differs is whether the misses have a common cause.

When the spread is estimated too

The formulas assume the standard error is known. With the spread estimated in each study the original’s interval is a t interval, wider than ±1.96\pm 1.96 standard errors, and the replication’s estimate varies as before.

How often a replication of the same size lands inside a 95% t interval, counted. With both studies estimating their own spread, a replication's mean lands inside the original's t interval 87.81% of the time at 5 observations, 85.51% of the time at 10 observations, 84.46% of the time at 20 observations, 83.79% of the time at 50 observations. The known-spread value is 83.42%.
Fig. 6 How often a same-sized replication’s mean lands inside the original’s 95% t interval when both studies estimate their own spread, counted over forty thousand pairs of normal samples at each size. The rule marks the known-spread value.
observations in each study captured
5 87.81%
10 85.51%
20 84.46%
50 83.79%

At small samples the capture is higher than the known-spread 83.42%, because the t interval’s extra width is spent on the same event. It falls towards the known-spread value as the samples grow, and by fifty observations it is within half a point. So the five-times-in-six figure is the large-sample value and a slightly pessimistic guide for small studies, and the direction of the gap is the opposite of the one a reader might guess.

What selection does to it

Originals are rarely a random draw. The winner’s curse describes the ones that are published: selected for reaching significance, and therefore selected for having landed on the side of the truth that makes the effect look larger.

How often a replication lands inside the interval of a significant original, by the original studies' power. Only originals that reached p < 0.05 are kept. At 8% power their intervals capture a replication's estimate 52.49% of the time; at 52% power, 83.78%. Unselected intervals capture 83.42% whatever the power.
Fig. 7 Among original studies kept only if they reached p < 0.05, the average chance that each one’s interval contains an unselected replication’s estimate, against the power of the original studies. The dashed rule is the unselected value.
power of the originals captured, among significant originals
7.9% 52.49%
17.0% 66.94%
50.0% 83.42%
80.0% 86.82%
93.8% 85.63%

Among significant results from underpowered studies, a replication lands inside the original interval about half the time. The selected originals are the ones that landed far above the truth — at 17% power the average significant estimate is more than twice the true effect — and the capture curve above says what an interval centred two standard errors off does to the next estimate.

At 50% power the capture is exactly the unselected value, and the reason is a symmetry: the selection keeps the half of originals that landed above the truth, and the capture curve is symmetric about zero. Above 50% power the capture among significant originals is slightly higher than unselected, because the originals discarded are the ones that landed far below the truth, and those were the ones least likely to capture. Selection helps capture when it discards large errors and hurts it when it keeps them, and which it does is set by the power.

The replication-failure literature often reads a replication landing outside the original interval as a sign that the original was a false positive. At 17% power, a third of perfectly genuine replications of perfectly genuine effects land outside, and the numbers above say why.

What to put beside the interval

None of this requires a new method. It requires saying which question an interval is being asked.

If the question is where the parameter is, the confidence interval answers it and its 95% is honest. If the question is where the next study’s estimate will fall, the answer is a prediction interval with a multiplier of 1.961+1/k1.96\sqrt{1 + 1/k} standard errors, and for the commonest case — a replication the same size — that is 2.772 rather than 1.96. If the question is whether a replication that has already happened disagrees with the original, the natural statistic is the difference of the two estimates divided by its own standard error, se1+1/k\mathrm{se}\sqrt{1 + 1/k}, read against a normal. That is an ordinary two-sample comparison, and landing outside the original interval is only a coarse and miscalibrated version of it.

The miscalibration can be stated exactly. For a same-sized replication, landing outside the original’s 95% interval happens 16.58% of the time when nothing at all is different; as a test of “the two studies disagree” it has a size of 16.58%, more than three times the 5% a reader would assume from the word “95%”. For a replication a tenth of the original’s size the size of that test is 55.46%. The comparison of two estimates is already a solved problem, and the capture reading replaces it with a worse one because the interval was sitting on the page. The two-sample comparison also gives a direction and a size to the disagreement, where “outside the interval” gives a yes or a no, and it treats the original and the replication symmetrically, as two measurements of one thing, rather than treating the first as the standard the second is held to.

How this relates to replicating significance

The p-value a replication gets asks a neighbouring question — the chance that an exact replication of a result at p = 0.05 is significant again — and finds it is exactly one half under the plain model. That is a statement about the replication’s estimate crossing a fixed threshold. Capture is a statement about the replication’s estimate falling within a band around the original’s.

The two are complementary readings of the same pair of estimates, and both correct the same intuition. A single study’s result is not a prediction of the next study’s result, because the next study carries its own error, and every summary of the first that is read as a forecast of the second is short by exactly that error.

What the formulas establish, and what they do not

The average capture is the conditional capture averaged over where originals land, and the integral agrees with the closed form to six decimals — 83.42% for a same-sized replication. A count with the spread estimated at a hundred observations in each study agrees with it within its own counting error.

Capture falls below one half exactly where the original’s own interval stops containing the truth, to three decimals, so the share of originals below one half is the interval’s own miss rate.

Replications of one original miss together. The mixed distribution of the count sums to one, and at five, ten and twenty replications it puts more probability than independence does on both extremes — none outside and all outside — which is the signature of a shared cause rather than a property of any particular number of replications.

What does not survive is a 95% interval read as a 95% prediction for a replication’s estimate. For a same-sized replication it is 83.42%, and the interval that does predict at 95% has a multiplier of 2.772.

Not claimed: anything about replications that differ from the original in more than sampling error — a different population, a changed protocol, a true effect that varies between studies. Each of those widens the gap between the two estimates, so every capture above is an upper bound on what a real replication should be expected to achieve, not an estimate of it. The selection figures also assume a single threshold at p < 0.05 and no other filtering, which is the simplest model of publication rather than a measurement of it.

Still open: two intervals read together

Capture compares one interval with one estimate. The commoner reading compares two intervals — a treatment group’s and a control’s, this year’s and last year’s — and asks whether they overlap.

The same difference of two estimates governs that reading, and the multiplier of 2.772 already hints that overlap and a 5% test are not the same thing. How far apart two intervals have to be for their difference to be significant, how that depends on the ratio of their widths and on any correlation between the estimates, and what error bars of other kinds say when they touch, are the questions of two intervals that overlap.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Capture percentageConfidence intervalCoveragePrediction intervalReplicationSampling variationStatistical powerThe winner's curse