Five times in six
Worth reading first: Twenty intervals and one expected miss.
Twenty intervals and one expected miss makes the point every account of a confidence interval makes: the 95% belongs to the procedure, and it describes how often intervals built this way contain the parameter. The parameter is never seen. What is seen, eventually, is another study of the same thing, and the reading that quietly replaces the procedural one is about that: “if this is right, a replication should land inside the interval nineteen times in twenty.”
That reading is about a different quantity, and it has a different number.
Two uncertain estimates, not one
Take an estimate with a known standard error, and its 95% interval . A replication of the same size produces with the same standard error, independently.
The interval contains the parameter 95% of the time because is within 1.96 standard errors of it 95% of the time. It contains the replication’s estimate when , and the difference of two independent estimates has standard error . So
Five times in six, not nineteen times in twenty. The shortfall is not a defect of the interval, which does exactly what it promises. It is the second study’s own sampling error, which the parameter does not have and the replication does. A reader who expects a faithful replication to land inside a 95% interval nineteen times in twenty, and sees it land outside one time in six, is liable to conclude that something is wrong with one of the studies when both are behaving normally.
It depends on where the original landed
The average hides a wide spread, because the original estimate itself is somewhere random.
If the original landed exactly on the truth, its interval is centred on the parameter and the replication’s estimate is inside with probability 95%: the replication’s error alone decides it. If the original landed standard errors away, its interval is off-centre and the capture probability is
| original’s error | chance the replication lands inside |
|---|---|
| 0 | 95.00% |
| 1 standard error | 82.99% |
| 2 standard errors | 48.40% |
| 3 standard errors | 14.92% |
The hero figure is this curve, with the distribution of where originals land drawn under it. The 83.42% is its average.
One property of the curve is exact and worth pausing on. The capture probability falls below one half at an original error of 1.9599 standard errors — the same 1.96 that defines the interval, to three decimals — so the originals that capture a replication less than half the time are essentially the ones whose own intervals missed the parameter, and they are 5.00% of all originals. An interval that has missed the truth is also an interval that is more likely than not to miss the next estimate. A reader holding one cannot tell which kind it is, which is the same predicament the single interval in front of them always presents, now with a second number attached.
A smaller replication lands inside less often
Replications are frequently not the same size as the study they replicate. A pilot is replicated by a larger trial; a large survey is checked against a small follow-up.
A replication times the original’s size has standard error , so the difference has standard error and the capture is :
| replication size | captured |
|---|---|
| a tenth | 44.54% |
| a quarter | 61.93% |
| half | 74.22% |
| the same | 83.42% |
| twice | 89.05% |
| four times | 92.04% |
| ten times | 93.83% |
A replication a tenth the size of the original lands inside the original’s interval less than half the time, with nothing wrong in either study. A replication ten times the size approaches but does not reach 95%, and only an infinitely large one does, because only then is its estimate the parameter.
The direction matters when a small study is used to check a large one. The large study’s interval is narrow, the small study’s estimate is noisy, and a result “outside the original interval” is the expected outcome rather than a sign of disagreement.
A pilot, and the trial that checks it
The two directions of that asymmetry meet in the most ordinary replication there is: a pilot study followed by a trial four times its size.
When the trial is read against the pilot’s interval, the replication is the larger study and its estimate is nearly the parameter. The capture curve becomes almost an indicator of whether the pilot’s interval contained the truth: 99.99% when the pilot landed on it, 97.26% one standard error away, 46.81% two away and 1.88% three away. Averaged, 92.04%. A trial landing outside the pilot’s interval is then close to a direct statement that the pilot’s interval missed, which is what a reader instinctively takes it to mean — and here the instinct is nearly right.
Read the other way — a pilot-sized study checking the trial’s interval — the curve is low and flat: 67.29% when the trial was exactly right, 61.49% one standard error off, 46.82% two off. The small study’s own noise dominates, and whether it lands inside says little about whether the trial was right. Averaged, 61.93%.
So the same event, “the replication fell outside the original interval”, is strong evidence against the original when the replication is larger and weak evidence when it is smaller. A report of a replication that does not say which was larger cannot be read at all.
The interval a replication would need
If the question is where the replication’s estimate will land, the interval to draw is a prediction interval for it, and the calculation above gives its multiplier directly: standard errors of the original.
For a same-sized replication the multiplier is 2.772, so the interval that captures a replication’s estimate 95% of the time is 41% wider than the confidence interval. For a replication a quarter the size it is 4.383, more than twice as wide. For one four times the size it is 2.191.
The number 2.772 turns up again in the essay on overlapping intervals, for a related reason: it is the separation at which two same-sized 95% intervals just touch, measured in standard errors of their difference. Both are the same fact about the difference of two independent estimates.
Ten replications of one study
A single replication either lands inside or it does not, and 83.42% says how often. The more informative situation is several replications of the same original — a multi-site programme, a set of registered replications — and a tally of how many fell outside.
The obvious calculation treats each replication as an independent trial with a 16.58% chance of landing outside and reads the tally against a binomial. That calculation is wrong in a way that matters, because the replications share one thing: the original’s interval. If the original landed far from the truth, every replication is likely to land outside it; if it landed close, nearly none will.
The count is a binomial mixed over where the original landed, and the mixing changes both ends of it:
- none of ten outside: 31.40%, against 16.32% if the replications were independent;
- at least three outside: 23.48%, against 22.23% — nearly the same in the middle;
- at least five outside: 8.72%, against 1.51%;
- at least eight outside: 1.56%, against about two in a hundred thousand.
Replications that share an original miss together. A programme in which eight of ten replications land outside the original interval looks, against the independent binomial, like something that essentially cannot happen by chance. It happens once in sixty-four programmes of perfectly faithful replications of a perfectly genuine effect, and when it does the explanation is not the replications. It is that the original landed badly, which is the thing the programme was set up to find out and which a 95% interval says will happen one time in twenty.
This is the same structure as errors that arrive together when tests share data: the rate per comparison is unchanged, and the distribution of how many occur at once is far wider than independence predicts. The correct reading of a replication tally is against the mixture, and a tally that is extreme against the binomial is much less extreme against it.
Twenty intervals, each with its own replication
The opposite arrangement is the one the twenty intervals already draws: twenty independent studies, each with its own interval, each replicated once.
Now the captures are independent, because nothing is shared, and the count of replications inside is binomial with twenty trials: 16.68 on average. At least nineteen of twenty land inside only 13.24% of the time, and all twenty 2.66% of the time. A collection of twenty genuine findings that each replicated “inside the original interval” would be a remarkable coincidence, not a vindication; the expected outcome is that three or four do not.
The two arrangements are worth holding side by side because they are so easily confused. Many studies, one replication each: a narrow binomial around 83%. One study, many replications: a mixture that is often all-inside and sometimes mostly-outside. The per-replication rate is the same in both. What differs is whether the misses have a common cause.
When the spread is estimated too
The formulas assume the standard error is known. With the spread estimated in each study the original’s interval is a t interval, wider than standard errors, and the replication’s estimate varies as before.
| observations in each study | captured |
|---|---|
| 5 | 87.81% |
| 10 | 85.51% |
| 20 | 84.46% |
| 50 | 83.79% |
At small samples the capture is higher than the known-spread 83.42%, because the t interval’s extra width is spent on the same event. It falls towards the known-spread value as the samples grow, and by fifty observations it is within half a point. So the five-times-in-six figure is the large-sample value and a slightly pessimistic guide for small studies, and the direction of the gap is the opposite of the one a reader might guess.
What selection does to it
Originals are rarely a random draw. The winner’s curse describes the ones that are published: selected for reaching significance, and therefore selected for having landed on the side of the truth that makes the effect look larger.
| power of the originals | captured, among significant originals |
|---|---|
| 7.9% | 52.49% |
| 17.0% | 66.94% |
| 50.0% | 83.42% |
| 80.0% | 86.82% |
| 93.8% | 85.63% |
Among significant results from underpowered studies, a replication lands inside the original interval about half the time. The selected originals are the ones that landed far above the truth — at 17% power the average significant estimate is more than twice the true effect — and the capture curve above says what an interval centred two standard errors off does to the next estimate.
At 50% power the capture is exactly the unselected value, and the reason is a symmetry: the selection keeps the half of originals that landed above the truth, and the capture curve is symmetric about zero. Above 50% power the capture among significant originals is slightly higher than unselected, because the originals discarded are the ones that landed far below the truth, and those were the ones least likely to capture. Selection helps capture when it discards large errors and hurts it when it keeps them, and which it does is set by the power.
The replication-failure literature often reads a replication landing outside the original interval as a sign that the original was a false positive. At 17% power, a third of perfectly genuine replications of perfectly genuine effects land outside, and the numbers above say why.
What to put beside the interval
None of this requires a new method. It requires saying which question an interval is being asked.
If the question is where the parameter is, the confidence interval answers it and its 95% is honest. If the question is where the next study’s estimate will fall, the answer is a prediction interval with a multiplier of standard errors, and for the commonest case — a replication the same size — that is 2.772 rather than 1.96. If the question is whether a replication that has already happened disagrees with the original, the natural statistic is the difference of the two estimates divided by its own standard error, , read against a normal. That is an ordinary two-sample comparison, and landing outside the original interval is only a coarse and miscalibrated version of it.
The miscalibration can be stated exactly. For a same-sized replication, landing outside the original’s 95% interval happens 16.58% of the time when nothing at all is different; as a test of “the two studies disagree” it has a size of 16.58%, more than three times the 5% a reader would assume from the word “95%”. For a replication a tenth of the original’s size the size of that test is 55.46%. The comparison of two estimates is already a solved problem, and the capture reading replaces it with a worse one because the interval was sitting on the page. The two-sample comparison also gives a direction and a size to the disagreement, where “outside the interval” gives a yes or a no, and it treats the original and the replication symmetrically, as two measurements of one thing, rather than treating the first as the standard the second is held to.
How this relates to replicating significance
The p-value a replication gets asks a neighbouring question — the chance that an exact replication of a result at p = 0.05 is significant again — and finds it is exactly one half under the plain model. That is a statement about the replication’s estimate crossing a fixed threshold. Capture is a statement about the replication’s estimate falling within a band around the original’s.
The two are complementary readings of the same pair of estimates, and both correct the same intuition. A single study’s result is not a prediction of the next study’s result, because the next study carries its own error, and every summary of the first that is read as a forecast of the second is short by exactly that error.
What the formulas establish, and what they do not
The average capture is the conditional capture averaged over where originals land, and the integral agrees with the closed form to six decimals — 83.42% for a same-sized replication. A count with the spread estimated at a hundred observations in each study agrees with it within its own counting error.
Capture falls below one half exactly where the original’s own interval stops containing the truth, to three decimals, so the share of originals below one half is the interval’s own miss rate.
Replications of one original miss together. The mixed distribution of the count sums to one, and at five, ten and twenty replications it puts more probability than independence does on both extremes — none outside and all outside — which is the signature of a shared cause rather than a property of any particular number of replications.
What does not survive is a 95% interval read as a 95% prediction for a replication’s estimate. For a same-sized replication it is 83.42%, and the interval that does predict at 95% has a multiplier of 2.772.
Not claimed: anything about replications that differ from the original in more than sampling error — a different population, a changed protocol, a true effect that varies between studies. Each of those widens the gap between the two estimates, so every capture above is an upper bound on what a real replication should be expected to achieve, not an estimate of it. The selection figures also assume a single threshold at p < 0.05 and no other filtering, which is the simplest model of publication rather than a measurement of it.
Still open: two intervals read together
Capture compares one interval with one estimate. The commoner reading compares two intervals — a treatment group’s and a control’s, this year’s and last year’s — and asks whether they overlap.
The same difference of two estimates governs that reading, and the multiplier of 2.772 already hints that overlap and a 5% test are not the same thing. How far apart two intervals have to be for their difference to be significant, how that depends on the ratio of their widths and on any correlation between the estimates, and what error bars of other kinds say when they touch, are the questions of two intervals that overlap.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name confidence interval, coverage, statistical power
- A ratio whose interval has to be the whole line — both name confidence interval, coverage, the winner's curse
- A simulation that stops when it looks settled — both name confidence interval, coverage, statistical power
- A tenth as wide, and both of them right — both name confidence interval, coverage, prediction interval
- Intervals for the findings — both name confidence interval, coverage, the winner's curse
- Right for the wrong reason — both name confidence interval, coverage, statistical power
Named objects
A flat tag is an object no other essay names yet.
Capture percentageConfidence intervalCoveragePrediction intervalReplicationSampling variationStatistical powerThe winner's curse