When the looking happens
Run a trial. Test the accumulated data five times as it comes in, at the usual 5% threshold, and stop the first time the test passes. Under a true null, that procedure rejects 14.0% of the time.
Look ten times and it rejects 19.1%. Look once and it rejects 4.9%, which is what it should.
What a p-value is relative to
The result seems paradoxical until the definition is read carefully, and then it is inevitable.
A p-value is the probability, under the null, of observing a result at least as extreme as the one obtained. That phrase quantifies over outcomes that could have occurred — and which outcomes could have occurred depends on how the experiment was going to be run.
An experiment fixed at four hundred observations has one set of possible outcomes: the four hundred are collected, one test is performed, and the statistic is compared with 1.96.
An experiment that tests five times and stops early has a much larger set: it could stop at eighty with an extreme early result, or run to four hundred, and the ways of ending up “significant” include five separate opportunities.
So the two experiments have different reference sets, and the same observed data has a different p-value in each. This is not a subtlety about interpretation. It is what the definition says.
Why the rate climbs
Each look is a fresh chance to cross the boundary, and the chances accumulate.
They do not accumulate the way independent tests would — the looks are strongly correlated, because each is computed from data that includes everything the previous one saw. So the rate does not follow 1 − 0.95ᵏ, which would give 23% at five looks. It gives 14%, which is lower.
The measured sequence, from one look to ten: 5.1%, 8.1%, 10.7%, 12.3%, 14.2%, 15.3%, 16.5%, 17.5%.
The increments shrink. The first extra look costs three percentage points; the eighth costs one. That is the correlation working: by the eighth look, most of the data is shared with the seventh, and the additional opportunity is nearly the same opportunity.
Two consequences worth taking.
A couple of looks is not harmless. Two looks nearly doubles the error rate, and two interim analyses is an entirely normal design.
And continuous monitoring is worse than the table suggests. The sequence is still climbing at ten. In the limit of testing after every single observation with no stopping boundary, the rate goes to 100% — a random walk crosses any fixed threshold eventually, so a trial that monitors indefinitely and stops when significant will always stop, whatever the truth.
That last case is not hypothetical. It is what an analyst does when they check the dashboard every morning and stop when the result turns significant.
Ten looks are worth four independent tests
The inflation is easier to size against the case where the looks share nothing.
Ten independent tests at five per cent reject at least once with probability , which is 40.1%. Ten looks at accumulating data reject 19.1% of the time.
Solving gives .
Ten looks at a growing dataset are worth about four independent chances. Consecutive looks share almost all of their rows — the ninth and tenth differ by one tenth of the data — so the overlap removes six of the ten.
It is worth setting that beside the other multiplicity this collection measures. Twenty analyses of one dataset are worth about sixteen and a half independent chances; ten looks at a growing one are worth four. Analyses of the same data overlap far less than looks at the same data, because they differ in what they compute rather than in how much they have.
The trials that cross
The figure makes the mechanism visible in a way the rate cannot.
Forty paths, all with a true effect of zero, all wandering as accumulating averages wander. Most stay within the boundary. Several touch it at some point — usually early, when the statistic is noisiest because it is computed from the least data.
That early crossing is the characteristic failure. At the first interim look the sample is smallest and the statistic is most variable, so it is the look most likely to produce an extreme value by chance. A design that tests early at the nominal level is spending most of its error budget at the moment the data is weakest.
Which is exactly the moment an investigator most wants to look, because stopping early is where the savings are.
Nobody has to cheat
The parallel with twenty analyses of nothing is exact and worth drawing, because both results are about honest procedures.
Every test performed is correct. Every p-value at every look is a valid p-value for the data available at that look. The investigator has not manipulated anything, discarded anything or chosen a favourable analysis.
What has happened is that the reported quantity — the p-value at the moment of stopping — is not the quantity the threshold was calibrated for. The threshold 0.05 is the right cut for a single test with a fixed sample size, and the statistic being compared with it is the minimum p-value across five correlated looks. That statistic has a different null distribution, and 14% of the time it falls below 0.05.
So this is the forking-paths argument with the branching along the time axis rather than across analyses, and it is the sharpest version of it, because the number of branches is known. In the forking-paths case nobody can say how many analyses were available. Here the number of looks is a fact about the protocol, which makes the inflation computable — and, as the next essay shows, correctable.
The asymmetry that makes it worse
One feature of how monitoring works in practice compounds the problem, and it is not in the arithmetic above.
The simulation stops when the boundary is crossed in either direction. Real monitoring is often asymmetric: an investigator watching for a positive result will stop when it appears and continue when it does not.
That is a stopping rule with one absorbing state, and it is more damaging than the two-sided version, because continuing after a discouraging look costs nothing while stopping after an encouraging one locks in the result. The rule “keep going until it is significant, or until the budget runs out” has an error rate bounded only by how much budget there is.
It also usually goes unrecorded. A protocol saying “the analysis was performed on the final sample” is consistent with several earlier looks that did not produce anything, and nothing in the data reveals them.
Which makes this, like multiplicity, a problem whose fixable version requires the plan to have been written down.
What this does not say
Three things, since the result is easy to over-read.
It does not say interim analyses are bad. They are valuable and often ethically required — a trial should stop early if a treatment is clearly harmful or clearly beneficial. The problem is testing at the nominal level, not looking.
It does not say the data is contaminated by being looked at. The observations are exactly what they were. What changes is the reference distribution the final statistic should be compared against, and that is a property of the plan rather than of the data.
And it does not apply to a Bayesian analysis in the same way. A posterior depends on the likelihood of the data observed, and the likelihood does not contain the stopping rule. That is a genuine and often-cited difference between the frameworks, and it comes with its own caveat: a Bayesian who stops when the posterior crosses a threshold has a procedure whose frequentist error rate is inflated exactly as above, even though each individual posterior is correct.
The third point is the one most often stated too strongly. Optional stopping does not invalidate a posterior; it does invalidate any frequentist guarantee attached to a procedure built on one.
Why this belongs on this site
The result is a measurement of something usually argued about, and the measurement is small enough to check.
The claim “a p-value depends on the sampling plan” sounds like philosophy. The demonstration is a simulation of trials where nothing is happening, a count of how many reach significance, and a comparison with the nominal rate — and the site’s gate pins the single-look case at exactly 5% so that the inflation being reported is not an artefact of the machinery.
That pinning is the part worth noticing. A demonstration showing 14% at five looks proves nothing on its own; it might be a broken simulation. A demonstration showing 4.9% at one look and 14.0% at five is evidence, because the first number establishes that the apparatus is calibrated.
Every measurement in this field is built that way: the corrected boundaries are checked against the one-look case where the answer is known to be 1.96, and the inflation is reported beside a control that has to come out right.
The same data, two p-values, worked through
An explicit example, because “the p-value depends on the plan” is the sort of statement that is agreed with and not absorbed.
Two investigators collect identical data: four hundred observations, final z statistic 2.1.
The first fixed the sample at four hundred in advance and analysed once. Their p-value is the two-sided normal tail beyond 2.1, which is 0.036. Significant.
The second planned five interim analyses and this was the last one — the earlier four did not cross. Their reported statistic is the outcome of a five-look procedure, and the relevant question is how often such a procedure produces a crossing under the null. It does so 14% of the time at the nominal boundary, so the threshold that gives a genuine 5% is not 1.96 but 2.41. Their statistic of 2.1 does not reach it. Not significant.
Same four hundred numbers. Same arithmetic on them. Different conclusions, and both correct.
The instinct that something has gone wrong here is worth examining, because the two conclusions are not contradictory. They are answers to different questions: how surprising is this data under a plan that looks once, and how surprising is this data under a plan that looks five times. The second plan had more opportunities to produce a surprising-looking result, so it takes more to be surprising.
That is also why the p-value cannot be computed from the data alone, and why a dataset handed over without its protocol is not sufficient for a frequentist analysis.
What a reader can do about it
The uncomfortable part: the inflation is invisible in the output, and the information needed to detect it is usually absent.
A published result reports a p-value and a sample size. It rarely reports how many times the data was examined during collection, and there is nothing in the numbers that distinguishes a single-look analysis from a five-look one.
Three partial signals.
A trial that stopped early for efficacy. This is usually stated, and it means interim analyses happened. The correct question is then which boundary was used, and a well-run trial will name it — O’Brien–Fleming or Pocock or an alpha-spending function. A trial that stopped early and does not name a boundary has probably not used one.
A sample size that does not match the protocol. Where protocols are public, a study that stopped at a different n than planned either had a stopping rule or had a problem, and either is worth knowing about.
And a p-value just under the threshold in a monitored study. Under a five-look procedure, a nominal 0.04 is not significant at a genuine 5%, so a marginal result from a monitored trial is the case where the distinction changes the conclusion.
None of that is satisfactory. The real fix is on the reporting side, and it is one line: state the number of interim analyses and the boundary used.
Where else the same structure appears
Optional stopping is the named case, and the underlying pattern — a data-dependent decision about when to stop collecting — turns up in places where nobody thinks of it as a stopping rule.
Collecting until the budget runs out is fine. The stopping time is independent of the data.
Collecting until the result is clear is not, and it is common in laboratory work where “clear” is informally assessed and the sample grows until it is achieved.
Adding a replication cohort when the first is marginal is a stopping rule with two looks, and it is standard practice presented as diligence.
Running an A/B test until the dashboard shows significance is continuous monitoring, and it is the industrial version of this problem. A platform that lets anyone check a running test at any time and stop it when a green light appears has built the 100% case into its interface.
That last one is worth dwelling on because the scale is large and the fix is well understood. Sequential methods designed exactly for this exist, and the reason they are not universal in online experimentation is that the naive dashboard is easier to build and its failures are invisible.
The common thread is the same as everywhere in this field: the decision about when to stop is part of the analysis, and treating it as an operational detail is how the error rate goes unaccounted for.
The boundary that would have been right
Since the problem is stated, the shape of the answer is worth previewing.
The inflation happens because five opportunities are each given the full 5%. The repair is to give them 5% between them — raise the boundary so that the chance of crossing it at any of the five looks totals 5%, rather than the chance of crossing at each look being 5%.
For five equally spaced looks the constant boundary that does this is 2.41 rather than 1.96, and the measured error rate under it is 4.8%. The problem is solved, and it is solved by arithmetic that was available before the trial began.
What the raised boundary costs, how to spend the budget unevenly so that stopping early requires more evidence than stopping late, and what a design that can stop early buys in return, are the next essay.
The summary
Testing five times at the nominal level rejects a true null 14% of the time, ten times 19%, and continuously 100%. One look gives exactly 5%, which is what establishes that the rest of the numbers are real.
Nothing in the data changed across those cases. The observations are the same, the arithmetic on them is the same, and every individual p-value is correct. What differs is the set of outcomes the procedure could have produced, and a p-value is a statement about that set.
So a p-value is not a property of a dataset. It is a property of a dataset and a plan, and reporting the first without the second is not enough for a reader to know what the number means.
One more way to see it
A thought experiment that settles the intuition for most people who resist the result.
Imagine a coin-flipping game where the null is that the coin is fair. An investigator flips it, tracking the running excess of heads, and is allowed to stop whenever they like and report the p-value at that moment.
Under the null the running excess is a random walk with no drift. A random walk with no drift is recurrent: it returns to any level infinitely often given enough steps, and it will eventually wander two standard errors from zero, then three, then any bound.
So the investigator who is patient enough will always be able to stop at a significant result, with a genuinely fair coin, without doing anything except waiting. The p-value they report is correct for the data they hold. The procedure that produced it rejects a true null with probability one.
That is the limiting case of the table at the top of this essay, and it is the clearest demonstration that the inflation has nothing to do with dishonesty and everything to do with which set of outcomes the threshold was calibrated against.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
- A block size that changes
- A schedule that reads the mean
- A width promised for a difference
- A width the trial has to stop for
- Choosing n after looking
- Dropping the losers
- Randomising towards the winner
- Spending the error rate
- Stopping on the arms
- Stopping when it is precise enough
- The estimate after the choice
- The experiments that could have happened
- The interval after a stop it chose
- The interval after the choice
- The rule that cannot see the mean
- Two degrees of freedom, one total
- What a design chosen from the data costs
- What a two-arm rule may not pool
- What the blindfold costs
- A simulation that stops when it looks settled
- The effect a stopped trial reports
- The outcomes a trial could have stopped with
- A boundary for giving up
- A look the trend asked for
- The trials that stopped early
- A width rule on skewed outcomes
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The rank is a decision — both name error rate, random walk, sequential testing
- A boundary for giving up — both name error rate, interim analysis
- Counting what is still wandering — both name error rate, random walk
- Estimating how many nulls are true — both name error rate, p-value
- The estimate after the choice — both name error rate, interim analysis
- The experiments that could have happened — both name error rate, p-value
Named objects
A flat tag is an object no other essay names yet.
Error rateInterim analysisOptional stoppingp-valuePosteriorRandom walkSequential testing