Concept

Optional stopping — where it appears

Ending collection when the result looks convincing, which inflates the error rate because the moment of stopping is chosen from the data. The repair is to stop on a quantity independent of the one being reported, which for a normal sample is available exactly.

Named by 10 essays across 7 fields — each of them below, with the objects they name alongside it.

The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%.

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

stop · Stopping
What the interim sees, at an effect of 1. The same 60 observations, estimated two ways. Keeping the arms separate gives 0.995, which is σ. Pooling them without separating the arms — the price of staying blind to the comparison — gives 1.114, against the identity √(1 + Δ²/4σ²) = 1.118. The sample size is proportional to the variance, so a blinded design at this effect asks for 25% more units than it needs, and it does so systematically rather than by chance.

Choosing n after looking

Re-estimating the sample size from an interim is the one adaptation with a defence, and the defence is exactly what it costs: an analyst kept blind to the arms measures a spread that contains the effect, so the design overshoots by 1 + Δ²/4σ². Re-estimating the effect instead breaks the error rate.

adaptive · Stopping
Testing at 0.05 every time the data is looked at. The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.

When the looking happens

A p-value is defined relative to a sampling plan, so the same data means different things under different stopping rules. Testing five times at the nominal level rejects a true null 14% of the time, and no observation in the dataset changed.

sequential · Stopping
Two promises, and no rule here keeps both. A fixed-width procedure promises two things: that the interval covers at its nominal rate, and that it is no wider than the width asked for. Over 1500 runs of the modelled weighting, a rule that stops when the interval it will report is short enough keeps the width — only 2.0% of runs come out wider than 0.34 — and covers at 91.13% against a nominal 95%. A rule that stops on a width predicted from the within-arm sums of squares covers at 94.80% and comes out wider than promised on 42.3% of runs. The two promises are in conflict because keeping the second one exactly requires conditioning on the very quantity that has to be independent of the stopping time for the first.

Stopping on the arms

The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.

stop · Width
The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs a fair coin rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 155 of the 999 re-randomisations reach it, and the p-value is (1 + 155)/(1 + 999) = 0.1560. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 1.946.

The test that needs the rule

A randomisation test assumes almost nothing about the data and one thing about the experiment. Tell it a fair coin produced an allocation that an adaptive rule produced — which is what every off-the-shelf permutation routine does — and it rejects 8.0% of true nulls where knowing the rule gives 4.0%.

exact · Reference
Exact coverage, at every block size. Coverage of the interval each rule reports, at a nominal 95%, over 2,500 runs each with a standard error of 0.44 points. The blinded rule stops on the within-block contrasts and reports an interval built from the block means, and those two are independent whatever the rule does — so the interval is an ordinary t interval on b − 1 degrees of freedom and its coverage is exact. It is exact at every block size drawn. The interval a practitioner writes at the purely sequential rule's stopping time covers 91.72%, and Stein's two-stage rule is exact for the same reason as the blinded rule and spends 2.10 times the observations to be so. The bars are truncated at 86% so the differences can be seen.

The rule that cannot see the mean

A sequential rule stops when its own estimate of the spread is small, which is more often on the samples whose spread came out low — so the interval afterwards is short. There is a way to keep updating the estimate and stop being able to see the mean at all.

blind · Stopping
What a fixed-width interval covers, by the number of blocks the trial ran before it stopped. Two thousand runs of each rule, the modelled weighting, a promise of 0.34. Reading its report: 4–8 blocks, 22.3% of runs, 78.2%; 9–12 blocks, 16.6% of runs, 90.4%; 13–16 blocks, 18.4% of runs, 96.2%; 17–20 blocks, 17.4% of runs, 96.0%; 21–28 blocks, 17.9% of runs, 96.4%; 29–36 blocks, 7.4% of runs, 99.3% — 91.45% overall. Reading the arms: 4–8 blocks, 0.0%, none; 9–12 blocks, 0.9%, 94.4%; 13–16 blocks, 30.4%, 95.6%; 17–20 blocks, 50.0%, 93.9%; 21–28 blocks, 18.0%, 94.4%; 29–36 blocks, 0.7%, 92.3% — 94.50% overall.

The trials that stopped early

A fixed-width trial that stops when its own interval is short enough covers 91.45% — an average of 78.2% among the 22.3% of runs that stop within eight blocks and 96% to 99% among those that run longer. Widening every interval by 17.1% brings the average to 95% and leaves the early stops at 85.6%, while 92.8% of runs now report an interval wider than the width they promised. Even doubling every interval leaves the early stops short.

stop · Width
What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

pace · Stopping
What the interval covers, after a design that read the data. 800 experiments of 12 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 91.3% against 92.6%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 94.3% and 94.9%.

What a design chosen from the data costs

Two fields on this site measured what happens when a rule reads the data, and the error rate broke both times. A design that reads the data to decide where to put its runs breaks nothing — and the control that proves it also finds what the real shortfall is.

robust · Local design
The chance of crossing later from each interim z, under six schedules with O'Brien–Fleming-type spending boundaries. Exact. At an interim |z| of 1.0: end only 3.64%, +0.6 3.41%, +0.75 3.20%, +0.9 3.41%, every 0.125 3.00%, every 0.05 2.88%. At 2.5: end only 38.30%, +0.6 45.80%, +0.75 45.36%, +0.9 41.70%, every 0.125 50.08%, every 0.05 53.24%. The heavy line is the largest of the six at each z.

A look the trend asked for

Under an O'Brien–Fleming-type spending function, every schedule of looks fixed in advance spends exactly 5.0000%. A committee that adds a look at three quarters of the trial whenever the interim z is 1.5 or more spends 5.2323% — 5.315% counted over a hundred thousand trials — and the most a committee choosing among six schedules could spend is 5.4390%.

sequential · Stopping

Named alongside it

The objects these essays reach for when they reach for this one.

BlindingCoverageStopping ruleFixed-width intervalError rateMonte CarloSequential analysisAdaptive designBlockingDegrees of freedomIndependenceInterval width

All concepts