Stopping rules

The outcomes a trial could have stopped with

A trial that stops at its second look with z = 3.3 has a two-sided p-value of 0.000969, 0.000987, 0.00187 or 0.0421, depending on how the outcomes it could have stopped with are ordered. One of the four orderings does not change when the looks the trial never reached are replanned, and the same one gives a trial that ran to the end with z = 6 a p-value of 0.0256.

Worth reading first: When the looking happens.

An O’Brien–Fleming trial of four hundred observations, looked at after every eighty, stops at its second look: a hundred and sixty observations, z = 3.3, past the boundary of 3.226. The trial reports a p-value. If it had been a fixed trial of a hundred and sixty observations, the p-value would be the chance of a z at least as large in either direction with no effect, 2(1 − Φ(3.3)) = 0.000967.

It was not a fixed trial. A p-value is relative to the plan that produced the data, and the plan here could have produced outcomes a fixed trial cannot: a stop at eighty observations, a stop at two hundred and forty with a smaller z, a trial that ran to four hundred and ended on 2.9. “At least as extreme as what was observed” has to rank all of those against the stop that happened, and the data do not say how.

The estimate such a trial reports was the first place this mattered — the trials that stop at a hundred and sixty observations report 1.83 times the effect on average — and the repair it named, an estimate that knows the rule, needs the same ranking. So does a confidence interval. This essay computes four standard rankings, exactly, and what each does to one trial.

Ordered stagewise: the outcomes at least as extreme as stopping at 160 observations with z = 3.3. Each column is one look of an O'Brien–Fleming trial; above the boundary a trial stops there. Highlighted are the outcomes that count as at least as extreme as the observed one when outcomes are ordered stagewise: at 80, z ≥ 4.56 (probability 2.54 × 10⁻⁶ with no effect); at 160, z ≥ 3.30 (probability 4.82 × 10⁻⁴ with no effect); at 240, none; at 320, none; at 400, none. The two-sided p-value is 9.69 × 10⁻⁴.
Fig. 1 The outcomes of an O’Brien–Fleming trial, one column per look; above a boundary, a trial stops there. Highlighted are the outcomes that count as at least as extreme as stopping at a hundred and sixty observations with z = 3.3 when earlier stops rank first.

Four answers to “further out”

An outcome of a sequential trial is a pair: the look at which it ended and the statistic there. Four ways of ordering pairs are in standard use, and each is a defensible reading of “more evidence against no effect”.

Stagewise. A trial that stopped earlier is more extreme than any trial that stopped later, and within a look a larger z is more extreme. The reasoning is that crossing a boundary sooner is stronger evidence, because the boundary was set higher there to make it so. In the figure the highlighted outcomes are every crossing at the first look and the crossings at the second look with z of at least 3.3; nothing at a later look counts, however large its z.

By the estimate. An outcome is more extreme when its estimate of the effect, the sum divided by the observations, is larger, whatever look it came from. A stop at a hundred and sixty with z = 3.3 is an estimate of 5.218 on the information scale, so an outcome at two hundred and forty is as extreme when its z is at least 4.04, at three hundred and twenty 4.67, and at four hundred 5.22.

By z. An outcome is more extreme when its z is larger, whatever the look. It is the fixed-trial ranking applied without regard to when.

By the sum. An outcome is more extreme when its partial sum is larger. The sum grows with the observations, so a later look reaches a given sum at a smaller z: 2.69 at two hundred and forty, 2.33 at three hundred and twenty, 2.09 at four hundred.

Nothing about the trial chooses between them. Each is a total order on the same set of outcomes, each produces a valid p-value — under no effect, the chance of a p-value at or below any level is that level — and each is a choice made before the p-value can exist.

One trial, four p-values

The p-value under an ordering is the probability with no effect of the highlighted region, doubled for two sides. Each region’s probability is an integral over the same recursion that computes the boundaries: the density of the partial sum on the region where the trial is still running, carried from look to look.

Ordered stagewise, the region is the whole first-look boundary, which trials cross with probability 0.00000254 with no effect, and the part of the second look above 3.3, with probability 0.000482. Doubled, 0.000969. Ordered by the estimate the regions at later looks add almost nothing, because a z of 4.67 at three hundred and twenty observations is rare, and the p-value is 0.000987. Ordered by z it is 0.00187, nearly twice as large, because a z of 3.3 at a later look is not rare at all.

Ordered by the sum: the outcomes at least as extreme as stopping at 160 observations with z = 3.3. Each column is one look of an O'Brien–Fleming trial; above the boundary a trial stops there. Highlighted are the outcomes that count as at least as extreme as the observed one when outcomes are ordered by the sum: at 80, z ≥ 4.67 (probability 1.53 × 10⁻⁶ with no effect); at 160, z ≥ 3.30 (probability 4.82 × 10⁻⁴ with no effect); at 240, z ≥ 2.69 (probability 0.0031 with no effect); at 320, z ≥ 2.33 (probability 0.0070 with no effect); at 400, z ≥ 2.09 (probability 0.0104 with no effect). The two-sided p-value is 0.0421.
Fig. 2 The same outcome, with outcomes ordered by their partial sum. A later look reaches the observed sum at a smaller z, so most crossings at the last three looks count as at least as extreme.

Ordered by the sum, it is 0.0421. The highlighted region now takes in most of the crossings at the last three looks — probabilities of 0.00315, 0.00703 and 0.0104 with no effect — because a trial that ran on to four hundred observations and crossed at a z of 2.1 has accumulated more total evidence, in this ordering’s sense, than one that stopped at a hundred and sixty on 3.3.

The same trial, the same data, a p-value from 0.000969 to 0.0421 — a factor of forty-three, all of it an ordering. The first three agree to within a factor of two and the fourth is an outlier, but it is not a mistake: it is the ordering that answers the question “how much evidence has accumulated”, and the other three answer “how early and how hard was the boundary crossed”.

Where the orderings part company

One O'Brien–Fleming trial's two-sided p-value under four orderings of its outcomes. Stopped at 80 with z = 4.6: stagewise 4.22 × 10⁻⁶, by the estimate 4.22 × 10⁻⁶, by z 8.75 × 10⁻⁶, by the sum 0.0470; a fixed design 4.22 × 10⁻⁶. Stopped at 160 with z = 3.3: stagewise 9.69 × 10⁻⁴, by the estimate 9.87 × 10⁻⁴, by z 0.0019, by the sum 0.0421; a fixed design 9.67 × 10⁻⁴. Stopped at 240 with z = 2.8: stagewise 0.0057, by the estimate 0.0060, by z 0.0089, by the sum 0.0308; a fixed design 0.0051. Stopped at 320 with z = 2.5: stagewise 0.0168, by the estimate 0.0176, by z 0.0207, by the sum 0.0236; a fixed design 0.0124. Stopped at 400 with z = 3.0: stagewise 0.0258, by the estimate 0.0130, by z 0.0046, by the sum 4.25 × 10⁻⁴; a fixed design 0.0027. Stopped at 400 with z = 6.0: stagewise 0.0256, by the estimate 1.52 × 10⁻⁴, by z 3.19 × 10⁻⁹, by the sum below 10⁻¹⁵; a fixed design 1.97 × 10⁻⁹.
Fig. 3 Two-sided p-values of six outcomes under the four orderings, on a log scale, with a fixed trial’s p-value for the same z as a grey tick and 0.05 dashed.

Six outcomes, set side by side, show where the disagreement lives.

Early stops on large z are where stagewise and the estimate agree. A trial that stops at the first look with z = 4.6 has a p-value of 0.0000042 under both — and 0.0470 ordered by the sum, because eighty observations carry a small sum whatever their z, and nearly every crossing at a later look carries a larger one.

Stops near a later boundary are where all four come together. Stopping at three hundred and twenty observations with z = 2.5 gives 0.0168, 0.0176, 0.0207 and 0.0236: a crossing that barely cleared a boundary near the end is not far out under any ordering.

Trials that ran to the last look are where they separate the other way. A trial ending at four hundred with z = 3.0 has a p-value of 0.0258 stagewise, 0.0130 by the estimate, 0.00464 by z and 0.000425 by the sum. The ranking by the sum, which put the early stop last, puts this outcome first.

So no ordering is uniformly the lenient one. Each concentrates its judgement on a different kind of trial, and a reader who knows only the p-value does not know which kind of trial produced it.

All four agree at 5%, and nowhere else

Four rankings of the same outcomes would be alarming if they disagreed about which trials are significant at the level the design was built to hold. They do not, and the reason is worth having.

Twenty thousand trials with no effect were run and each outcome’s p-value computed under all four orderings. 5.16% of the trials crossed a boundary. Under every ordering, 5.16% had a p-value at or below 0.05 — and they were the same trials, in 100.000% of cases: no trial was significant under one ordering and not under another, and none was significant without having crossed. At 0.01 the four orderings reject 0.96%, 0.97%, 0.97% and 0.98% of the trials, and at 0.001 between 0.09% and 0.10%. Each is within its standard error of the level, which is what it means for each to be a p-value.

The agreement at 5% is a property of the O’Brien–Fleming shape rather than a coincidence of these trials. Its boundary is c/tc/\sqrt{t} on the zz scale, which makes it c/tc/t as an estimate and a constant cc as a sum — so written in any of the three currencies the orderings use, no interim boundary is lower than the last one. Every trial that crossed therefore ranks ahead of every trial that did not, under each ordering, and the crossings are exactly the outcomes whose tail is 2.5% or less.

So the choice of ordering never changes whether a trial of this design is reported as significant at 5%. It changes everything reported beside that decision: the p-value, which ranges over a factor of forty-three for one trial; the interval; and the adjusted estimate. Those are the numbers that travel on into a review, a guideline or the next trial’s power calculation, and the decision does not travel without them.

A trial that ran to the end

The p-value of a trial that ran to its last look, against the z it ended with. Exact. Ordered stagewise, every trial that stopped at an interim look counts as more extreme than any trial that reached the last one, so the p-value cannot fall below 0.0256 however large the final z: at z = 6 it is 0.0256, where ordering by z gives 3.19 × 10⁻⁹ and a fixed design 1.97 × 10⁻⁹.
Fig. 4 The two-sided p-value of a trial that reached its last look, against the z it ended with, under three orderings and for a fixed trial of four hundred observations.

The stagewise ordering has one consequence that is worth seeing on its own, because it is the reason the ordering is sometimes rejected.

Every trial that stopped at an interim look ranks ahead of every trial that reached the last one. With no effect, the four interim looks together are crossed upward with probability 0.01279, so a trial that reached four hundred observations cannot have a stagewise p-value below twice that, 0.0256, whatever it found there. At z = 3.0 the p-value is 0.0258. At z = 6.0 it is still 0.0256 — where ordering by z gives 3.19 × 10⁻⁹ and a fixed trial of the same size 1.97 × 10⁻⁹.

A final z of 6 is overwhelming evidence by any reading a reader would recognise, and the stagewise p-value ranks it below a trial that stopped at three hundred and twenty observations on 2.29. The ordering is doing what it says: it treats the timing of a crossing as the primary evidence and the size of the statistic as secondary. For a trial that stops early that is reasonable, because the boundary was built to make early crossings hard. For a trial that ran to the end it throws away the thing that happened.

Ordering by the estimate falls between, at 0.000152 for z = 6: a final estimate that large outranks every early stop except the ones whose estimates were larger still. Every figure in this section is two-sided, and every one of them is a valid p-value.

The looks the trial never reached

The p-value of stopping at 160 observations with z = 3.3, under two plans for the looks it never reached. The planned trial looks three more times; the other plan has the same first two looks and then one final analysis at 1.964. stagewise: 9.69 × 10⁻⁴ planned, 9.69 × 10⁻⁴ under the other plan; by the estimate: 9.87 × 10⁻⁴ planned, 9.69 × 10⁻⁴ under the other plan; by z: 0.0019 planned, 0.0018 under the other plan; by the sum: 0.0421 planned, 0.0371 under the other plan.
Fig. 5 The p-value of the same stop, at a hundred and sixty observations with z = 3.3, under the planned five looks and under a plan with the same first two looks and a single final analysis.

The trial stopped at its second look. Suppose its protocol had instead said: two interim looks on the same boundaries, then one final analysis at four hundred observations, with its boundary solved to spend what is left of the 5% — which comes to 1.964. The first two looks are identical under both plans, the trial stopped at the second, and the data are the same numbers.

Stagewise, the p-value is 0.000969 under both plans, to every digit. Ordered by the estimate it moves from 0.000987 to 0.000969; by z, from 0.00187 to 0.00183; by the sum, from 0.0421 to 0.0371.

The reason is visible in the first figure. Stagewise, nothing at a later look is ever more extreme than an earlier stop, so the later looks’ boundaries never enter the calculation. Every other ordering counts some outcomes at later looks as at least as extreme, and the probability of those outcomes depends on the boundaries that would have been in force there — boundaries of looks that did not happen, for a trial that no longer exists.

That is not a curiosity. A reference distribution is the set of experiments that could have been run, and a sequential trial’s could-have-run set extends past its own stop. A monitoring committee that stops a trial is often not the body that wrote down what the later looks would have been; a design with a spending function does not fix when its later looks happen at all. In either case three of the four p-values cannot be computed, and one can. The stagewise p-value is the only one of the four that depends on nothing after the stop.

Five intervals for the same trial

Five 95% intervals for one trial that stopped at 160 observations with z = 3.3. The mean at the stop is 0.2609. x̄ ± 1.96/√n: 0.1059 to 0.4158, estimate 0.2609; stagewise: 0.1059 to 0.4158, estimate 0.2609; by the estimate: 0.1029 to 0.3857, estimate 0.2459; by z: 0.0751 to 0.3551, estimate 0.2043; by the sum: 0.0038 to 0.6297, estimate 0.1061. The four ordered intervals each cover exactly 95% of trials; they are not the same interval.
Fig. 6 Five 95% intervals for the trial that stopped at a hundred and sixty observations with z = 3.3, with each ordering’s median-unbiased estimate, and the mean at the stop dashed.

An ordering gives an interval by the same arithmetic, run backwards over the effect: the lower limit is the effect at which the observed outcome sits in the top 2.5% of outcomes, the upper limit the effect at which it sits in the bottom 2.5%, and the median-unbiased estimate the effect at which it sits in the middle. Each limit is found by bisection on the recursion.

For the stop at a hundred and sixty observations with z = 3.3, the mean at the stop is 0.2609 and the naive interval, the mean plus and minus 1.96 over the square root of the observations, runs from 0.1059 to 0.4158. The stagewise interval runs from 0.1059 to 0.4158, with a median-unbiased estimate of 0.2609. Ordered by the estimate, 0.1029 to 0.3857 around 0.2459; by z, 0.0751 to 0.3551 around 0.2043; by the sum, 0.0038 to 0.6297 around 0.1061.

The first of those is the unexpected one. At an early O’Brien–Fleming stop the stagewise interval is the naive interval, to four places, and its adjusted estimate is the unadjusted mean. The only outcomes ranked ahead of this one are crossings at the first look, and at every effect in range those are so improbable that ordering stagewise at the second look is ordering by z at the second look alone. The estimate that knows the rule, under the ordering that needs nothing after the stop, has nothing to adjust — and the trials that stop at this look report 1.83 times the true effect on average.

Exactly 95%, and not at the look

How often five 95% intervals contain the true effect of 0.16, by where the trial stopped. stop at 80 · 0.09%: x̄ ± 1.96/√n 0.00%, stagewise 0.00%, by the estimate 0.00%, by z 0.00%, by the sum 100.00%. stop at 160 · 11.4%: x̄ ± 1.96/√n 78.61%, stagewise 78.82%, by the estimate 79.14%, by z 87.25%, by the sum 99.37%. stop at 240 · 32.7%: x̄ ± 1.96/√n 98.99%, stagewise 100.00%, by the estimate 99.89%, by z 97.26%, by the sum 97.48%. stop at 320 · 28.5%: x̄ ± 1.96/√n 100.00%, stagewise 100.00%, by the estimate 100.00%, by z 99.78%, by the sum 96.53%. end at 400 · 27.3%: x̄ ± 1.96/√n 90.86%, stagewise 90.83%, by the estimate 90.83%, by z 90.83%, by the sum 88.58%. every trial: x̄ ± 1.96/√n 94.65%, stagewise 95.00%, by the estimate 95.00%, by z 95.00%, by the sum 95.00%. Counted over 20,000 trials, the naive interval covers 94.58%.
Fig. 7 How often each interval contains a true effect of 0.16, among the trials that ended at each look and among all trials, with the naive interval as an open circle.

That is not a contradiction, and resolving it is the most useful thing the orderings teach.

Each ordered interval covers a true effect of exactly 95% of trials. It is 95% by construction — the observed outcome’s rank at the true effect is uniform — and computing the coverage directly, from the stretch of outcomes whose interval contains 0.16, returns 95.0000% for all four. The naive interval covers 94.65%, computed exactly, and 94.58% over twenty thousand counted trials. Across all trials the naive interval is short by a third of a point.

Among the trials that stopped at a particular look, nothing covers 95%. Of the trials that stop at eighty observations the stagewise, estimate and z intervals cover 0% — every one of those trials reports at least 0.510, and no interval around it reaches 0.16 — and among those that stop at a hundred and sixty, 78.82%, 79.14% and 87.25%. At two hundred and forty and three hundred and twenty observations the same intervals cover between 97.26% and 100%, and at four hundred, 90.83%. The overall 95% is an average of groups that are each wrong, which is the same shape the estimate at the stop had.

The ordering by the sum is the exception that confirms it. Its intervals cover 100% at eighty observations and 99.37% at a hundred and sixty — by being the interval that ran from 0.0038 to 0.6297 — and pay for it with 88.58% at four hundred. Every ordering spends its 95% somewhere; none can spend it evenly across looks, because the look a trial stopped at is itself evidence about the effect.

So the orderings repair the statement about the procedure: the p-value is a p-value and the interval covers 95% of the trials that could have run. They cannot repair the statement about the trials that stopped at this look, because that statement is about a group selected by the effect, and an interval after a stop that read the data is conditioned on the reading whatever it is computed from.

What a report after a sequential stop should say

The ordering, by name. A p-value, an interval or an adjusted estimate after a sequential stop is not defined until an ordering is, and the range here is a factor of forty-three on one trial. “Adjusted for the interim analyses” does not say which adjustment.

Stagewise, unless there is a reason. It is the only one of the four computable without the plan for looks that did not happen, which is the situation of every trial whose later looks were never fixed, and at the first two looks it agrees with the estimate ordering to within two per cent. Its weakness is the trial that ran to the end, where it puts a floor under the p-value equal to what the interim looks spent; a trial that reached its last look with a large statistic should report that floor for what it is.

The look, beside the interval. A 95% interval from a stop at a hundred and sixty observations covers 95% of trials and 79% of trials that stopped there. That is not a defect any ordering fixes, and the reading of 95% as a statement about a procedure is the only reading under which the interval is right.

What is exact here, and what is counted

Exact. Every p-value, every per-look probability, every interval and median-unbiased estimate, the final boundary of 1.964 for the two-look plan, and every coverage, overall and by look. Exact means quadrature on four hundred and one points a look and bisection to well below the last digit shown; the four overall coverages of 95.0000% are the check that the interval construction and the coverage calculation, which share the recursion and nothing else, agree.

Counted. The naive interval’s coverage over twenty thousand trials, 94.58% against 94.65%.

Particular to this trial. One O’Brien–Fleming design with five equal looks and a known variance, and one observed stop. The disagreements between orderings are general; their sizes are this trial’s.

Still open: a boundary for giving up

Every outcome ordered here was a stop for benefit or a trial that ran to the end, because the trial had only boundaries for benefit — and their mirror images, which no trial with a real effect ever reaches. Most trials monitored this way also carry a boundary for stopping because the effect looks absent, and that boundary changes more than the ordering. It removes trials that would have crossed later, which costs power; it removes trials that would have crossed by chance, which gives back error rate; and whether the error rate is given back depends on whether anyone is bound to obey it. A boundary for giving up counts all three, and the difference between a futility rule that binds and one that does not turns out to be a difference in where the benefit boundary may sit.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formConditional distributionConfidence intervalCoverageGroup-sequential designInterim analysisMedian-unbiasedO'Brien–Flemingp-valueReference distributionStagewise orderingStopping rule