Stopping rules

Spending the error rate

The repair for interim testing is to spend 5% across the looks rather than at each one. The boundaries are solvable rather than quotable, and a trial that can stop early uses 298 observations where a fixed design uses 400 — at a cost of half a point of power.

Worth reading first: When the looking happens.

Looking five times at the nominal boundary rejects a true null 14% of the time. The repair is not to stop looking. It is to raise the boundary so that the five opportunities share the 5% between them rather than each taking it.

Four boundaries for 5 looks, all spending 5% in totaltest at 0.05 every look: 1.96, 1.96, 1.96, 1.96, 1.96. Pocock — a constant, higher boundary: 2.41, 2.41, 2.41, 2.41, 2.41. O'Brien–Fleming — strict early, nearly nominal at the end: 4.55, 3.22, 2.63, 2.27, 2.03. Bonferroni across looks: 2.58, 2.58, 2.58, 2.58, 2.58. Every one except the first spends the same total error rate; they differ in when they spend it.234512345lookcritical value of |z|test at 0.05 every lookPocockO'Brien–FlemingBonferroni across lookssolved for, not quotedthe same 5%, spent differently
Fig. 1 Four boundaries for five looks. Three of them spend exactly 5% in total and differ only in when they spend it.

The budget framing

The useful way to think about it is as an allocation problem.

A trial has 5% of error rate to spend. A fixed design spends all of it at one analysis, at a boundary of 1.96. A design with five looks must divide it — and the division is a genuine choice, because spending more early makes early stopping easier and leaves less for the end.

Two classical allocations, and they sit at opposite ends.

Pocock spends equally across the looks: one constant boundary, used at every analysis. For five looks that constant is 2.41.

O’Brien–Fleming spends almost nothing early and nearly everything at the end. Its boundary is c·√(K/k) — very high at the first look, falling to close to nominal at the last. For five looks the final boundary is about 2.03, barely above 1.96.

Both hold the overall error rate at 5%. Measured across twenty thousand null trials: Pocock 4.78%, O’Brien–Fleming 4.80%.

Solved, not quoted

A methodological point that matters more than it might, and which the site’s rule required.

These constants are usually looked up in a table. The tables are correct and they cover the standard cases — five equally spaced looks, 5% two-sided — and a site whose premise is that nothing is asserted without being counted cannot take a number from a table for a case it might change.

So the boundaries here are solved for. The z statistics across looks are correlated, because each is built from the same accumulating sum, so the crossing probability is not a product and has no closed form. The solver draws forty thousand null paths once, then bisects on the boundary multiplier until the fraction of paths crossing anywhere equals exactly 5%.

The check that the solver works is the case where the answer is known: with one look, the boundary must come out at 1.96, and it does. That single assertion is what licenses the other numbers, and without it a plausible-looking 2.41 would be unverifiable.

It also means the boundaries are correct for whatever number of looks the figure is dragged to, rather than correct for five and interpolated elsewhere.

The five boundaries, written out

The two allocations are easier to compare as five numbers each than as a description of a shape.

O’Brien–Fleming’s rule is cK/kc\sqrt{K/k} with the final boundary at 2.03, which fixes c and gives

4.54, 3.21, 2.62, 2.27, 2.03

against Pocock’s flat 2.41, 2.41, 2.41, 2.41, 2.41.

They cross between the third look and the fourth, which is the whole geometry of the choice: before the crossing O’Brien–Fleming is stricter, after it Pocock is.

What each spends at the first look

Converting the first boundary into a nominal p-value says how differently the budget is being allocated, and the difference is not a matter of degree.

Pocock’s first look stops the trial at a nominal two-sided p of 2(1Φ(2.41))2(1-\Phi(2.41)), which is 0.0160nearly a third of the whole five per cent budget, spent at the first analysis.

O’Brien–Fleming’s first look requires 2(1Φ(4.54))=5.6×1062(1-\Phi(4.54)) = 5.6\times10^{-6}. Six millionths.

A ratio of about 2,900 between the two at the same analysis of the same trial, from two rules that both spend exactly 5% in total.

What the looks cost at the end

The other side of the same budget is the final boundary, and it prices the whole design against a trial that never looks.

A fixed design uses 1.96. O’Brien–Fleming’s final boundary is 2.03, three and a half per cent higher; Pocock’s is 2.41, twenty-three per cent higher.

At an effect of three standard errors that is a final-look power of 85.1% for the fixed design, 83.4% for O’Brien–Fleming and 72.2% for Pocock — so the price of being allowed to look five times, paid entirely at the last analysis, is 1.7 points of power under one rule and 12.9 under the other.

Neither figure is the whole story, because both rules can stop early and a trial that stops early has spent fewer patients. But it does say what a reader should expect from the two: O’Brien–Fleming is very nearly a fixed design that happens to have four chances to stop early, and Pocock is a genuinely different design that has traded a large piece of its final power for a real chance of stopping at the first look.

The two null rates measured — 4.78% and 4.80% on twenty thousand trials, a standard error of 0.15 points apart — confirm only that both spend the budget. They say nothing about how, which is the entire content of the choice.

Which allocation to choose

The two spend the same total and behave very differently, and the choice follows from what stopping early is for.

O’Brien–Fleming is the usual default in clinical trials, for two reasons that reinforce each other. Stopping a trial early is a serious action — it ends recruitment, fixes the evidence base, and often determines practice — so it should require overwhelming evidence, and a first-look boundary above 4 supplies that. And because it spends so little early, the final analysis is barely penalised: a trial that runs to completion is tested at almost the boundary it would have used with no interim analyses at all.

That second property is what makes it nearly free. The design buys the option to stop early and gives up almost nothing if it does not exercise it.

Pocock stops earlier and more often, because its boundary is much lower at the first looks. That is appropriate where early stopping is cheap and desirable — an online experiment, a manufacturing check — and where the final analysis being penalised matters less than getting an answer quickly.

Bonferroni across the looks is also available and is the wrong answer for the same reason it is the wrong answer in the multiplicity field: it ignores the correlation between the looks and is conservative for it. Measured at 3.09% against a nominal 5%, it spends only three fifths of the budget it was given, and the unspent portion is power thrown away.

Five looks: what each rule delivers at an effect of 0.16. test at 0.05 every look: rejects a true null 14.4%, finds a real effect 92%, uses 204 observations on average. Pocock: rejects a true null 4.9%, finds a real effect 82%, uses 256 observations on average. O'Brien–Fleming: rejects a true null 4.9%, finds a real effect 88%, uses 299 observations on average. Bonferroni across looks: rejects a true null 3.1%, finds a real effect 77%, uses 274 observations on average.
Fig. 2 What each rule delivers: the error rate it holds, the power it achieves, and the observations it uses.

What early stopping buys

The reason to accept a raised boundary at all, and it is measurable.

A fixed design with four hundred observations, at a true effect of 0.16, has 89.4% power and uses four hundred observations every time. An O’Brien–Fleming design with five looks on the same maximum sample has 88.8% power and uses 297.6 observations on average.

A hundred observations saved, per trial, for six tenths of a percentage point of power.

The saving comes from the trials where the effect is real and large enough to cross the boundary early — those stop and stop counting. The trials where nothing is happening run to the end and use the full four hundred, which is the right allocation of effort.

That trade is unusually favourable and it explains why group-sequential designs are standard in trials despite the extra machinery. The cost is a fraction of a percentage point of power; the benefit is a quarter of the sample, which in a clinical trial is a quarter of the subjects exposed to an inferior treatment and a quarter of the cost.

The generalisation: alpha spending

The two classical boundaries are special cases of something more flexible, and the generalisation is what is actually used.

An alpha-spending function states how much of the error budget has been spent by any point in the trial, as a function of the fraction of information collected. Pocock and O’Brien–Fleming correspond to two particular spending functions, and any increasing function from 0 to α defines a valid design.

The practical advantage is that a spending function does not require the number and timing of the looks to be fixed in advance. Real trials do not recruit on schedule, and interim analyses happen when a data monitoring committee meets rather than at exact fractions of the sample. A spending function accommodates that: whatever fraction of the information has arrived, the function says how much alpha may be spent, and the boundary is computed from what remains.

What it cannot accommodate is choosing the timing based on the results. The flexibility is in the calendar, not in the data — deciding to hold an extra interim analysis because the trend looks promising reintroduces exactly the problem the design was built to prevent, and no spending function corrects for it.

Where this field meets the others

The stopping-rule problem is the same shape as two others on this site, and setting them side by side is worth doing because the remedies differ in an instructive way.

Multiple comparisons is branching across analyses. The number of branches is declared, the correction is a division, and it works.

Sequential testing is branching across time. The number of branches is declared, the correction is a solved boundary, and it works.

The garden of forking paths is branching across analyses that were never declared. The number of branches is unknown, no correction is available, and nothing works.

So the two fixable problems are exactly the two where somebody wrote down in advance how many opportunities there would be. That is not a coincidence about statistics; it is the whole content of what pre-specification buys. A declared multiplicity is an arithmetic problem and an undeclared one is not a problem that can be solved, and the machinery in this field only exists because trials have protocols.

O'Brien–Fleming — strict early, nearly nominal at the end — forty trials of a true null. Forty trials in which nothing is happening, monitored at 5 interim points. The boundary is 4.55, 3.22, 2.63, 2.27, 2.03. 1 of the forty cross it somewhere and would be reported as significant.
Fig. 3 Forty null trials against the O’Brien–Fleming boundary — very strict early, close to nominal at the end.

What to report

Short, and it closes the field.

The number of interim analyses, planned and performed. Without this a reader cannot know which reference distribution applies.

The spending function or boundary used, by name. “O’Brien–Fleming” is sufficient and it is one word.

The look at which the trial stopped, if it stopped early, and the boundary in force at that look. A trial that stopped at the second of five interim analyses crossed a boundary of about 3.0, and that is a much stronger result than the nominal p-value would suggest.

And the final sample size against the planned maximum. The gap is what the design bought.

All four are facts about the protocol rather than about the data, all four are known before the analysis, and between them they let a reader reconstruct what the reported p-value means — which, as the previous essay established, is not determined by the data alone.

Testing at 0.05 every time the data is looked at. The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.
Fig. 4 And the reason all of it is necessary: the rate the uncorrected procedure reaches, climbing with every look.
Four boundaries for 10 looks, all spending 5% in total. test at 0.05 every look: 1.96, 1.96, 1.96, 1.96, 1.96, 1.96, 1.96, 1.96, 1.96, 1.96. Pocock — a constant, higher boundary: 2.56, 2.56, 2.56, 2.56, 2.56, 2.56, 2.56, 2.56, 2.56, 2.56. O'Brien–Fleming — strict early, nearly nominal at the end: 6.60, 4.67, 3.81, 3.30, 2.95, 2.70, 2.50, 2.33, 2.20, 2.09. Bonferroni across looks: 2.81, 2.81, 2.81, 2.81, 2.81, 2.81, 2.81, 2.81, 2.81, 2.81. Every one except the first spends the same total error rate; they differ in when they spend it.
Fig. 5 The boundaries at ten looks. The same 5%, divided ten ways, with O’Brien–Fleming starting above five standard errors.

Why the boundary has to be higher, in one paragraph

The intuition, since the arithmetic above is a numerical solve and numerical solves are unpersuasive.

Under the null, the accumulating z statistic is a random walk. A walk that is checked once has one chance to be beyond 1.96 at that moment. A walk that is checked five times has five chances, and although the checks are correlated — the walk does not teleport between them — they are not the same check, because the walk moves in between.

Raising the boundary to 2.41 makes each individual crossing rarer, by enough that five chances at the higher boundary total the same 5% that one chance at 1.96 gives. That is the entire content of the Pocock constant, and the reason it cannot be written in closed form is that “five correlated chances” requires knowing the joint distribution of the walk at five times, which is a multivariate normal integral.

O’Brien–Fleming makes the same total from a different shape: because the walk is most variable early — relative to its own standard error, the early statistic is the noisiest — most of the crossing risk is concentrated at the first looks, and a boundary that is very high there removes most of the inflation while costing almost nothing at the end.

That is why the O’Brien–Fleming final boundary is 2.03 rather than 2.41. It has spent so little of the budget early that nearly all of it remains for the last analysis.

The measurement that makes the design worth it

The comparison at the heart of the field, restated because it is easy to lose among the boundaries.

design error rate power observations used
fixed, one look 4.9% 89.4% 400
naive, five looks 14.0%
O’Brien–Fleming, five looks 4.8% 88.8% 297.6

The middle row is the one to avoid and the one that happens by default. The bottom row is the same trial with a boundary computed in advance: the error rate is back where it belongs, the power is essentially unchanged, and a quarter of the observations are never needed.

The naive row has no power figure because a procedure that does not hold its error rate has no meaningful power — power is the rate of correct rejection for a test at a stated level, and this one is not at its stated level.

That omission is deliberate and it is the point. A procedure that fails its own size cannot be credited with detecting anything, because its detections include the 14% it produces from nothing.

What the field concludes

Two essays, and they divide the same way the multiplicity field does: what goes wrong, and what fixes it.

Testing five times at the nominal level rejects a true null 14% of the time, ten times 19%, continuously 100%, and one look exactly 5% — the last being the control that makes the others credible.

A p-value is a property of a dataset and a plan. Two investigators with identical data and different stopping rules have different p-values, and both are right.

Raising the boundary fixes it exactly, at 2.41 for five equally-spaced looks, and the boundaries are solvable rather than requiring a table.

And the design pays for itself: 297.6 observations against 400, for six tenths of a percentage point of power.

The through-line with the rest of the site is the one the thread is named for. The rule by which data was going to be collected is part of the result, not context around it — and the number reported at the end cannot be interpreted without it.

Five looks: what each rule delivers at an effect of 0.25. test at 0.05 every look: rejects a true null 14.4%, finds a real effect 100%, uses 122 observations on average. Pocock: rejects a true null 4.9%, finds a real effect 100%, uses 149 observations on average. O'Brien–Fleming: rejects a true null 4.9%, finds a real effect 100%, uses 211 observations on average. Bonferroni across looks: rejects a true null 3.1%, finds a real effect 100%, uses 160 observations on average.
Fig. 6 The same four rules against a larger real effect, where every design stops earlier and the savings are bigger.

What the boundaries look like at other numbers of looks

The constants above are for five looks, and the shape of their dependence on the number of looks is worth knowing because it determines how much flexibility a design can afford.

Pocock’s constant rises slowly with the number of looks — it has to, since each additional look is another chance to cross, but the looks are increasingly correlated so each one adds less. Going from two looks to ten raises the constant much less than doubling it.

O’Brien–Fleming’s final boundary rises hardly at all. Because the shape puts almost all the strictness early, adding looks mostly adds very high early boundaries that are rarely crossed, and the last analysis stays close to 1.96 however many interim looks precede it.

That is the practical argument for the O’Brien–Fleming shape over Pocock’s, beyond the ethical one about early stopping. The number of interim analyses can be increased at almost no cost to the final test. A design with ten looks and an O’Brien–Fleming boundary is tested at the end at a threshold barely distinguishable from the fixed-sample one, so the monitoring is nearly free.

Under Pocock, by contrast, every extra look raises the constant boundary everywhere, including at the end, so more monitoring makes the final analysis strictly harder to pass.

The failure mode that survives all of this

One thing no boundary corrects, and it is the one to watch for.

The design controls the error rate for the looks that were planned. An investigator who holds an unplanned extra analysis — because a conference deadline arrives, or because the trend looks promising, or because a monitoring committee asks — has spent alpha that the spending function did not allocate.

Spending functions handle unplanned timing gracefully, which is their advantage over fixed boundaries. They do not handle unplanned looks, and the distinction is exactly the one that matters: adjusting for an analysis that happened at 43% of information rather than 40% is routine, and adding a sixth analysis to a five-look design is not.

The tell is the same as everywhere else in this area. A decision about the analysis that was influenced by seeing the data cannot be corrected for afterwards, because the correction needs to know how many such decisions were available and that number is not recoverable.

So the field’s machinery, sophisticated as it is, rests on the same foundation as the multiple-comparison procedures: it converts a declared multiplicity into a controlled error rate, and it has nothing to offer an undeclared one.

That is not a limitation of the boundaries. It is the boundary of what any frequentist correction can reach, and recognising where it falls is the difference between a design that controls its error rate and one that performs the appearance of doing so.

The first is a property of the protocol; the second is a property of the methods section, and only the protocol constrains what the trial was allowed to do.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Alpha spendingError rateGroup-sequential designO'Brien–Flemingp-valueThe Pocock boundary