Stopping rules

A boundary for giving up

Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.

Worth reading first: When the looking happens.

Every trial measured so far in this series could stop for one reason: the evidence for benefit crossed a boundary. A trial with no effect almost never does, so it runs to the end. The O’Brien–Fleming trial of four hundred observations, looked at after every eighty, uses 398.6 observations on average when there is nothing to find.

That is the case most monitoring committees are least willing to pay for. A trial of a treatment that does not work exposes patients to it, spends money, and delays the next trial, and the data usually say so long before the last look. So most monitored trials carry a second boundary, below the first, at which the trial stops because the effect looks absent. This essay adds three such boundaries to one trial and counts what each saves, what each costs, and what each does to the error rate the benefit boundary was solved to hold.

The trial here tests for benefit only. Its benefit boundary spends a one-sided 2.5% — the upper half of the two-sided 5% the boundaries in this series were solved for — and is the same boundary: 4.562, 3.226, 2.634, 2.281 and 2.040.

Forty trials at a true effect of 0.16, under the rule "power at the trend < 10%". The upper line is the benefit boundary (4.56, 3.23, 2.63, 2.28, 2.04); the lower line is where the rule stops a trial for futility (0.40 at 80, 0.66 at 160, 0.95 at 240, 1.31 at 320). Of forty trials with a real effect, 29 cross for benefit and 11 are stopped for futility.
Fig. 1 Forty trials with a real effect of 0.16 per observation, under a rule that stops a trial when conditional power at the observed trend falls under 10%. Trials stopped for benefit are marked above, trials stopped for futility below.

Three reasons to give up

Stop if z is below zero. The simplest rule in use: at any interim look, a trial whose estimate points the wrong way stops. It asks nothing about how far the wrong way, or how early.

Stop if conditional power at the design effect is under 10%. Conditional power is the chance that the trial will cross at its last look, given where the partial sum is now and an assumed effect for the rest of the trial. Assuming the effect the trial was designed to detect — 0.16 per observation here — this rule stops only a trial whose data are so bad that even the hoped-for effect, for every remaining observation, would leave it under a 10% chance of success.

Stop if conditional power at the observed trend is under 10%. The same calculation with the effect for the rest of the trial taken to be the one estimated so far. It is the version a committee reaches for when it wants to ask “if things carry on as they are”, and it is the one that believes the interim estimate.

Three futility boundaries under one benefit boundary. The benefit boundary is 4.562, 3.226, 2.634, 2.281, 2.040. A trial stops for futility at the first four looks when z is below: stop if z < 0: 0.000, 0.000, 0.000, 0.000; power at the design effect < 10%: -3.726, -1.380, -0.065, 0.925; power at the trend < 10%: 0.400, 0.662, 0.952, 1.312.
Fig. 2 The z below which each rule stops the trial at the first four looks, under the benefit boundary they share. The design-effect rule’s boundary starts far below zero and rises; the trend rule’s is above zero from the first look.

On the z scale the three rules are three boundaries. “Below zero” is zero at every look. The design-effect rule stops at −3.726 at eighty observations, −1.380 at a hundred and sixty, −0.065 at two hundred and forty and 0.925 at three hundred and twenty. The trend rule stops at 0.400, 0.662, 0.952 and 1.312 — above zero at every look, including the first, where no other rule would stop anything with a positive estimate.

The difference is where each takes its effect from. At the first look the estimate of the effect has a standard error of 0.112, seven tenths of the effect the trial is looking for. The trend rule reads that estimate as the truth for the other three hundred and twenty observations, so a z of 0.400 — an estimated effect of 0.045 — is enough to call the trial hopeless.

What giving up saves

What each futility rule costs at a true effect of 0.16 and saves when there is none. no futility rule: power 88.45%, 398.6 observations on average with no effect and 297.2 at 0.16; 0.00% of trials with the real effect stopped for futility. stop if z < 0: power 83.25%, 195.5 observations on average with no effect and 273.3 at 0.16; 8.58% of trials with the real effect stopped for futility. power at the design effect < 10%: power 88.33%, 287.9 observations on average with no effect and 294.6 at 0.16; 2.73% of trials with the real effect stopped for futility. power at the trend < 10%: power 75.22%, 134.0 observations on average with no effect and 244.5 at 0.16; 21.28% of trials with the real effect stopped for futility.
Fig. 3 Power at a real effect of 0.16, and the mean number of observations with no effect and at 0.16, for the trial with no futility rule and under each of the three.

With no effect, “below zero” stops 72.66% of trials before the end and brings the mean number of observations from 398.6 to 195.5. The design-effect rule stops 82.45% — most of them at the last two interim looks, where its boundary has climbed — and uses 287.9. The trend rule stops 94.52% and uses 134.0, a third of what the trial without a rule spends on nothing.

The error rate falls as well, and it has to. A futility stop removes trials that would otherwise have had later chances to cross for benefit, and some of those would have crossed by chance. Obeyed, the three rules hold the one-sided rate at 2.205%, 2.482% and 1.761% where the benefit boundary was solved for 2.5%. That is the sense in which a futility rule is “conservative”: the trial rejects a true null less often than it says.

What giving up costs

The same removal happens when the effect is real, and then the trials removed are ones that would have succeeded.

At an effect of 0.16, power without a futility rule is 88.45%. “Below zero” takes it to 83.25%, a loss of 5.20 points, and stops 8.58% of trials with the real effect — 7.62% at the very first look, where eighty observations point the wrong way by chance. The design-effect rule takes power to 88.33%, a loss of 0.12 points, stopping 2.73% of real effects. The trend rule takes it to 75.22%, a loss of 13.23 points: it stops 21.28% of trials with a real effect of the size the trial was built to find, 15.12% of them at the first look.

The power each futility rule gives up, across true effects. Exact, non-binding. stop if z < 0: at most 5.64%, at an effect of 0.14; power at the design effect < 10%: at most 0.19%, at an effect of 0.12; power at the trend < 10%: at most 14.24%, at an effect of 0.14.
Fig. 4 The power each rule gives up against the same trial with no futility rule, across true effects per observation.

Across effects the loss has a peak, and the peak is not at the design effect. “Below zero” loses at most 5.64 points at an effect of 0.14; the trend rule at most 14.24 points at an effect of 0.14; the design-effect rule never more than 0.19 points, at 0.12. At no effect there is no power to lose, and at a large effect almost no trial looks hopeless at any look; the loss lives in between, where interim estimates are noisy enough to fall below a boundary and the effect is real enough for that to matter. That is the region a power calculation is usually aimed at, so a futility rule’s price is largest exactly where trials are designed to operate.

How many observations a trial uses under each futility rule, across true effects. Exact, non-binding. With no effect: no futility rule 398.6, stop if z < 0 195.5, power at the design effect < 10% 287.9, power at the trend < 10% 134.0. At 0.16: 297.2, 273.3, 294.6, 244.5.
Fig. 5 The mean number of observations before the trial ends, across true effects, with no futility rule and under each of the three.

The saving runs the other way. At no effect the rules differ by more than two hundred and sixty observations; at the design effect they use 297.2, 273.3, 294.6 and 244.5 — the trend rule’s saving there is fifty-three observations, bought with thirteen points of power. At an effect of 0.30 all four curves meet, because every trial crosses for benefit early and no futility boundary is ever reached.

Why the design-effect rule is nearly free

The design-effect rule saves 110.7 observations on a trial with no effect and costs 0.12 points of power at the effect the trial was built for. The trend rule saves 264.6 and costs 13.23. The ratio between them is the whole argument for choosing the effect conditional power is computed at.

A rule computed at the design effect stops only when the data would be hopeless even under the hypothesis the trial exists to test. Early on, almost nothing is that hopeless — its boundary at eighty observations is a z of −3.726, which a trial with no effect reaches about once in ten thousand — so it waits, and does its stopping at the last two interim looks, where the remaining observations are too few to rescue a poor start. A rule computed at the trend stops whenever the data so far look poor, and early data look poor often, whatever the truth. It converts the noise of eighty observations into a decision about four hundred.

Put as an exchange rate — observations saved when there is no effect, for each point of power given up at the design effect — “below zero” saves 203.1 observations for 5.20 points, 39 a point. The trend rule saves 264.6 for 13.23, 20 a point. The design-effect rule saves 110.7 for 0.12, about 890 a point. The trend rule saves the most observations in total and pays the most for each one.

The conditional-power arithmetic itself is the same in both; one number differs. That number is a judgement — what effect to assume for the part of the trial not yet run — and the counts say which judgement costs what. A committee that believes the interim trend is not being more realistic than one that holds to the design effect. It is using an estimate with a standard error the size of the thing being estimated as though it had none.

What a trial stopped for futility can still contain

A futility stop is reported as a result — the treatment did not do what was hoped — and the stopped trial’s own interval says how much that result covers.

Take the trials with a real effect of 0.16 that each rule stops. Under the trend rule their mean estimate of the effect is 0.0016, essentially nothing, which is what they were stopped for. But the ordinary 95% interval around each one’s estimate, from the observations it had when it stopped, still reaches up to 0.16 in 83.10% of them. Four in five of the trials the trend rule abandons cannot, on their own data, rule out the effect they were built to find. At the first look, where most of its stops happen, the share is 83.5%.

Under “below zero” the stopped trials report a mean of −0.0468, and 59.66% of their intervals still reach 0.16: two thirds of the stops at the first look, where an interval from eighty observations is ±0.219 wide on either side, and none of the few at later looks. Under the design-effect rule the stopped trials report 0.0212, and only 4.68% of their intervals reach 0.16. When that rule gives up, the data it gives up on exclude the effect the trial was built for in more than nineteen stops in twenty.

With no effect at all the picture holds. Of the trend rule’s futility stops, 54.54% still have intervals reaching 0.16; of the design-effect rule’s, 0.68%.

This is the difference between the rules seen from the report rather than from the design. A futility stop under the trend rule says that the estimate so far is small, and nothing more. A futility stop under the design-effect rule says that the estimate so far is small enough to exclude what the trial was looking for. Only the second is evidence that the treatment lacks the effect the trial was about, and the stopped trial’s estimate is as selected as a stop for benefit’s, truncated from above instead of from below. The two kinds of futility stop are written up in the same words.

Binding and non-binding

The benefit boundary with the rule "power at the trend < 10%" non-binding and binding. Non-binding, the benefit boundary is 4.562, 3.226, 2.634, 2.281, 2.040 and power at 0.16 is 75.22%. Binding, it falls to 4.250, 3.005, 2.454, 2.125, 1.901 and power is 78.42%. A binding design whose futility rule is then ignored rejects a true null 3.523% of the time.
Fig. 6 The benefit boundary with the trend rule non-binding — solved as if there were no futility rule — and binding, solved with the futility rule in force, beside the binding futility boundary.

The figures so far kept the benefit boundary where it was solved without any futility rule. That is a non-binding design: the futility rule is advice, the trial may continue past it, and the benefit boundary holds 2.5% whether it does or not. It pays for that freedom with the conservatism counted above — 1.761% under the trend rule, where 2.5% was allowed.

A binding design spends the full 2.5% by solving the benefit boundary with the futility rule in force. Removing the trials that would drift up later leaves room to lower the boundary, and it falls — under the trend rule from 4.562, …, 2.040 to 4.250, 3.005, 2.454, 2.125, 1.901. Power at 0.16 rises from 75.22% to 78.42%, recovering 3.19 of the 13.23 points the rule cost. Under “below zero” the last benefit boundary falls to 1.988 and power rises to 84.09%; under the design-effect rule, to 2.037 and 88.39%.

The price of the recovered power is a promise. A binding design is valid only if every futility stop is obeyed. A binding trend-rule design whose futility rule is ignored altogether rejects a true null 3.523% of the time; a binding “below zero” design, 2.851%; a binding design-effect design, 2.518%.

The error rate of a futility design when its futility stops are overruled. Exact. stop if z < 0, binding: 2.500% when obeyed, 2.663% when half its stops are overruled, 2.851% when all are; power at the design effect < 10%, binding: 2.500% when obeyed, 2.509% when half its stops are overruled, 2.518% when all are; power at the trend < 10%, binding: 2.500% when obeyed, 2.898% when half its stops are overruled, 3.523% when all are; power at the trend < 10%, non-binding: 1.761% when obeyed, 2.048% when half its stops are overruled, 2.500% when all are.
Fig. 7 The chance of crossing for benefit with no effect, against the share of futility stops that are overruled, for the three binding designs and the non-binding trend-rule design.

Committees do not ignore futility rules altogether; they overrule some of them — for a subgroup that looks promising, a secondary endpoint, a sponsor’s case that the next look will be different. Overruling half its futility stops takes the binding trend-rule design to 2.898% and the binding “below zero” design to 2.663%; the design-effect design barely moves, to 2.509%, because it had few stops to overrule. The non-binding trend-rule design, overruled half the time, spends 2.048%, and overruled every time it spends exactly the 2.5% it was solved for — never more, however often it is overruled.

So the choice between binding and non-binding is not a technicality in the protocol. It is a choice between three points of power and an error rate that depends on the conduct of a committee nobody outside the trial observes.

Two routes to the same rows

Power and observations under four futility rules, counted over 80,000 trials and computed exactly. no futility rule: power 88.45% exact, 88.64% counted; 297.2 observations exact, 297.0 counted. stop if z < 0: power 83.25% exact, 83.36% counted; 273.3 observations exact, 272.9 counted. power at the design effect < 10%: power 88.33% exact, 88.52% counted; 294.6 observations exact, 294.3 counted. power at the trend < 10%: power 75.22% exact, 75.53% counted; 244.5 observations exact, 244.9 counted.
Fig. 8 Power at a real effect of 0.16 and the mean number of observations under each rule, computed exactly and counted over eighty thousand trials a rule.

Every number above is exact: the density of the partial sum carried from look to look on the region above the futility boundary, with the grid split at the boundary, where a share of carried-on trials makes the density jump. Halving or doubling the grid leaves the error rates unchanged to five decimal places.

The same four designs were also run observation by observation, eighty thousand trials each, on a block of seeds fixed before any was drawn. Counted power is 88.64% with no futility rule, 83.36% under “below zero”, 88.52% under the design-effect rule and 75.53% under the trend rule, against 88.45%, 83.25%, 88.33% and 75.22% exact; the counted mean numbers of observations are 297.0, 272.9, 294.3 and 244.9 against 297.2, 273.3, 294.6 and 244.5. Each count sits within about two of its standard errors, which at eighty thousand trials is 0.11 to 0.15 points of power and 0.3 observations.

What a futility rule should state

Whether it binds. A non-binding rule and a binding one on the same boundary are different designs with different benefit boundaries, different power and different error rates, and a report that names the rule without saying which has not described the trial.

The effect its conditional power assumes. “Conditional power under 10%” is not a rule until the effect is named, and the two natural choices here differ by thirteen points of power at the effect the trial was built to detect.

What a futility stop is not. A trial stopped for futility at eighty observations has not shown the effect to be absent. Its estimate is truncated from above — it stopped because its estimate was low — and its interval, from eighty observations, is wide enough to contain the effect it was looking for. An absence of evidence is a number with a width, and a futility stop reports the absence and usually omits the width. In the orderings of a stopped trial’s outcomes every futility stop ranks with the least extreme outcomes, which is right for a p-value and says nothing about how large an effect the stopped data could still contain.

What is exact here, and what is counted

Exact. The benefit and futility boundaries; every error rate, power, share of futility stops and mean number of observations, obeyed or overruled, binding or not; and the curves across effects. Quadrature on panels split at each futility boundary, four hundred and one points a look.

Counted. The forty trials in the first figure and the eighty thousand a rule in the last.

Particular to this trial. One benefit boundary, five equal looks, a one-sided 2.5%, a normal outcome with known variance, and conditional power computed at the last look alone, which is how it is computed in practice and ignores the interim benefit looks still to come. Other futility rules, spent from a function of their own, would give other numbers; the trade between what a rule saves with no effect and what it costs with one is the part that carries.

Still open: a look the trend asked for

Every schedule in this series was fixed before the trial began: five looks, at eighty, a hundred and sixty, two hundred and forty, three hundred and twenty and four hundred observations. Spending the error rate named the design that removes that constraint — a spending function, which computes each look’s boundary from the fraction of information reached, so a committee may look whenever it meets — and named the one freedom it does not grant: choosing when to look because of what the trend looks like. Futility rules make that temptation concrete, since a trial near a futility boundary is exactly the trial a committee wants to look at again soon. A look the trend asked for computes what that choice spends: the error rate of a committee that adds a look when the interim result is close, and the worst any committee choosing among a menu of schedules could do.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formConditional powerError rateFutilityGroup-sequential designInterim analysisMonte CarloO'Brien–FlemingSample sizeStagewise orderingStatistical powerStopping rule