Stopping rules

A look the trend asked for

Under an O'Brien–Fleming-type spending function, every schedule of looks fixed in advance spends exactly 5.0000%. A committee that adds a look at three quarters of the trial whenever the interim z is 1.5 or more spends 5.2323% — 5.315% counted over a hundred thousand trials — and the most a committee choosing among six schedules could spend is 5.4390%.

Worth reading first: When the looking happens.

Spending the error rate ended on a distinction it stated and did not measure. A spending function lets a trial’s interim analyses happen whenever a monitoring committee meets, at whatever fraction of the information has arrived, because each look’s boundary is computed from what the function allows to have been spent by then. What it cannot accommodate, that essay said, is choosing the timing on the results: an extra analysis held because the trend looks promising “reintroduces exactly the problem the design was built to prevent, and no spending function corrects for it.”

Both halves are checkable, and the second has a size. This essay measures it on the trial the series has used throughout — four hundred observations, two-sided 5% — with a first look at half of them, two hundred observations, and then whatever the committee decides to do next.

The chance of crossing later from each interim z, under six schedules with O'Brien–Fleming-type spending boundaries. Exact. At an interim |z| of 1.0: end only 3.64%, +0.6 3.41%, +0.75 3.20%, +0.9 3.41%, every 0.125 3.00%, every 0.05 2.88%. At 2.5: end only 38.30%, +0.6 45.80%, +0.75 45.36%, +0.9 41.70%, every 0.125 50.08%, every 0.05 53.24%. The heavy line is the largest of the six at each z.
Fig. 1 From a look at half the observations, the chance that a trial with no effect goes on to cross at a later look, against its z there, for six schedules of later looks. The heavy line is the largest of the six at each z.

Any schedule chosen in advance spends exactly 5%

An O’Brien–Fleming-type spending function allows 2(1Φ(1.96/t))2\,(1 - \Phi(1.96/\sqrt{t})) of the error rate to have been spent by the time a fraction t of the observations is in. At half the observations that is 0.557%, and the first look’s boundary is the z that spends it: 2.772.

After that first look, six schedules are on offer: one more look at the end; a look at six tenths of the observations and then the end; one at three quarters and then the end; one at nine tenths and then the end; a look every eighth of the trial; and a look every twentieth. Each has its own boundaries, solved one look at a time from what the function allows at that look minus what the earlier looks spent. One more look at the end is 1.979. A look at three quarters and then the end is 2.298 and 2.043. A look every twentieth runs from 2.723 at the second look down to 2.167 at the last.

Six schedules after a first look at half the observations, each with O'Brien–Fleming-type spending boundaries. Exact. The first look's boundary is 2.772 in every schedule. end only: 2.772, 1.979, spending 5.0000%; +0.6: 2.772, 2.593, 2.001, spending 5.0000%; +0.75: 2.772, 2.298, 2.043, spending 5.0000%; +0.9: 2.772, 2.090, 2.088, spending 5.0000%; every 0.125: 2.772, 2.535, 2.351, 2.212, 2.104, spending 5.0000%; every 0.05: 2.772, 2.723, 2.636, 2.554, 2.481, 2.414, 2.355, 2.301, 2.252, 2.208, 2.167, spending 5.0000%.
Fig. 2 The boundaries of the six schedules against the fraction of observations reached, all starting from the same first look.

Carrying the density of the partial sum through the looks, each of the six spends 5.0000% — to every digit shown, which is the check that the boundaries and the recursion agree. That is the whole of the freedom a spending function grants, and it is a real freedom: a committee may meet at a hundred and twenty observations or at three hundred and sixty, twice or ten times, and the trial holds its level. What it has to be is chosen by something other than the data: a calendar, a recruitment milestone, a committee’s diary.

Why a committee would want to choose

The first figure is the reason the freedom gets used differently in practice. From the look at half the observations, what matters is the chance that the trial goes on to cross at some later look, given the z it has there, and that chance depends on the schedule.

With no effect and an interim |z| of 1.0, one more look at the end crosses 3.64% of the time, a look at three quarters and then the end 3.20%, a look every twentieth 2.88%. A trial far from its boundary is best served by the fewest looks: every extra one spends a little of the remaining error rate at a moment the trial has almost no chance of using it, and raises the boundary where it might.

Close to the boundary it is the other way round. At an interim |z| of 2.0, one more look at the end crosses 21.29% of the time and a look every twentieth 24.45%. At 2.5 it is 38.30% for one more look, 45.36% for a look at three quarters, and 53.24% for a look every twentieth. A trial that is nearly there gains most from being looked at soon and often, because the next look catches it while it is still near the line.

Every curve is correct. Any one schedule, applied whatever the interim z turned out to be, averages its curve against the distribution of that z and comes to exactly 5%. The heavy line — the largest of the six at each z — is not a schedule. It is what a committee gets by choosing one after seeing where the trial is, and the arithmetic of the next three sections is the arithmetic of that line.

A committee that looks again when it is close

The mild version is a rule a reasonable committee might follow without thinking of it as a rule. At half the observations, if |z| is at least 1.5 — promising, not yet across 2.772 — add a look at three quarters; otherwise carry on to the end.

With no effect, the look is added in 12.96% of trials, and the trial’s error rate is 5.2323%: 0.23 points above the level. No decision inside it is wrong. Every look it takes has the right spending boundary for the schedule it is on, and a reader of its report sees a spending function, a set of looks and a boundary at each, all correctly computed. What is wrong is that which schedule it is on is a function of the interim z, and the error rate of a procedure is a property of every branch it could have taken. A fixed-width trial meets the same rule in its block sizes, which may be anything provided they are functions of the contrasts, and two natural schedules that read the mean break it.

The error rate of adding a look when the interim trend is close, against how close. Exact. O'Brien–Fleming-type spending: 5.0000% when every trial or no trial adds the look, largest at 5.2347% when it is added from |z| = 1.6; Pocock-type spending: 5.0000% when every trial or no trial adds the look, largest at 5.1606% when it is added from |z| = 1.5.
Fig. 3 The error rate of adding a look at three quarters whenever the interim |z| is at least a trigger, against the trigger, under both spending functions.

The size depends on where the trigger sits, and it has to vanish at both ends. Adding the look for every trial is simply the three-quarters schedule, and adding it for none is one more look at the end; both are fixed and spend 5.0000%. Between, the rate rises to a peak of 5.2347% with the trigger at 1.6, where the look is added for almost exactly the trials it helps and withheld from the ones it would hurt. Under Pocock-type spending, whose early boundaries are lower and whose remaining budget is smaller, the peak is 5.1606%, with the trigger at 1.5.

The rule does not buy what it looks as though it should. At a real effect of 0.16 per observation its power is 88.94%, against 89.01% for one more look at the end. Whatever a committee hopes to gain by looking again at a close trial, it is not the chance of finding a real effect; it is the chance of finding it sooner, and the cost is paid in the error rate rather than in power.

Cancelling a planned look is the same rule

The committee in the last section added a look. The more common story runs the other way. A protocol plans a look at three quarters of the trial, and when the trial reaches half its observations with a trend that is going nowhere, the committee cancels it: there is nothing to see, an analysis costs time and money, and every look unblinds somebody.

Written down, the two stories are one procedure. The planned-then-cancelled trial takes the look at three quarters when |z| at half the observations is at least the level below which the committee cancels, and goes straight to the end otherwise. That is, word for word, the rule that adds a look from the same level, and its error rate is the same number — 5.2323% with the line drawn at 1.5.

It is worth dwelling on because the two stories do not feel alike. Adding a look to a promising trial sounds like a committee helping its result along; cancelling a look on an unpromising one sounds like a committee saving effort. The first figure says why they are the same. At an interim |z| of 1.0 the look at three quarters lowers the chance of crossing later from 3.64% to 3.20%, so a committee that cancels it for such trials raises their chance a little, exactly as a committee that adds it for a trial at 2.5 raises that trial’s chance a lot. Whether a decision is described as more monitoring or less, the error rate reads only which schedule each interim value was sent to.

The most a committee could spend

The schedule that spends the most, at each interim z, under O'Brien–Fleming-type spending. Exact. For |z| up to the first look's boundary of 2.772, the committee that spends the most sends the trial to: end only for |z| from 0.00 to 1.42; +0.9 for |z| from 1.43 to 1.73; every 0.05 for |z| from 1.74 to 2.77. The trial's error rate is then 5.4390%.
Fig. 4 The distribution of the interim z with no effect, shaded by the schedule that gives the largest chance of crossing later at each value — the choice of a committee that spends as much as it can.

The worst case asks what a committee could reach by choosing, at every interim value, whichever of the six schedules gives the largest chance of crossing later. The error rate is then 5.4390%.

The choice has three bands. The trial goes to one more look at the end for an interim |z| up to 1.42; to a look at nine tenths and then the end from 1.43 to 1.73; and to a look every twentieth from 1.74 up to the first look’s boundary. The bands where the committee looks more carry little probability — a trial with no effect has |z| above 1.43 at the first look about one trial in seven — and that is why the inflation is fractions of a point rather than whole points. Under Pocock-type spending the worst is 5.3233%.

It is worth being clear about what this worst case is. It is not a committee acting in bad faith. It is the error rate of the reasoning “this trial is close, so it should be looked at again soon”, taken to its limit on one menu of schedules — and every step of that reasoning is what a conscientious committee is asked to do.

What an early spender leaves to choose with

Pocock-type spending, αlog(1+(e1)t)\alpha \log(1 + (e - 1)\,t), has already spent 3.10% of the 5% by half the observations, against 0.557% under the O’Brien–Fleming-type function, and its first-look boundary is correspondingly lower, 2.157 against 2.772. Every number a committee can move is smaller under it: the rule that adds a look from 1.5 spends 5.1606% rather than 5.2323%, and the worst on the six schedules is 5.3233% rather than 5.4390%.

The explanation that fits is where the remaining budget sits. Every schedule a committee can choose after the first look is a way of spending what the function has not yet spent, and a function that front-loads its spending leaves less of it: 1.90% after half the observations under Pocock-type spending, 4.44% under the O’Brien–Fleming type. That is an explanation, and two spending functions are the only evidence for it here.

If it holds, it cuts against the usual preference in an unexpected place. The property that makes O’Brien–Fleming the usual default — almost nothing spent early, nearly everything held for the end — is also what leaves the most for a committee’s timing to act on. The price is small and it is real, and it is paid only by trials whose looks were chosen by their data.

A longer menu reaches further

The worst a committee can spend, as the schedules it may choose among grow. one look more: O'Brien–Fleming-type spending 5.000%, Pocock-type spending 5.000%. or a late one: O'Brien–Fleming-type spending 5.146%, Pocock-type spending 5.077%. or any of three: O'Brien–Fleming-type spending 5.259%, Pocock-type spending 5.191%. or four more looks: O'Brien–Fleming-type spending 5.366%, Pocock-type spending 5.260%. or ten more: O'Brien–Fleming-type spending 5.439%, Pocock-type spending 5.323%. or twenty more: O'Brien–Fleming-type spending 5.463%, Pocock-type spending 5.345%.
Fig. 5 The worst error rate a committee can reach, as schedules are added to the ones it may choose among, under both spending functions.

The worst case belongs to its menu. With only one more look at the end on offer there is nothing to choose and the rate is 5.000%. Offering a look at nine tenths as well brings the worst to 5.146%; any of three single looks, 5.259%; a look every eighth, 5.366%; a look every twentieth, 5.439%; and a look every fortieth — twenty more looks — 5.463%. Under Pocock-type spending the same menus reach 5.077%, 5.191%, 5.260%, 5.323% and 5.345%.

Each addition buys less than the one before. Nothing here computes the limit — a committee free to look continuously after the first look, choosing each next look on the data — and the sequence says only that the limit is above 5.463%. The shrinking steps suggest it is not far above, and that is a suggestion, not a result.

A hundred thousand trials against the same rates

Three monitoring rules, counted over 100,000 trials each and computed exactly. O'Brien–Fleming, never adds: counted 5.087% ± 0.069%, exact 5.0000%; a look added in 0.00% of trials. O'Brien–Fleming, adds from 1.5: counted 5.315% ± 0.071%, exact 5.2323%; a look added in 12.96% of trials. Pocock, adds from 1.5: counted 5.256% ± 0.071%, exact 5.1606%; a look added in 10.38% of trials.
Fig. 6 The error rate of three monitoring rules counted over a hundred thousand trials with no effect, beside the exact rate.

The same rules were run observation by observation, a hundred thousand trials with no effect each, on one block of seeds. The O’Brien–Fleming-type trial that never adds a look rejects 5.087% ± 0.069, against exactly 5%. The one that adds a look from |z| = 1.5 rejects 5.315% ± 0.071, against 5.2323%. The Pocock-type trial that does the same rejects 5.256% ± 0.071, against 5.1606%. Each count sits within one and a half of its standard errors.

The three counts share their trials, so they drift to the same side together, and the sharper comparison is the difference between the first two: counted, adding the look raises the rate by 0.228 points; exact, by 0.2323. The trials that never trigger the rule are the same trials in both rows, and the difference is the part of the count that belongs to the rule.

What the inflation is, and what it is not

It is small. Choosing when to look on the trend adds 0.23 points for a committee that looks again when a trial is close, and 0.44 points at the worst on a menu of ten more looks — against the jump from 5% to 14% that testing at the nominal level at five looks produces. The spending function has done most of its work, and a monitoring committee’s judgement about timing is not where most of a sequential trial’s risk lives.

It is not visible in the report. A published trial states its spending function, the fractions at which it looked and the boundary at each, and all of them will be correct. Nothing in that list says the third look was held because the second was close. The four orderings of a stopped trial’s outcomes make the same point from the other side: only the stagewise p-value can be computed without the schedule of looks the trial never took, and even that one assumes the looks it did take were not chosen from its data. So does the interval a precision rule reports after the stop it chose, which claims 95% and covers 90% with every number it prints correct.

It cannot be removed afterwards, and it can be removed in advance. A rule written into the protocol — add a look at three quarters if |z| at half the observations is at least 1.5 — is a fixed function of the data. Its error rate is the 5.2323% computed here, and its boundaries could be solved to spend exactly 5% instead, which makes it part of the design. The same decision made in a committee room, unrecorded, is a branch of a garden nobody mapped, and the forking paths that result have no correction because nobody knows how many there were.

The same structure runs through the rest of this series. A futility boundary is safe to overrule when it was never counted on and costly when it was; a look is safe to schedule freely when its timing ignores the data and costly when it does not. In both, the freedom that is safe is the freedom the error rate was computed to allow. A sample size re-estimated at an interim is the same distinction on a third dial: re-estimated blind to the arms it has a defence, and re-estimated from the effect it breaks the error rate.

What a monitoring plan should say

When looks are added, as a rule. A spending function is chosen so that a committee need not fix its meeting dates, and that is its purpose. If extra looks may also be requested because a result is close, the plan should say what “close” is, so that the procedure’s error rate is a number rather than an estimate of a committee’s habits.

Which looks were taken, and why. A report that lists the fractions at which a trial was analysed should say which of them were scheduled and which were requested, and on what grounds.

The spending function by name, and the first look’s boundary. Everything above depends on how much was spent before the choice was made — 0.557% under the O’Brien–Fleming-type function at half the observations — and a Pocock-type function, having spent more early, leaves less to be spent by choosing.

What is exact here, and what is counted

Exact. The spending boundaries of every schedule, each computed by bisection on the forward recursion; the chance of crossing later from every interim value, computed backwards look by look; every error rate, every trigger’s rate, the worst case on each menu, and the power at 0.16. The six fixed schedules’ 5.0000% is the check that the backward and forward computations agree.

Counted. A hundred thousand trials for each of the three rules in the last figure.

Particular to this trial. One first look, at half the observations; one choice, made there; a menu of named schedules; two spending functions; a normal outcome with known variance. A committee that can choose again at its second look has more branches than this, and its worst case is larger.

Still open: a rule for adding looks that spends exactly 5%

The rule a committee would actually write down — add a look at three quarters when the interim |z| is at least 1.5 — has an error rate of 5.2323%, and nothing stops a protocol from giving it boundaries that spend 5% instead. The added look, and the final look on the branch that has it, would each need to be higher than the spending function alone would make them, and by an amount that depends on the trigger. What that costs is the unanswered part: raising only the boundaries of trials that were already close is a penalty aimed exactly at the trials most likely to have a real effect, and whether the power it gives up is a few hundredths of a point or a whole point is a calculation the same backward recursion can make and has not made.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Alpha spendingClosed formError rateGroup-sequential designInterim analysisMonte CarloO'Brien–FlemingOptional stoppingThe Pocock boundaryStagewise orderingStatistical powerStopping rule