Concept

Group-sequential design — where it appears

A trial analysed at a few planned looks, with a boundary at each that stops it early if crossed. The boundaries are solved so that the looks together spend the stated error rate, and stopping on them changes what the estimate and interval at the stop mean.

Named by 5 essays across one field — each of them below, with the objects they name alongside it.

Also named here as o'brien–fleming — the same set of essays touches all of them, so they are one junction rather than several.

Four boundaries for 5 looks, all spending 5% in total. test at 0.05 every look: 1.96, 1.96, 1.96, 1.96, 1.96. Pocock — a constant, higher boundary: 2.41, 2.41, 2.41, 2.41, 2.41. O'Brien–Fleming — strict early, nearly nominal at the end: 4.55, 3.22, 2.63, 2.27, 2.03. Bonferroni across looks: 2.58, 2.58, 2.58, 2.58, 2.58. Every one except the first spends the same total error rate; they differ in when they spend it.

Spending the error rate

The repair for interim testing is to spend 5% across the looks rather than at each one. The boundaries are solvable rather than quotable, and a trial that can stop early uses 298 observations where a fixed design uses 400 — at a cost of half a point of power.

sequential · Stopping
Forty O'Brien–Fleming trials at a true effect of 0.16, with the boundary written as an effect. The dashed line is the smallest effect a trial can report and still stop at each look: 0.510 at 80 observations, 0.255 at 160 observations, 0.170 at 240 observations, 0.128 at 320 observations, 0.102 at 400 observations. The true effect is 0.16, so at 3 of the five looks a trial cannot stop without reporting more than it. 29 of these forty trials stop before the last look, each marked where it stopped.

The effect a stopped trial reports

An O'Brien–Fleming trial at 88.45% power holds its error rate exactly and reports an effect 9.6% too large on average. The 11.39% of trials that stop at the second look report 1.83 times the truth, the ones that cross at the last look report 0.80 times it, and pooling every trial by its size gives the truth back to the last digit.

sequential · Stopping
Ordered stagewise: the outcomes at least as extreme as stopping at 160 observations with z = 3.3. Each column is one look of an O'Brien–Fleming trial; above the boundary a trial stops there. Highlighted are the outcomes that count as at least as extreme as the observed one when outcomes are ordered stagewise: at 80, z ≥ 4.56 (probability 2.54 × 10⁻⁶ with no effect); at 160, z ≥ 3.30 (probability 4.82 × 10⁻⁴ with no effect); at 240, none; at 320, none; at 400, none. The two-sided p-value is 9.69 × 10⁻⁴.

The outcomes a trial could have stopped with

A trial that stops at its second look with z = 3.3 has a two-sided p-value of 0.000969, 0.000987, 0.00187 or 0.0421, depending on how the outcomes it could have stopped with are ordered. One of the four orderings does not change when the looks the trial never reached are replanned, and the same one gives a trial that ran to the end with z = 6 a p-value of 0.0256.

sequential · Stopping
Forty trials at a true effect of 0.16, under the rule "power at the trend < 10%". The upper line is the benefit boundary (4.56, 3.23, 2.63, 2.28, 2.04); the lower line is where the rule stops a trial for futility (0.40 at 80, 0.66 at 160, 0.95 at 240, 1.31 at 320). Of forty trials with a real effect, 29 cross for benefit and 11 are stopped for futility.

A boundary for giving up

Adding "stop if z is below zero" to an O'Brien–Fleming trial costs 5.20 points of power at the effect it was designed for and halves the observations a trial with no effect uses. Stopping when conditional power at the observed trend falls under 10% costs 13.23 points and stops 21.28% of trials with a real effect. Making that rule binding lowers the benefit boundary from 2.040 to 1.901, and a binding rule that is then ignored rejects a true null 3.523% of the time instead of 2.5%.

sequential · Stopping
The chance of crossing later from each interim z, under six schedules with O'Brien–Fleming-type spending boundaries. Exact. At an interim |z| of 1.0: end only 3.64%, +0.6 3.41%, +0.75 3.20%, +0.9 3.41%, every 0.125 3.00%, every 0.05 2.88%. At 2.5: end only 38.30%, +0.6 45.80%, +0.75 45.36%, +0.9 41.70%, every 0.125 50.08%, every 0.05 53.24%. The heavy line is the largest of the six at each z.

A look the trend asked for

Under an O'Brien–Fleming-type spending function, every schedule of looks fixed in advance spends exactly 5.0000%. A committee that adds a look at three quarters of the trial whenever the interim z is 1.5 or more spends 5.2323% — 5.315% counted over a hundred thousand trials — and the most a committee choosing among six schedules could spend is 5.4390%.

sequential · Stopping

Named alongside it

The objects these essays reach for when they reach for this one.

O'Brien–FlemingClosed formInterim analysisStopping ruleError rateMonte CarloThe Pocock boundaryStagewise orderingStatistical powerAlpha spendingConditional distributionp-value

All concepts