The effect a stopped trial reports
Worth reading first: When the looking happens.
Spending the error rate across five looks repairs the rate at which a trial rejects a true null, and the repair is exact. Four hundred observations, looked at after every eighty, against an O’Brien–Fleming boundary of 4.562, 3.226, 2.634, 2.281 and 2.040: the two-sided rate is 5%, where testing at the nominal level at every look was 14%. At a true effect of 0.16 per observation the same trial crosses for benefit 88.45% of the time and uses 297.2 observations on average. A recursion over the partial sum returns those numbers to four places, where twenty thousand simulated trials had given 88.8% and 297.6.
None of them is the number a trial reports. A trial that stops publishes an estimate of the effect — the mean of the observations it collected — and the boundary was never asked to do anything about that. It was built to decide when to stop. Deciding when to stop on the data is a selection on the data, and the estimate is computed from the same data after the selection has been made.
A boundary on z is a floor on the effect
A look after n observations stops the trial when reaches the boundary b. Written in the units of the estimate instead, it stops when reaches . For O’Brien–Fleming’s five looks that is 0.510 at eighty observations, 0.255 at a hundred and sixty, 0.170 at two hundred and forty, 0.128 at three hundred and twenty and 0.102 at four hundred.
Set those beside the truth. At each of the first three looks the smallest effect a trial can stop on is larger than the effect it is estimating — more than three times larger at the first, sixty per cent larger at the second, six per cent at the third. So no trial that stops at any of those looks can report the true effect. Every one of them reports more, and nothing about any of them was done wrong: the boundary is the correct boundary, the observations are honest draws, and the estimate is the ordinary mean of what was observed.
That is the whole mechanism, and it is worth seeing before any averaging, because averaging hides it. The rule “stop when the evidence is strong enough” is the rule “stop when the estimate is large enough”, drawn on a different axis. In the figure the trials that stop early are the ones whose running estimate happened to be high at a look; the trials running low carry on, and the boundary falls to meet them later. A trial’s stopping time and its estimate are two readings of one random walk, and the rule couples them.
What the trials that stop at each look report
The boundary is a floor, and the trials that clear it do not land on it. Among the trials that stop at each look the mean reported effect, computed exactly, is 0.5407 at eighty observations — 3.38 times the truth — on 0.09% of trials; 0.2931 at a hundred and sixty, 1.83 times, on 11.39%; 0.2065 at two hundred and forty, 1.29 times, on 32.70%; 0.1580 at three hundred and twenty, 0.99 times, on 28.55%; and 0.1274, 0.80 times, among the 15.73% that cross only at the last look. The trials that reach four hundred observations without crossing at all report 0.0763, less than half the effect they are about.
The pattern is monotone and it is not subtle. The earlier a trial stops, the more it reports, and the trials that run long report less. A reader holding one stopped trial knows which row it came from, because the look at which a trial stopped is the one fact every stopped trial announces — and the row fixes the size of the overstatement far better than the confidence interval around the estimate does.
The fourth look deserves a second reading. Its mean of 0.1580 is almost exactly the truth, and it would be a mistake to take that as the look where the estimate is honest. It is the look where two selections cancel. The trials that reach it were selected for having been low enough not to stop at the first three looks, and the ones that stop there are selected for being high enough to cross; both are in force, and at this effect they happen to balance. At an effect of 0.12 the same look reports 1.27 times the truth, and at 0.20 it reports 0.82 times. A row that is right at one effect is a coincidence of two biases, and a coincidence does not travel.
Every group is wrong and the whole is nearly right
Weighted by how often each happens, the rows average to 0.1753: a bias of 0.0153, or 9.6% of the true effect. That is a real bias and a modest one, and the density shows why it is modest. The curves on either side of the truth largely offset each other. The trials that stop early are fewer and far above; the trials that end at the last look are many and below.
A mean of about the right size assembled from groups each of which is wrong is what makes this bias easy to miss. Anybody who sees every trial of the design, whatever each did, sees estimates spread around something close to the truth and concludes that the design is nearly unbiased. Anybody who sees one trial sees a draw from one of the curves, and it is never the average curve.
The density makes a second point the means cannot. A trial stopped at a hundred and sixty observations cannot report anything below 0.255, whatever the truth was. Its estimate is not a noisy reading of the effect with an error that could go either way; it is a reading truncated from below, at a place fixed by the design before any data arrived. An interval quoted around that estimate as though the sampling distribution were a normal centred on the truth describes a distribution the trial did not have.
The total is right and the ratio is not
There is an exact statement of where the bias lives, and it matters because it says which repair works.
Write N for the number of observations when the trial ends and S for their sum. The running difference S − θn is a sum of independent increments with mean zero, and the trial ends at a time bounded by four hundred observations, so its expectation when the trial ends is zero. That is Wald’s identity:
At the design effect both sides are 47.5565 — 0.16 times the 297.23 observations a trial uses on average — and the identity holds whatever the stopping rule is. What it does not say is that E[S/N] = θ, and S/N is what each trial reports. The ratio of the expectations is the truth. The expectation of the ratio is 0.1753.
So pooling trials by size recovers the effect. Summing every observation across every trial of the design, stopped or not, and dividing by the total number of observations gives 0.16000. That is what an analysis does when it weights each trial by its sample size, and it is why a pooled analysis of a complete set of such trials is not dragged by the early stops the way each trial’s own report is. The early stops are large estimates from small trials, and weighting by size gives them the small weight a small trial should have.
The repair has a condition, and the condition is the complete set. Pool only the trials that crossed for benefit and the size-weighted estimate is 0.1754, which is no better than the plain average across every trial. The selection is now on the outcome rather than on the timing, and Wald’s identity says nothing about a subset chosen by its result.
Twenty thousand trials against the same numbers
Every number so far is an integral. The density of the partial sum at each look is the previous look’s density convolved with a normal increment, and a stopped group’s mean is that density integrated past the boundary. A calculation that conditions on thin regions — the first look holds less than one trial in a thousand — is exactly the kind whose errors do not show, so the same trials are also run observation by observation, twenty thousand times, and counted.
Across all of them the mean reported effect is 0.1756 ± 0.00046, against 0.1753 exact; among the trials that crossed for benefit it is 0.1881 ± 0.00043, against 0.1883. At each look the counts agree to within their own error: 0.2925 ± 0.0007 on the 2,339 trials that stopped at a hundred and sixty observations, against 0.2931; and 0.0756 ± 0.0005 on the 2,213 that never crossed, against 0.0763. Only eighteen of the twenty thousand stop at the first look, and they report 0.5384 ± 0.0055 against 0.5407.
Two routes to the same number earns its keep here for a specific reason. The exact route is where a mistake in the conditioning would hide — a boundary applied on the wrong scale, an increment variance off by one look — and the counted route shares nothing with it except the rule. They agree at every look, including the one eighteen trials reach, and that agreement is what licenses the rows above.
Where on the range of effects the bias lives
The 9.6% belongs to one effect. Computed across effects from nothing to 0.36 per observation, the bias rises, peaks and falls, and it does so differently for each boundary.
With no interim looks it is zero everywhere, because the mean of a fixed four hundred observations is unbiased whatever it turns out to be. At no effect it is zero under every design, because the design is symmetric: a trial is as likely to stop on the lower boundary as on the upper, and the two truncations cancel. Under O’Brien–Fleming it climbs to 0.0163 at an effect of 0.20, which is 8.2% of it, falls to 0.0131 at 0.33, and turns up again at 0.36 as the first look begins to take trials. Under Pocock it climbs to 0.0350 at 0.19 — 18.4% of the effect, more than twice O’Brien–Fleming’s worst.
The rise is the selection gaining force: at small effects few trials stop early, so few estimates are truncated. The fall is the selection losing its grip in a different way. At large effects nearly every trial stops at one of the same one or two looks, and a truncation that removes little of an estimate’s distribution moves its mean little. At 0.36 the first look, which needs 0.510, takes 8.98% of O’Brien–Fleming trials and the second takes 81.81%; most trials report from a curve whose floor of 0.255 sits well below where they are.
As a share of the effect the bias is large where the effect is small: 6.5% at 0.05 under O’Brien–Fleming, where only one trial in six crosses. That is the region a trial enters when the effect assumed at the design stage was optimistic, and the effect a design assumes and the effect a trial meets are not the same number. The bias is a function of the second.
The trials that get written up
A positive trial is not a random trial. The estimate that reaches a guideline, a press release or a review restricted to significant results is the estimate among trials that crossed for benefit, and that group carries two selections at once: the one a fixed design has too, and the one the stopping adds.
At the design effect the trials that cross report 1.065 times the truth with no interim looks, 1.177 times under O’Brien–Fleming and 1.346 times under Pocock. The first of those is the winner’s curse at 89% power, small because nearly every trial crosses. The gap between it and the other two is what looking added to the report, and it is there in trials whose error rates are exactly right.
At an effect of 0.04, where few trials cross, the multiples are 3.07, 3.57 and 4.96. A trial that crosses there got lucky under any design. A trial that crosses early there got lucky at a moment when luck was large, with few observations behind it, and its report carries both. It is the same shape as the estimate of an arm chosen for being ahead and the arm carried forward after the losers were dropped: a quantity used to make a decision, reported afterwards as though no decision had been made with it.
Pocock’s early spending moves the bias forward
The two classical boundaries spend the same 5% and differ in when they spend it. The estimate inherits the difference.
Pocock’s constant boundary of 2.413 is a floor of 0.270 on the effect at eighty observations — well below O’Brien–Fleming’s 0.510, and so reachable by far more trials. 16.30% of Pocock trials stop at the first look, against 0.09%, and they report 2.06 times the truth. Across the design the Pocock trial stops sooner, on 253.1 observations on average rather than 297.2; crosses for benefit less often, 82.45% against 88.45%; and reports an effect 20.7% too large where O’Brien–Fleming’s is 9.6% too large.
So the property that makes O’Brien–Fleming the usual default — that it spends almost nothing early — has a second benefit it was not chosen for. A boundary that is very hard to cross early rarely truncates an estimate built from a small sample, and the few estimates it does truncate belong to so few trials that they barely move the mean. Stopping sooner is paid for twice: once in power at the last look, which the boundaries were compared on, and once in the number the stopped trial reports.
What a stopped trial should say about its estimate
Three things follow for a report, and none of them needs a new calculation.
The look at which the trial stopped, and the boundary in force there. The protocol facts that make a p-value readable make an estimate readable too. A trial that stopped at a hundred and sixty of four hundred observations reports an estimate that could not have been below 0.255, and a reader who knows that knows which curve the number was drawn from.
The bias at the design effect, computed before the trial starts. It is a property of the boundary and the effect assumed in planning, and it is the same kind of integral as the power calculation the protocol already contains. A design that expects to stop early should say how much its early stops will overstate.
The estimate kept apart from the decision. The boundary answered whether to stop. It did not answer how large the effect is, and the mean at the stop answers that second question with a bias that depends on the answer to the first. The same holds for a trial that stops when its estimate of the spread is small and for a simulation stopped when its estimate reaches a target: each stops on something computed from the data and reports something computed from the same data, and the report is conditioned on the stop.
None of this is an argument against stopping early. A trial that stops at the second look on a real effect has spared three fifths of its planned observations, and in a clinical trial those are patients. The argument is narrower: the decision was right and the number attached to it is high, by an amount that can be stated in advance.
What is exact here, and what is counted
Exact. The boundaries, solved by bisection on the recursion and equal to the published 2.413 and 2.040; every probability of stopping at a look; every mean reported at a look; the overall bias and its curve across effects; Wald’s identity at the design effect; the pooled 0.16000 and the pooled 0.1754 among trials that crossed. Exact here means exact up to quadrature on four hundred and one points a look, and halving or doubling that grid leaves the O’Brien–Fleming constant unchanged to six decimal places.
Counted. The forty paths in the first figure and the twenty thousand trials in the fifth, which agree with the exact means at every look.
Particular to this design. Four hundred observations of unit variance, five equally spaced looks, a two-sided 5% boundary, a normal outcome and a known variance. The mechanism — a boundary on z is a floor on the estimate — does not depend on any of these. Every number does.
Still open: an estimate that knows the rule
The mean at the stop is biased because it ignores the rule, and the obvious response is an estimate that does not. There are several and they disagree. A conditional estimate asks what effect would make the observed mean the average among trials that stopped at this look. A median-unbiased estimate asks what effect would put the observed outcome in the middle of the outcomes the trial could have produced.
The second needs an ordering of those outcomes. Is a trial that stopped at the second look with a z of 3.3 more extreme than one that stopped at the fourth with a z of 2.5, or less? A fixed design never has to answer, because every trial ends at the same place and a larger z is simply more extreme. A sequential design has to answer before it can report a p-value, a confidence interval or an adjusted estimate. The outcomes a trial could have stopped with takes that question up: four natural orderings give four different p-values for one stopped trial, and only one of them can be computed without the boundaries at looks the trial never reached.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A look the trend asked for — both name closed form, group-sequential design, interim analysis, monte carlo, o'brien–fleming, the pocock boundary, statistical power, stopping rule
- A coverage table with its own error — both name closed form, monte carlo, sample size, statistical power
- The sample is a condition — both name closed form, conditional distribution, monte carlo, selection bias
- Three mechanisms and one dataset — both name closed form, conditional distribution, monte carlo, selection bias
- A block size that changes — both name monte carlo, sample size, stopping rule
- A lead that a heavy tail keeps — both name closed form, monte carlo, selection bias
Named objects
A flat tag is an object no other essay names yet.
Closed formConditional distributionGroup-sequential designInterim analysisMonte CarloO'Brien–FlemingThe Pocock boundarySample sizeSelection biasStatistical powerStopping ruleThe winner's curse