A look added by a rule
Worth reading first: When the looking happens.
A look the trend asked for measured what a spending function does not cover. A trial of four hundred observations looks at half of them; a committee then adds a look at three quarters if the interim |z| is at least 1.5, and otherwise carries on to the end. Every boundary it uses is the one the O’Brien–Fleming-type spending function gives for the schedule it is on, and the trial spends 5.2323% rather than 5%, because which schedule it is on depends on the interim result.
The essay ended on the repair a protocol could make. The committee’s rule is a fixed function of the data once it is written down, so it can be given boundaries that spend exactly 5%. The added look and the final look on its branch would have to be higher than the spending function alone makes them, and the question left open was what that costs — whether raising the boundaries of trials that were already close, which are the trials most likely to have a real effect, gives up a few hundredths of a point of power or a whole point.
It gives up about four tenths of a point, whichever of three corrections is used. And the reason the cost is that size, rather than larger or smaller, says something about the uncorrected rule that the earlier essay’s numbers only hinted at.
Three ways to spend exactly 5%
The trial is the earlier essay’s: a normal outcome with known variance, four hundred observations, a first look at two hundred with the O’Brien–Fleming-type boundary 2.772, and then one of two schedules. Straight to the end has a final boundary of 1.979; a look at three quarters and then the end has boundaries 2.298 and 2.043. Each is what the spending function gives, and each, applied to every trial, spends exactly 5%. The committee’s rule sends a trial to the second schedule when its interim |z| is at least 1.5, and that is what spends 5.2323%.
A protocol that writes the rule down can change the boundaries so that the rule, not either schedule on its own, spends 5%. Three changes are natural.
Raise the added branch. Multiply the added look’s boundary and the final boundary on that branch by one factor, and leave the other branch alone. The factor that spends exactly 5% is 1.0205, which moves the boundaries from 2.298 and 2.043 to 2.345 and 2.085. The whole correction is charged to the trials that were close at the interim.
Raise everything after the first look. Multiply every later boundary, on both branches, by one factor. The factor is 1.0099, which moves the straight-to-the-end boundary from 1.979 to 1.999 and the added branch’s to 2.321 and 2.063. The correction is shared between the close trials and the rest.
Preserve the plan’s chance at every interim value. The planned schedule — the one the protocol would have used without the committee — gives each interim value a chance of crossing later. For each interim value that triggers the added look, raise that branch’s boundaries by whatever factor makes its chance of crossing later exactly equal to the planned one. Then, interim value by interim value, the rule spends what the plan spent, and the total is 5% by construction whatever the trigger. This is the conditional-error principle that adaptive designs use to change a trial’s course safely, applied to the smallest change there is: one look.
Each is computed exactly, by the backward recursion the earlier essay used for the chance of crossing later from every interim value, with the factors found by bisection; each spends 5.0000% to the digits shown.
What each costs
At an effect of 0.16 per observation, the effect the trial was sized for, the uncorrected rule has power 88.94%. Raising the added branch gives 88.56%; raising every later boundary, 88.55%; preserving the planned chance at each interim value, 88.56%. The three corrections differ by a hundredth of a point, and each gives up 0.38 points relative to the uncorrected rule.
The expected number of observations barely moves either: 304.9 uncorrected, and 306.0, 305.4 and 306.3 under the three corrections. Raising a boundary makes a trial a little less likely to stop at three quarters, and a trial that does not stop there runs to the end, so every correction costs a few observations as well as a little power. With no effect the corrections use 397.6, 397.5 and 397.6 observations against the uncorrected rule’s 397.4, since a trial with nothing to find rarely stops early under any of them.
That the three agree has an explanation that fits, though only these three corrections are evidence for it. Each removes the same 0.2323 points of excess error, and at an effect the trial is sized for, error removed near a boundary and power removed near it come from nearly the same trials: the ones whose path hovers around the boundary at three quarters. Where the correction is charged — to the close trials alone, to all of them, or interim value by interim value — changes which of those trials lose their crossing, and hardly changes how many.
The power that was never there
The number worth dwelling on is the uncorrected rule’s 88.94%. Beside it, the plan straight to the end has 89.01% power at 338.9 expected observations, and a look at three quarters for every trial has 88.38% at 300.7. The uncorrected rule, at 304.9 observations, appears to have almost the power of the plan that never stops early and almost the economy of the plan that stops as early as it can. That combination is what makes the committee’s instinct feel like good design: looking again at a close trial seems to buy speed without costing power.
The correction says where that combination came from. Corrected, the rule has 88.56% power at 306.3 observations — between the two fixed schedules on both counts. The 0.38 points the correction removed were the power the rule had bought with its 0.23 points of excess error, and none of it was a property of the timing. A rule that spends more than 5% has more power than one that spends 5% for the same reason a one-sided test at 6% has more power than one at 5%, and comparing it with fixed schedules at 5% was comparing across error rates.
A point of error for a point and a half of power
The same correction can be made at every trigger, and the pattern it shows is steadier than any one trigger’s numbers.
The excess error rises and falls with the trigger, from 0.059 points at a trigger of 0.5 to its peak of 0.232 at 1.5 and back to 0.058 at 2.5, as the earlier essay found: a trigger at either end adds the look for almost every trial or almost none, and a schedule applied to every trial or to none spends exactly 5%. The power the correction takes back follows it, from 0.075 points to 0.379 and back to 0.096, and the ratio of the two hardly moves: 1.59, 1.63, 1.63 and 1.64 points of power for each point of excess error from a trigger of 1 upwards, and 1.26 at 0.5.
That ratio is the exchange rate between error and power at this effect and this design, and it is the reason the question the earlier essay left — a few hundredths of a point or a whole point? — has the answer it has. However the committee’s rule is set, its correction costs about one and a half times the error it removes, and the most any single-look rule on this trial spends above 5% is about a quarter of a point. A monitoring rule that added more error would cost correspondingly more power to make exact, and the earlier essay’s worst case — a committee choosing among six schedules at every interim value, 0.44 points above 5% — would, if the same exchange rate held, cost about seven tenths of a point — an extrapolation from the ratio, not a measurement. The same exchange rate prices every data-triggered look this series has measured, and it says that the look’s apparent power was never cheap: it was paid for in error at a rate of about two for three.
Where the exact rule sits
The hero figure’s curve is the corrected rule at five triggers, and the grey line is what a protocol could get without any rule at all by flipping a coin between the two fixed schedules in the right proportion. Every point of the corrected curve is above the line, by 0.017 points of power at a trigger of 0.5, 0.091 at 1.5 and 0.138 at 2.0.
So the committee’s rule, once written down and made exact, is a genuinely better design than mixing the fixed schedules — slightly. It sends to the extra look exactly the trials for which a look at three quarters is most likely to end the trial early, which is the use of data a design is allowed to make when the error rate is accounted for. Under Pocock-type spending the same holds with a larger margin at a trigger of 1.5, 84.86% against 84.69% on the line, because more of that function’s budget has been spent by half the observations and the choice of branch matters more for what is left. The gain is small in either case, a tenth or two of a point at the same expected size, and it is the whole of what the committee’s instinct was worth once the error rate stopped paying for it.
The boundary a close trial is given
The preserved-chance correction does not raise the added look’s boundary by a fixed amount. For a trial triggered at an interim |z| of 1.50 it barely changes it — to 2.290 from the spending function’s 2.298, fractionally lower, because for a trial that far from the first boundary a look at three quarters is worth slightly less than going straight to the end and the correction gives some of it back. For a trial whose interim |z| was 2.77, right at the first boundary, it raises it to 2.459. The trials closest to crossing are asked for the most, because they are the ones for which the extra look adds most to the chance of crossing, and preserving the plan’s chance means taking that addition back.
This is the sense in which the lead’s worry — a penalty aimed exactly at the trials most likely to have a real effect — is right and does not matter. The correction is aimed at those trials; it has to be, since they are where the excess error came from. But the trials most likely to have a real effect are also the ones most likely to cross even a raised boundary, and the power lost is the 0.38 points of the table rather than anything larger. The correction that spreads the penalty over every later look costs the same, which says the targeting is not what costs power.
What a protocol can write
A rule for adding looks can be part of the design. Written as a function of the interim result, it can be given boundaries that spend exactly the nominal error rate, and the three corrections here are each a sentence of protocol: a factor for the added branch, a factor for every later look, or the plan’s chance preserved at each interim value.
The preserved-chance version is the one to prefer. It costs the same power as the others and it is the only one whose error rate does not depend on the trigger being fixed in advance. Because it spends, at every interim value, exactly what the plan spent there, a committee could change the trigger — or add the look for reasons that have nothing to do with a threshold — without disturbing the error rate, provided the added look is given the preserved-chance boundary for the interim value the trial actually had. It is the rule that makes the committee’s discretion safe rather than the rule that removes it, and it is the same principle choosing n after looking arrives at for changing a sample size.
Say what the look reads. The earlier essay pointed to a schedule that reads the mean as the fixed-width trial’s version of this problem: block sizes may be any function of the contrasts and no function of the mean. A monitoring plan has the same division. A look triggered by the interim effect needs one of the corrections here; a look triggered by recruitment, by calendar or by a blinded quantity may not, and a plan that allows both kinds of look should say which kind each is, because only the first carries a price and only the report can say which was paid. The outcomes a trial could have stopped with adds the analysis’s half of the same point: a p-value computed after a data-triggered look needs the rule that triggered it, corrected or not, to be written down.
Report the power of the corrected rule, not the uncorrected one. The uncorrected rule’s power is higher by 0.38 points and none of that is real. A protocol comparing monitoring plans by power should compare them at the same error rate, which means correcting every data-triggered plan before comparing it.
The earlier essays in this series set the same boundary from other sides: spending the error rate granted the freedom to time looks independently of the data, when the looking happens priced looks taken at the nominal level instead, and a boundary for giving up found the same accounting for stopping early for futility. A looked-for look, a futility stop and a cancelled look are all branches of one procedure, and a correction of the kind here is what makes any of them part of the design rather than a departure from it.
The rates, powers and expected numbers of observations are exact: each is an integral over the first look’s continuation region of the chance of crossing later, or of the observations used, from each interim value, computed backwards look by look on a grid of 801 interim values and 201 nodes at each later look. The two uniform factors are found by bisection on the error rate; the preserved-chance factor by bisection at each of the interim values beyond the trigger. Not measured: a committee that can choose again at its second look, whose rule has more branches than this; a trigger based on anything other than the interim |z|, such as a conditional-power estimate, which is a function of it here but not in general; and the counted version of these rules, which the earlier essay ran for the uncorrected one and found within one and a half standard errors of exact.
Still open: a look added on a nuisance quantity
Every trigger here reads the effect: the look is added because |z| is large. Choosing n after looking and blinded, and still exact found that an adaptation reading only a nuisance quantity — a pooled variance, an event rate across both arms — leaves the error rate alone. A monitoring committee often wants to add a look for exactly such a reason: recruitment faster than planned, an event rate higher than expected, a variance larger than assumed, each of which changes when the information will arrive.
Whether a look added on a blinded quantity keeps the spending function’s guarantee exactly, or whether the correlation between a blinded quantity and the effect in a finite trial leaves a residue like the ones this series has measured, is a calculation of the same kind as these. It would say whether the one kind of monitoring decision committees most often make for administrative reasons is also the one that needs no correction.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The effect a stopped trial reports — both name closed form, group-sequential design, interim analysis, o'brien–fleming, the pocock boundary, statistical power, stopping rule
- A simulation that stops when it looks settled — both name closed form, error rate, interim analysis, statistical power, stopping rule
- An order that spends the error rate — both name closed form, error rate, statistical power
- Dropping the losers — both name adaptive design, error rate, interim analysis
- Randomising towards the winner — both name adaptive design, error rate, statistical power
- The estimate after the choice — both name adaptive design, error rate, interim analysis
Named objects
A flat tag is an object no other essay names yet.
Adaptive designAlpha spendingClosed formError rateGroup-sequential designInterim analysisO'Brien–FlemingThe Pocock boundaryStatistical powerStopping rule