Stopping rules

The variance that contains the effect

A monitoring committee that adds an interim look because the variance came in higher than planned is reading a nuisance quantity, and a look added on a nuisance quantity should leave the spending function's 5% alone. It does, exactly, when the variance is the within-arm one, and when the trigger is the calendar. It does not when the variance is the one computed with the arm labels hidden: that variance is the within-arm sum of squares plus the interim statistic squared, so it is high more often when the trend is strong. The excess is 0.069 points at a twenty-patient interim and falls as one over the square root of the interim's size — a tenth of what a trigger on the trend itself spends, and never zero.

Worth reading first: When the looking happens · A width promised for a difference.

A look the trend asked for found that a committee adding a look at three quarters of a trial whenever the interim ∣z∣|z| is 1.5 or more spends 5.2323% under a spending function built to spend 5%, and a look added by a rule found three ways to give such a look boundaries that spend 5% exactly. Every trigger in those essays read the effect. The look was added because the trend was close.

Committees more often add a look for reasons that are not the effect. Recruitment ran faster than planned, so the information will arrive sooner. The event rate is higher than assumed. The outcome’s variance came in larger than the sample size calculation used, so the trial is less informative than it was meant to be at this point and a later look would be worth more. Each of these is a nuisance quantity, and the standing advice is that an adaptation reading only a nuisance quantity leaves the error rate alone. The question is whether a nuisance quantity can carry any of the effect. For one of the most common of them it can, and the reason is an identity.

How much a look added at three quarters spends above 5%, by what the decision to add it reads. A two-look O'Brien–Fleming-type trial, interim at half the information, adds a look at three quarters with a stated chance under the null. Its error rate above 5%, in points, at chances of 0.1, 0.3, 0.5, 0.7, 0.9: triggered by a calendar or by the within-arm variance, 0.00000, 0.00000, 0.00000, 0.00000, 0.00000; by the blinded variance with 20 patients at the interim, 0.034, 0.064, 0.069, 0.054, 0.022; with 100, 0.014, 0.027, 0.030, 0.025, 0.012; with 1,000, 0.004, 0.008, 0.009, 0.008, 0.004; by the interim trend itself, 0.234, 0.163, 0.089, 0.042, 0.013. Exact integrals over the interim statistic.
Fig. 1 The error rate above 5%, in points, of a two-look O’Brien–Fleming-type trial that adds a look at three quarters with a stated chance under the null, by what the decision to add it reads: the calendar or the within-arm variance, the variance computed with the arm labels hidden at three interim sizes, and the interim trend itself at the same chance. Exact integrals over the interim statistic.

The variance computed with the labels hidden

Take a two-arm trial with a normal outcome, n1n_1 patients at the interim, half in each arm. The committee wants to know whether the outcome is noisier than the protocol assumed, and it is not supposed to see the arms, so it computes the variance of all n1n_1 outcomes with the labels hidden. This is what a blinded sample-size review computes, and it is called blinded because nobody who computes it can tell which arm is ahead.

Its sum of squares splits exactly into two parts. The within-arm part is the sum of squared deviations from each arm’s own mean, which is σ2\sigma^2 times a chi-square on n1−2n_1 - 2 degrees of freedom, independent of the arm means. The between-arm part is the squared difference of the arm means times n1/4n_1/4, and that is σ2\sigma^2 times the square of the interim statistic. So

(n1−1) Sblind2σ2=χn1−22+z12,\frac{(n_1 - 1)\,S^2_{\text{blind}}}{\sigma^2} = \chi^2_{n_1-2} + z_1^2,

exactly, under the null and under any effect. The variance nobody can read the effect from contains the interim statistic, squared. It cannot tell the committee which arm is ahead; it can tell it, a little, how far apart the arms are. Choosing n after looking met the same identity from the other side, as the overshoot a blinded sample-size re-estimate pays because its variance includes the effect.

The within-arm variance needs the labels — it is computed arm by arm — and contains only the chi-square. A committee looking at it sees the labels and is, in that sense, unblinded. It is the one that carries none of the effect.

Three reasons, three dependences

The three administrative reasons a committee gives for adding a look differ in exactly this respect. Recruitment ahead of plan is a fact about the calendar: the look is added because patients arrived faster, and nothing about their outcomes enters the decision. A spending function was built for this case. Spending the error rate is a promise that any schedule fixed without reference to the data — including one fixed by the pace of enrolment — spends the stated total, because the boundaries at each look are computed from the information that has arrived and the share of the error allotted to it.

A variance larger than assumed is a fact about the outcomes, and which variance is meant decides whether it is also a fact about the effect. The within-arm variance is independent of the arm means for a normal outcome, by the same theorem that makes a tt statistic’s numerator and denominator independent, and a look triggered on it is as safe as a look triggered by the calendar. The variance with the labels hidden is not, by the identity above.

An event rate higher than expected, for a binary outcome, is the third and is not measured here; the closing section says why it is a different calculation. What the first two establish is that “nuisance” is not a property of the quantity’s name. It is a property of the quantity’s joint distribution with the interim statistic, and two quantities both called the variance can sit on opposite sides of the line.

A trigger that reads ∣z1∣|z_1| without meaning to

A rule that adds a look when the blinded variance exceeds a threshold therefore adds it with a probability that depends on the interim statistic. With the threshold set so that the look is added half the time under the null, that probability is, given z1z_1,

P(add∣z1)=P(χn1−22>c−z12),P(\text{add} \mid z_1) = P\left(\chi^2_{n_1-2} > c - z_1^2\right),

which is lowest at z1=0z_1 = 0 and rises symmetrically as the trend strengthens in either direction.

The chance a look is added, given the interim statistic, when the trigger is the blinded variance. With the threshold set so the look is added half the time under the null: with 20 patients at the interim the chance is 43.4% at an interim statistic of zero and 70.7% at plus or minus two; with 100, 47.2% and 58.6%; with 1,000, 49.1% and 52.7%. A trigger on the within-arm variance adds it with chance one half at every interim value.
Fig. 2 The chance a look is added, given the interim statistic, when the trigger is a blinded variance above a threshold set to add it half the time under the null — with 20, 100 and 1,000 patients at the interim — beside a within-arm variance trigger, which adds it half the time at every interim value.

With twenty patients at the interim the look is added in 43.4% of trials whose interim statistic is zero and 70.7% of those at ±2\pm2. With a hundred, 47.2% and 58.6%. With a thousand, 49.1% and 52.7%. The trend’s share of the blinded variance is one degree of freedom in n1−1n_1 - 1, so a larger interim dilutes it — but it never removes it, and the trigger is always more likely to fire on the trials whose trend is strong.

That is the same shape as the rule that adds a look when ∣z1∣≥1.5|z_1| \ge 1.5, softened. The trend rule is a step from zero to one; the blinded-variance rule is a gentle slope. Both send the trials with strong trends to the schedule with the extra look more often than the trials with weak ones.

What the slope spends

Whether the slope costs anything depends on what the extra look does to a trial that reaches it. Under an O’Brien–Fleming-type spending function the boundaries of each schedule are set so the schedule spends 5% if it is chosen without reference to the data. A trial whose interim trend is strong is more likely to cross with an extra look than without; one whose trend is weak is less likely to. A trigger that sends more of the strong trials to the extra look spends more than 5%, by an amount that is an exact integral over the interim statistic.

The hero figure sets out that integral for every trigger. A calendar trigger and a within-arm variance trigger spend 5% exactly at every marginal chance of adding the look, to the integration’s accuracy: both add it with a probability that does not depend on z1z_1, so the trial is a fixed mixture of two schedules that each spend 5%. The blinded variance spends more. With twenty patients at the interim it adds 0.034 points of error at a marginal chance of one in ten, 0.064 at three in ten, 0.069 at one half, 0.054 at seven in ten and 0.022 at nine in ten; with a hundred patients, 0.012 to 0.030; with a thousand, 0.004 to 0.009. The excess is largest when the look is added about half the time, because then the trigger’s slope across z1z_1 is steepest where the trials are.

Beside the trend trigger at the same marginal chance, the blinded variance is a small residue. Adding the look one time in ten on the strongest interim trends spends 0.234 points above 5%; adding it on the blinded variance with a hundred patients at the interim spends 0.014, about a seventeenth as much. At one half the two spend 0.089 and 0.030. The nuisance trigger carries a fraction of the trend trigger’s damage, and how large a fraction depends on how many patients there are to dilute the trend.

The shape of the excess across the marginal chance is worth reading too. It is zero at both ends, because a trigger that never fires or always fires is a fixed schedule and spends 5%. It peaks near one half, and it is lopsided — 0.054 points at seven in ten against 0.034 at one in ten, with twenty patients at the interim — because the threshold that gives a high marginal chance sits below the chi-square’s centre, where its density, and so its sensitivity to the extra z12z_1^2, is greater than at the threshold that gives a low one.

How the residue shrinks

The dilution can be read directly. At the marginal chance where the residue is largest, it is 0.069 points with twenty patients at the interim, 0.042 with fifty, 0.030 with a hundred, 0.021 with two hundred, 0.015 with four hundred and 0.009 with a thousand.

The blinded variance's excess error against the size of the interim, at the marginal chance where it is largest. Adding the look with chance one half when the blinded variance is high spends 0.069, 0.042, 0.030, 0.021, 0.015, 0.009 points above 5% with 20, 50, 100, 200, 400, 1,000 patients at the interim. On logarithmic axes the points fall with slope −0.51: the residue shrinks as one over the square root of the interim size, because the interim statistic makes up one degree of freedom of all those the blinded variance is summed over, and the trigger's sensitivity to it is the chi-square's density at the threshold.
Fig. 3 The blinded variance’s excess error at a marginal chance of one half, against the number of patients at the interim, on logarithmic axes, with a dashed line falling as one over the square root. Exact integrals.

On logarithmic axes the six points fall on a line of slope −0.51. The residue shrinks as one over the square root of the interim size, not as one over the size, and the reason is in the formula for the trigger. The trend’s contribution to the blinded variance is one term among n1−1n_1 - 1, but the trigger turns on whether the total crosses a threshold, and how much one extra unit of z12z_1^2 moves that probability is the chi-square’s density at the threshold — about 1/4πn11/\sqrt{4\pi n_1} near the middle of the distribution. The trend’s influence on the decision falls with the square root of the interim, and so does the error it spends.

That decay is slow. A trial with four hundred patients at its interim — a large trial by most standards — still spends 0.015 points above 5% if its committee adds a look half the time on the blinded variance. The residue is negligible for any practical purpose at that size; it is not zero at any size, and a committee asked whether its blinded look needs a correction can answer with the exact number rather than with a principle.

For a reader of a trial report the practical meaning is a bound. When the looking happens showed that five looks at the nominal level spend 14%; an O’Brien–Fleming-type function reduces that to 5% by construction; a look added on the trend gives back a quarter of a point; and a look added on the blinded variance gives back between a hundredth and a fifteenth of a point, depending on the trial’s size. Of the ways a committee can spend more than it was allowed, this one is the smallest that is not zero.

When the effect is real

Under the null the trigger’s slope costs a little error. Under the effect the trial is powered for, it does something a committee might care about more: it makes the look’s addition a statement about the effect.

How often each trigger adds the look when the trial's planned effect is real, set to add it half the time under the null. Under the effect the trial is powered for, with the trigger set to add the look in half of null trials among those still running: blinded variance, 20 patients, 67.1%; blinded variance, 50 patients, 60.5%; blinded variance, 100 patients, 57.4%; blinded variance, 200 patients, 55.2%; blinded variance, 400 patients, 53.6%; blinded variance, 1,000 patients, 52.3%; within-arm variance or calendar, 50.0%; the interim trend, 92.1%. The within-arm variance and the calendar add it 50.0% of the time whatever the effect, so they carry none of it.
Fig. 4 How often each trigger adds the look, among trials still running at the interim, when the effect the trial is powered for is real — with each trigger set to add it half the time under the null.

The within-arm variance and the calendar add the look in 50.0% of running trials whether the effect is there or not; they carry none of it. The blinded variance adds it in 67.1% of running trials with twenty patients at the interim, 57.4% with a hundred and 52.3% with a thousand. The trend trigger, which reads the effect on purpose, adds it in 92.1%. A committee that sees the look being added more often than planned under a blinded-variance rule is seeing, in part, the effect — exactly the information a blinded procedure is meant to withhold from it.

None of this touches power in a way worth measuring. With the look added half the time the trial’s power is 88.695% under the within-arm trigger and 88.699% under the blinded one with twenty patients at the interim, against 89.01% with no added look at all. An extra look under an O’Brien–Fleming-type function costs a little power whatever triggers it; what triggers it moves the error rate, not the power.

What blinded means, and what exactness needs

The word blinded is doing two different jobs here, and the arithmetic separates them. In the sense a protocol uses it, a quantity is blinded when the person computing it cannot tell which arm is which. In the sense the spending function’s guarantee needs, a quantity is safe when it is independent of the interim statistic. The variance with the labels hidden is blinded in the first sense and not the second; the within-arm variance is safe in the second sense and not blinded in the first.

The rule that cannot see the mean found the same distinction in a fixed-width trial: a stopping rule built from within-arm contrasts cannot see the arm means at all, and that — not hiding the labels — is what made its coverage exact. A schedule that reads the mean found what breaks when a schedule is a function of anything else. A committee choosing its looks from a variance is in the same position. What protects the error rate is that the quantity be a function of within-arm deviations, and a lumped variance is not.

The practical arrangement follows from that. In most monitored trials the independent committee already receives the outcomes by arm in its closed report, so the within-arm variance is available to it at no cost to anyone’s blinding; it is the sponsor’s open report, written to keep the arms hidden, that carries the lumped figure. A decision to add a look belongs with whoever can compute the safe quantity, and the safe quantity is the one that needed the labels.

What a protocol can say

Trigger an added look on the calendar or on the within-arm variance. Both spend the spending function’s 5% exactly, at every marginal chance of adding the look. The within-arm variance needs the arm labels, which an independent monitoring committee holds in any case; the variance it reports to a sponsor can still be pooled across arms.

If the trigger must be the variance with the labels hidden, the cost is known. It is 0.069 points at a twenty-patient interim, 0.030 at a hundred and 0.009 at a thousand, at its worst, and it can be removed by preserving the conditional error at each interim value, as a look added by a rule did for the trend trigger, where the correction cost 0.38 points of power. The same correction for the blinded trigger has not been computed here; the residue it would remove is a tenth of the trend trigger’s or less.

Do not describe a blinded trigger as independent of the effect. It is less dependent than a trigger on the trend by a factor of ten or more, and under the planned effect it fires 67.1% of the time where the null would fire it half the time, at twenty patients.

Every rate here is an exact integral over the interim partial sum on a grid of 801 points, with the chance of crossing later under each schedule computed by backward recursion over the later looks’ continuation regions; the blinded trigger’s chance given z1z_1 is the chi-square tail in closed form, its threshold the χn1−12\chi^2_{n_1-1} point that gives the stated marginal chance. The calendar and within-arm triggers are checked to spend 5% to six decimals at three marginal chances, and the blinded trigger’s residue is checked to be positive and to shrink as the interim grows. A blinded trigger read as exact is refused: at twenty patients it spends 5.069%.

Still open: the event rate, and outcomes that are not normal

The identity that puts z12z_1^2 inside the blinded variance is exact for a normal outcome. Two of the commonest nuisance triggers are not of that kind. A pooled event rate across both arms, for a binary outcome, is uncorrelated with the difference in rates under the null but not independent of it — a pooled rate near one half leaves room for a large difference and a rate near zero does not — and the statistic’s own variance depends on it. And a skewed continuous outcome breaks the independence of the within-arm variance from the arm means, since a sample’s variance and its mean are correlated whenever the outcome’s third moment is not zero.

Whether a look added on the pooled event rate spends more or less than one added on the blinded variance of a normal outcome of the same size, and whether the within-arm variance of a lognormal outcome is still safe enough to trigger on, are both exact sums of the kind used here — the first over the two binomial counts, the second over the interim distribution of a skewed sample’s variance and mean — and neither has been computed.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Alpha spendingBlinded analysisGroup sequentialInterim analysisNuisance parameterStatistical powerStopping ruleType i error