The variance that contains the effect
Worth reading first: When the looking happens · A width promised for a difference.
A look the trend asked for found that a committee adding a look at three quarters of a trial whenever the interim is 1.5 or more spends 5.2323% under a spending function built to spend 5%, and a look added by a rule found three ways to give such a look boundaries that spend 5% exactly. Every trigger in those essays read the effect. The look was added because the trend was close.
Committees more often add a look for reasons that are not the effect. Recruitment ran faster than planned, so the information will arrive sooner. The event rate is higher than assumed. The outcome’s variance came in larger than the sample size calculation used, so the trial is less informative than it was meant to be at this point and a later look would be worth more. Each of these is a nuisance quantity, and the standing advice is that an adaptation reading only a nuisance quantity leaves the error rate alone. The question is whether a nuisance quantity can carry any of the effect. For one of the most common of them it can, and the reason is an identity.
The variance computed with the labels hidden
Take a two-arm trial with a normal outcome, patients at the interim, half in each arm. The committee wants to know whether the outcome is noisier than the protocol assumed, and it is not supposed to see the arms, so it computes the variance of all outcomes with the labels hidden. This is what a blinded sample-size review computes, and it is called blinded because nobody who computes it can tell which arm is ahead.
Its sum of squares splits exactly into two parts. The within-arm part is the sum of squared deviations from each arm’s own mean, which is times a chi-square on degrees of freedom, independent of the arm means. The between-arm part is the squared difference of the arm means times , and that is times the square of the interim statistic. So
exactly, under the null and under any effect. The variance nobody can read the effect from contains the interim statistic, squared. It cannot tell the committee which arm is ahead; it can tell it, a little, how far apart the arms are. Choosing n after looking met the same identity from the other side, as the overshoot a blinded sample-size re-estimate pays because its variance includes the effect.
The within-arm variance needs the labels — it is computed arm by arm — and contains only the chi-square. A committee looking at it sees the labels and is, in that sense, unblinded. It is the one that carries none of the effect.
Three reasons, three dependences
The three administrative reasons a committee gives for adding a look differ in exactly this respect. Recruitment ahead of plan is a fact about the calendar: the look is added because patients arrived faster, and nothing about their outcomes enters the decision. A spending function was built for this case. Spending the error rate is a promise that any schedule fixed without reference to the data — including one fixed by the pace of enrolment — spends the stated total, because the boundaries at each look are computed from the information that has arrived and the share of the error allotted to it.
A variance larger than assumed is a fact about the outcomes, and which variance is meant decides whether it is also a fact about the effect. The within-arm variance is independent of the arm means for a normal outcome, by the same theorem that makes a statistic’s numerator and denominator independent, and a look triggered on it is as safe as a look triggered by the calendar. The variance with the labels hidden is not, by the identity above.
An event rate higher than expected, for a binary outcome, is the third and is not measured here; the closing section says why it is a different calculation. What the first two establish is that “nuisance” is not a property of the quantity’s name. It is a property of the quantity’s joint distribution with the interim statistic, and two quantities both called the variance can sit on opposite sides of the line.
A trigger that reads without meaning to
A rule that adds a look when the blinded variance exceeds a threshold therefore adds it with a probability that depends on the interim statistic. With the threshold set so that the look is added half the time under the null, that probability is, given ,
which is lowest at and rises symmetrically as the trend strengthens in either direction.
With twenty patients at the interim the look is added in 43.4% of trials whose interim statistic is zero and 70.7% of those at . With a hundred, 47.2% and 58.6%. With a thousand, 49.1% and 52.7%. The trend’s share of the blinded variance is one degree of freedom in , so a larger interim dilutes it — but it never removes it, and the trigger is always more likely to fire on the trials whose trend is strong.
That is the same shape as the rule that adds a look when , softened. The trend rule is a step from zero to one; the blinded-variance rule is a gentle slope. Both send the trials with strong trends to the schedule with the extra look more often than the trials with weak ones.
What the slope spends
Whether the slope costs anything depends on what the extra look does to a trial that reaches it. Under an O’Brien–Fleming-type spending function the boundaries of each schedule are set so the schedule spends 5% if it is chosen without reference to the data. A trial whose interim trend is strong is more likely to cross with an extra look than without; one whose trend is weak is less likely to. A trigger that sends more of the strong trials to the extra look spends more than 5%, by an amount that is an exact integral over the interim statistic.
The hero figure sets out that integral for every trigger. A calendar trigger and a within-arm variance trigger spend 5% exactly at every marginal chance of adding the look, to the integration’s accuracy: both add it with a probability that does not depend on , so the trial is a fixed mixture of two schedules that each spend 5%. The blinded variance spends more. With twenty patients at the interim it adds 0.034 points of error at a marginal chance of one in ten, 0.064 at three in ten, 0.069 at one half, 0.054 at seven in ten and 0.022 at nine in ten; with a hundred patients, 0.012 to 0.030; with a thousand, 0.004 to 0.009. The excess is largest when the look is added about half the time, because then the trigger’s slope across is steepest where the trials are.
Beside the trend trigger at the same marginal chance, the blinded variance is a small residue. Adding the look one time in ten on the strongest interim trends spends 0.234 points above 5%; adding it on the blinded variance with a hundred patients at the interim spends 0.014, about a seventeenth as much. At one half the two spend 0.089 and 0.030. The nuisance trigger carries a fraction of the trend trigger’s damage, and how large a fraction depends on how many patients there are to dilute the trend.
The shape of the excess across the marginal chance is worth reading too. It is zero at both ends, because a trigger that never fires or always fires is a fixed schedule and spends 5%. It peaks near one half, and it is lopsided — 0.054 points at seven in ten against 0.034 at one in ten, with twenty patients at the interim — because the threshold that gives a high marginal chance sits below the chi-square’s centre, where its density, and so its sensitivity to the extra , is greater than at the threshold that gives a low one.
How the residue shrinks
The dilution can be read directly. At the marginal chance where the residue is largest, it is 0.069 points with twenty patients at the interim, 0.042 with fifty, 0.030 with a hundred, 0.021 with two hundred, 0.015 with four hundred and 0.009 with a thousand.
On logarithmic axes the six points fall on a line of slope −0.51. The residue shrinks as one over the square root of the interim size, not as one over the size, and the reason is in the formula for the trigger. The trend’s contribution to the blinded variance is one term among , but the trigger turns on whether the total crosses a threshold, and how much one extra unit of moves that probability is the chi-square’s density at the threshold — about near the middle of the distribution. The trend’s influence on the decision falls with the square root of the interim, and so does the error it spends.
That decay is slow. A trial with four hundred patients at its interim — a large trial by most standards — still spends 0.015 points above 5% if its committee adds a look half the time on the blinded variance. The residue is negligible for any practical purpose at that size; it is not zero at any size, and a committee asked whether its blinded look needs a correction can answer with the exact number rather than with a principle.
For a reader of a trial report the practical meaning is a bound. When the looking happens showed that five looks at the nominal level spend 14%; an O’Brien–Fleming-type function reduces that to 5% by construction; a look added on the trend gives back a quarter of a point; and a look added on the blinded variance gives back between a hundredth and a fifteenth of a point, depending on the trial’s size. Of the ways a committee can spend more than it was allowed, this one is the smallest that is not zero.
When the effect is real
Under the null the trigger’s slope costs a little error. Under the effect the trial is powered for, it does something a committee might care about more: it makes the look’s addition a statement about the effect.
The within-arm variance and the calendar add the look in 50.0% of running trials whether the effect is there or not; they carry none of it. The blinded variance adds it in 67.1% of running trials with twenty patients at the interim, 57.4% with a hundred and 52.3% with a thousand. The trend trigger, which reads the effect on purpose, adds it in 92.1%. A committee that sees the look being added more often than planned under a blinded-variance rule is seeing, in part, the effect — exactly the information a blinded procedure is meant to withhold from it.
None of this touches power in a way worth measuring. With the look added half the time the trial’s power is 88.695% under the within-arm trigger and 88.699% under the blinded one with twenty patients at the interim, against 89.01% with no added look at all. An extra look under an O’Brien–Fleming-type function costs a little power whatever triggers it; what triggers it moves the error rate, not the power.
What blinded means, and what exactness needs
The word blinded is doing two different jobs here, and the arithmetic separates them. In the sense a protocol uses it, a quantity is blinded when the person computing it cannot tell which arm is which. In the sense the spending function’s guarantee needs, a quantity is safe when it is independent of the interim statistic. The variance with the labels hidden is blinded in the first sense and not the second; the within-arm variance is safe in the second sense and not blinded in the first.
The rule that cannot see the mean found the same distinction in a fixed-width trial: a stopping rule built from within-arm contrasts cannot see the arm means at all, and that — not hiding the labels — is what made its coverage exact. A schedule that reads the mean found what breaks when a schedule is a function of anything else. A committee choosing its looks from a variance is in the same position. What protects the error rate is that the quantity be a function of within-arm deviations, and a lumped variance is not.
The practical arrangement follows from that. In most monitored trials the independent committee already receives the outcomes by arm in its closed report, so the within-arm variance is available to it at no cost to anyone’s blinding; it is the sponsor’s open report, written to keep the arms hidden, that carries the lumped figure. A decision to add a look belongs with whoever can compute the safe quantity, and the safe quantity is the one that needed the labels.
What a protocol can say
Trigger an added look on the calendar or on the within-arm variance. Both spend the spending function’s 5% exactly, at every marginal chance of adding the look. The within-arm variance needs the arm labels, which an independent monitoring committee holds in any case; the variance it reports to a sponsor can still be pooled across arms.
If the trigger must be the variance with the labels hidden, the cost is known. It is 0.069 points at a twenty-patient interim, 0.030 at a hundred and 0.009 at a thousand, at its worst, and it can be removed by preserving the conditional error at each interim value, as a look added by a rule did for the trend trigger, where the correction cost 0.38 points of power. The same correction for the blinded trigger has not been computed here; the residue it would remove is a tenth of the trend trigger’s or less.
Do not describe a blinded trigger as independent of the effect. It is less dependent than a trigger on the trend by a factor of ten or more, and under the planned effect it fires 67.1% of the time where the null would fire it half the time, at twenty patients.
Every rate here is an exact integral over the interim partial sum on a grid of 801 points, with the chance of crossing later under each schedule computed by backward recursion over the later looks’ continuation regions; the blinded trigger’s chance given is the chi-square tail in closed form, its threshold the point that gives the stated marginal chance. The calendar and within-arm triggers are checked to spend 5% to six decimals at three marginal chances, and the blinded trigger’s residue is checked to be positive and to shrink as the interim grows. A blinded trigger read as exact is refused: at twenty patients it spends 5.069%.
Still open: the event rate, and outcomes that are not normal
The identity that puts inside the blinded variance is exact for a normal outcome. Two of the commonest nuisance triggers are not of that kind. A pooled event rate across both arms, for a binary outcome, is uncorrelated with the difference in rates under the null but not independent of it — a pooled rate near one half leaves room for a large difference and a rate near zero does not — and the statistic’s own variance depends on it. And a skewed continuous outcome breaks the independence of the within-arm variance from the arm means, since a sample’s variance and its mean are correlated whenever the outcome’s third moment is not zero.
Whether a look added on the pooled event rate spends more or less than one added on the blinded variance of a normal outcome of the same size, and whether the within-arm variance of a lognormal outcome is still safe enough to trigger on, are both exact sums of the kind used here — the first over the two binomial counts, the second over the interim distribution of a skewed sample’s variance and mean — and neither has been computed.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A boundary for giving up — both name interim analysis, statistical power, stopping rule
- A simulation that stops when it looks settled — both name interim analysis, statistical power, stopping rule
- The effect a stopped trial reports — both name interim analysis, statistical power, stopping rule
- A block size that changes — both name nuisance parameter, stopping rule
- A statistic that is exact twice — both name nuisance parameter, statistical power
- The charge that is not a sum — both name statistical power, type i error
Named objects
A flat tag is an object no other essay names yet.
Alpha spendingBlinded analysisGroup sequentialInterim analysisNuisance parameterStatistical powerStopping ruleType i error