Designs that change while they run

Choosing n after looking

Re-estimating the sample size from an interim is the one adaptation with a defence, and the defence is exactly what it costs: an analyst kept blind to the arms measures a spread that contains the effect, so the design overshoots by 1 + Δ²/4σ². Re-estimating the effect instead breaks the error rate.

Worth reading first: When the looking happens · What a p-value does not say.

Every sample-size calculation needs a standard deviation nobody has. The formula gives n per arm for a stated power at a stated effect, and σ appears in it squared, so being out by 30% means being out by 70% in the number of units. The guess is usually taken from a previous study, an assumption, or nothing at all.

The obvious repair is to look. Run part of the trial, estimate σ from what has arrived, recompute n, and carry on. It is called an internal pilot, it is the one adaptation with a serious defence, and the defence turns out to be exactly what it costs.

What the interim sees, at an effect of 1The same 60 observations, estimated two ways. Keeping the arms separate gives 0.995, which is σ. Pooling them without separating the arms — the price of staying blind to the comparison — gives 1.114, against the identity √(1 + Δ²/4σ²) = 1.118. The sample size is proportional to the variance, so a blinded design at this effect asks for 25% more units than it needs, and it does so systematically rather than by chance.05001e+30.50011.502the interim estimate of σtrials out of 6,000σ = 1√(1 + Δ²/4σ²) = 1.118unblindedblinded6,000 interims of 30 per arm, effect 1blinded 1.114 against 1.118
Fig. 1 The same sixty interim observations, estimating σ two ways. One of the two distributions is centred on the truth and the other is centred on √(1 + Δ²/4σ²) times it.

The defence, and what it is a defence against

A stopping rule changes what a p-value means because it changes the set of outcomes the experiment could have produced. That is the sequential field’s whole content, and it applies to any decision made after looking at the data.

The internal pilot’s claim is that it does not look at the comparison. It looks at a nuisance parameter — the spread — and re-estimates a design quantity from it, leaving the hypothesis being tested and the analysis performed on it untouched. If that separation is real, then re-estimating n is not peeking in the sense the sequential field means.

Keeping it real is what blinding is for. An analyst who never learns which arm is which cannot use the comparison even accidentally, and the design is safe by construction rather than by discipline.

It is worth being precise about what “not looking at the comparison” has to mean, because the phrase is doing real work and it is easy to satisfy in letter only. It is not enough that no formal test was performed at the interim. The requirement is that the rule mapping interim data to the new sample size does not depend on the difference between the arms — so a rule that pools the arms qualifies, a rule that separates them and uses only their spreads very nearly qualifies, and a rule that looks at the gap does not, however informally it looks.

That distinction is the whole of the table below, and the three rules in it differ by exactly which function of the interim they are allowed to see.

What blinding costs, exactly

Pool two groups of equal size that differ by Δ and compute their spread about the grand mean. Each observation’s deviation now has two parts: its own noise, and half the difference between the arms. The expectation is

E[s²] = σ² + Δ²/4

which is an identity, not an approximation. The sample size is proportional to the variance estimate, so a blinded design asks for 1 + Δ²/4σ² times the units it needs.

Measured across twelve thousand interims of thirty per arm:

true effect blinded σ̂ √(1 + Δ²/4σ²) unblinded σ̂
0.5 1.026 1.031 1.000
1.0 1.115 1.118 1.000
2.0 1.416 1.414 0.995

At an effect of one standard deviation the blinded design asks for 25% more units than it needs. At two, for twice as many. The counted estimate and the identity agree to three decimal places at every effect, which is what an identity should give — and is worth checking rather than assuming, because the whole argument of this essay rests on the displacement being a known quantity rather than an approximation with a regime of validity.

The unblinded estimate is unaffected by the effect at any size, which is what makes it the right estimate of the nuisance parameter: it separates the arms, computes each group’s spread about its own mean, and the difference between the arms cancels.

What the interim sees, at an effect of 2. The same 60 observations, estimated two ways. Keeping the arms separate gives 0.995, which is σ. Pooling them without separating the arms — the price of staying blind to the comparison — gives 1.414, against the identity √(1 + Δ²/4σ²) = 1.414. The sample size is proportional to the variance, so a blinded design at this effect asks for 100% more units than it needs, and it does so systematically rather than by chance.
Fig. 2 At an effect of two standard deviations the two distributions barely overlap. The blinded estimate is not noisy — it is centred somewhere else, and the somewhere else has a closed form.

And what it buys

The error rate. Both rules were run twenty thousand times with no effect at all, and tested at a nominal 5%:

the interim looks at size mean n per arm
nothing — fixed design 4.98% 85.0
the pooled spread — blinded 5.01% 84.6
each arm’s spread — unblinded 5.22% 84.5
the spread and the observed effect 7.34% 835.5

The first three hold their level. The blinded rule holds it exactly; the unblinded rule is a fifth of a point high, which is at the edge of what twenty thousand trials can resolve and is the known small inflation of an internal pilot that separates the arms.

Notice also that at a true effect of zero the blinded rule costs nothing — Δ = 0, so the inflation factor is 1. Its penalty is a function of the effect, which means it is zero exactly where the design’s error rate is being measured and largest exactly where the trial is most likely to succeed anyway.

Four interims, and which of them looked at the comparison. 8,000 trials with no effect at all. Three of these rules change the sample size after looking at the data and hold the error rate at 4.63%, 4.81% and 4.91%; the fourth reaches 7.20%. The difference is not how much the design changed but what the interim was allowed to see: the first three re-estimate a nuisance parameter, and the fourth re-estimates the effect being tested.
Fig. 3 Four interims, and which of them looked at the comparison. Three hold the rate they claim; the fourth does not, and the difference is not how much the design changed.

The fourth rule, and why it is a different animal

The last row is the one that is not a nuisance-parameter adaptation at all. It re-estimates σ and the effect, and recomputes n to power the study for the effect it has just seen — which sounds like a sensible response to an interim showing something smaller than planned, and is the version most often reached for when a trial is in trouble.

It rejects a true null 7.34% of the time, and the mean sample size goes to 835 per arm against a planned 85, because a small observed effect demands an enormous n.

The mechanism is the sequential field’s: the sample size is now a function of the comparison, so the set of outcomes that produce a rejection is different from the set the critical value was computed for. Trials that happen to show a large effect early stop small and are tested at the nominal level; trials that show nothing are enlarged until they might. That is spending the error rate without budgeting for it, and the budget is what the corrected boundaries in that field exist to keep.

The repair is known and it is not this essay’s: a design that re-estimates on the effect must combine the two stages with pre-specified weights rather than pooling them, which restores the level and is the machinery the next essays are about.

What each rule spends, in units per arm. The same four rules, and what they ask for. A design that knew σ would use 85 per arm. The blinded rule asks for 84 because the spread it measures contains the effect; the unblinded rule asks for 84; and re-powering for the observed effect asks for 838, which is a different number every time and is the point rather than a side effect.
Fig. 4 What each rule spends. The blinded rule’s overshoot is a fact about the effect, and the effect-adaptive rule’s is a fact about how small the observed effect happened to be.

Where to put the interim

The rule says re-estimate; it does not say when. That is a design decision with the same shape as the pilot size in the allocation field — both ends are bad — and the two ends are bad for different reasons here.

Blinded re-estimation at a true effect of half a standard deviation, planned at 85 per arm:

interim at mean n the middle 90% of n power
10 per arm 90.2 48 to 145 86.5%
20 89.8 59 to 126 89.1%
40 89.7 68 to 114 89.4%
60 89.9 72 to 110 90.0%

The average sample size barely moves. What moves is its spread: an interim at ten per arm produces trials anywhere from 48 to 145 per arm, for a design that was planned at 85, and that variability costs three and a half points of power against an interim at sixty.

The mechanism is worth stating because it is not obvious that variability should cost anything when the average is right. Power is a concave function of n — the first extra units buy far more than the last — so a design whose n is scattered around the right value has less power than one that lands on it, by Jensen’s inequality. The trials that came out small lose more than the trials that came out large gain.

Pushing the interim later fixes that and introduces the constraint that ends the trade: a late interim cannot shrink a trial below what has already been spent. At sixty per arm, a design that discovers it needed fifty has spent sixty, and the re-estimation is one-sided by then. So the range of adjustments available narrows exactly as the estimate driving them improves, and the useful window is somewhere in the middle — a quarter to a half of the planned size is the range these numbers support.

What it does to power

The size table above is the safety question. The power table is the one a trial is actually run for, and it says something the size table does not.

At an effect of half a standard deviation the fixed design has 89.9% power at its planned 85 per arm. The blinded re-estimation, which spends 90 on average, has 89.1%. It spends 6% more units and has less power.

That is the concavity again, and it is the honest accounting for what an internal pilot costs when the original guess at σ happened to be right. The design pays for the insurance in two currencies: units, because blinding inflates the estimate, and power, because a variable n is worth less than a fixed n of the same average.

What it buys is protection against the case the table cannot show — the original σ being wrong. A fixed design planned at σ = 1 when the truth is 1.3 has 85 units where it needed 143, and its power is not 89.9% but far below. The re-estimating design finds that and fixes most of it. The comparison above is the premium, measured on the trials where the insurance was not needed.

What the interim sees, at an effect of 0.5. The same 60 observations, estimated two ways. Keeping the arms separate gives 0.995, which is σ. Pooling them without separating the arms — the price of staying blind to the comparison — gives 1.025, against the identity √(1 + Δ²/4σ²) = 1.031. The sample size is proportional to the variance, so a blinded design at this effect asks for 6% more units than it needs, and it does so systematically rather than by chance.
Fig. 5 The interim’s two estimates at the effect a trial is typically powered for. The blinded distribution is displaced by 3%, which is a 6% overshoot in units, and this is the case the rule is usually applied in.

Five patients an arm, whatever the effect

The blinding premium is quoted above as a percentage — 6% at half a standard deviation, 25% at one, 100% at two — and read that way it looks like a penalty that grows alarmingly with the effect. In the currency a trial is recruited in it does not grow at all.

A trial powered at 1 − β for a two-sided test at α needs about n=2σ2(zα/2+zβ)2/Δ2n = 2\sigma^2(z_{\alpha/2} + z_\beta)^2/\Delta^2 per arm. The blinding inflates that by Δ²/4σ², so the extra units are

nΔ2/4σ2  =  (zα/2+zβ)22n\,\Delta^2/4\sigma^2 \;=\; \frac{(z_{\alpha/2} + z_\beta)^2}{2}

in which neither Δ nor σ appears. At 90% power and 5% two-sided that is (1.96 + 1.282)²/2 = 5.25 units per arm, and it is 5.25 units per arm at every effect size and every spread.

The table confirms it three times over. At Δ = 0.5 the planned trial is 85 per arm and 6.25% of it is 5.3. At Δ = 1 the planned trial is a quarter of that, about 21 per arm, and 25% of it is 5.3. At Δ = 2 the planned trial is about 5 per arm and 100% of it is 5.3. One number, wearing three percentages.

That reframes the choice between blinded and unblinded re-estimation entirely. The question is not whether a trial can afford a percentage that might be 6% or might be 100%; it is whether it can afford about five more subjects per arm — five at 90% power, 3.9 at 80%, 7.2 at 95% — and the answer is almost always yes. The percentage is large exactly where the trial is small, which is where five extra units are cheapest to find.

It also says where the blinded rule genuinely is expensive: nowhere in the effect, and everywhere in the power. The premium goes as (zα/2+zβ)2(z_{\alpha/2} + z_\beta)^2, so demanding 99% power rather than 90% takes it from 5.25 to 9.9 per arm. A design that is already buying certainty is the one that pays most for blinding, which is the opposite of the reading the percentage column invites.

The spread falls as the root of the interim

The timing table’s useful column is the middle 90% of the re-estimated sample size — 48 to 145, 59 to 126, 68 to 114, 72 to 110 — and its widths are 97, 67, 46 and 38 at interims of 10, 20, 40 and 60 per arm.

Those ratios are √m and nothing else. 97/67 = 1.45 against √2 = 1.41; 97/46 = 2.11 against √4 = 2.00; 97/38 = 2.55 against √6 = 2.45. The scatter in the trial’s own size falls exactly as the precision of the variance estimate driving it, which is a reassuring finding rather than a surprising one and is worth having because it makes the timing decision arithmetic.

Halving the scatter takes four times the interim. Going from a tenth of the planned trial to a quarter of it cuts the scatter by a third and recovers most of the three and a half points of power the earliest interim gives away; going further recovers half a point more and runs into the one-sidedness that ends the trade. That is what puts the useful window where it is, and it is decided by a square root rather than by judgement.

Blinded or unblinded, then

The trade is now stated in numbers rather than in principle.

Blinded costs 1 + Δ²/4σ² in units and nothing in error rate. At an effect of half a standard deviation — a typical planning assumption — that is 6% more units; at one standard deviation, 25%.

Unblinded costs nothing in units and a fifth of a percentage point in error rate, plus an operational cost that this site’s arithmetic cannot measure: somebody has to see the arms separated partway through a trial, and the reason blinding exists is that people who see interim results behave differently afterwards.

For most trials the blinded rule is right, and the reason is the shape of the penalty rather than its size. The overshoot depends on Δ², so at the effects a trial is usually powered for it is small — 6% at half a standard deviation — and it grows only where the trial was going to succeed comfortably anyway. An unnecessarily large trial that finds a real effect is a much better failure than a correctly sized trial with an error rate nobody can state.

Four interims, and which of them looked at the comparison. 8,000 trials with no effect at all. Three of these rules change the sample size after looking at the data and hold the error rate at 5.00%, 4.85% and 5.09%; the fourth reaches 7.17%. The difference is not how much the design changed but what the interim was allowed to see: the first three re-estimate a nuisance parameter, and the fourth re-estimates the effect being tested.
Fig. 6 The same four rules with the interim twice as late. Everything is stabler and the ordering is unchanged, because the rule that looked at the comparison is not made safe by looking later.

The one-sidedness nobody plans for

There is an asymmetry in how these designs get used that no simulation here can capture, and it is worth naming because it changes what the numbers above mean in practice.

A re-estimation rule is two-sided arithmetic: if σ̂ comes in low, n falls. Trials do not shrink. A protocol amendment that reduces the sample size has to be justified to a committee that has already approved the larger one, the recruitment is often already contracted, and the downside of stopping short is visible while the downside of continuing is not.

So the rule as implemented is usually max(planned, re-estimated), and that is a different rule with different properties. Its average sample size is above both the planned and the re-estimated one; its error rate is unaffected, since nothing about the comparison entered; and its overshoot is now the sum of two things — the blinding penalty, which has a closed form, and the truncation, which does not.

The honest way to design for that is to plan the smaller trial and let the rule enlarge it, rather than planning conservatively and re-estimating on top. Otherwise the conservatism is applied twice and neither application knows about the other.

Testing at 0.05 every time the data is looked at. The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.
Fig. 7 The sequential field’s headline, for comparison: testing at the nominal level at each of five looks rejects a true null far more often than 5%. The internal pilot changes the sample size after looking and does not, because what it looked at was not the comparison.

What this shares with the rest of the site

The blinded rule is a plug-in estimate feeding a design decision, which is the shape the allocation field ends on and the shape the pooling field is built on. What makes this one the mildest of the three is worth naming, because it is the only case on this site where the plug-in habit comes out looking respectable.

The quantity enters linearly. n is proportional to σ², and the estimate of σ² is unbiased, so the average sample size is nearly right — where the allocation rule needed a ratio, whose expectation is not the ratio of the expectations, and the pooled interval needed a spread that appears in the width.

The error is one-sided and known. The blinded overshoot has a closed form and a sign, so a designer who has assumed an effect for the power calculation already knows how much to discount.

And overshooting is not a failure. A design that asks for 25% more units than it needs still tests at the level it claims and has more power than planned. The plug-in failures elsewhere on this site produce intervals that miss and rules that lose; this one produces a trial that is larger than it had to be.

That is the whole reason internal pilots survive as a technique while response-adaptive randomisation — which looks similar, and is the next essay — does not survive contact with a time trend.

The refusal, which is the design that adapts nothing

The check behind this essay runs the same simulation, the same test and the same machinery with the adaptation switched off. That fixed design rejects a true null 4.98% of the time across forty thousand trials and uses 85 units per arm on every single one of them, because nothing about it depends on the data.

Without that row the whole table would be uninterpretable. An inflation measured against a harness that was itself slightly wrong is not an inflation, and this site has recorded that mistake in enough places to write the control first.

What each rule spends, in units per arm. The same four rules, and what they ask for. A design that knew σ would use 85 per arm. The blinded rule asks for 85 because the spread it measures contains the effect; the unblinded rule asks for 85; and re-powering for the observed effect asks for 1103, which is a different number every time and is the point rather than a side effect.
Fig. 8 What each rule spends with a later interim. The three nuisance-parameter rules are close to the design that adapts nothing; the effect-adaptive rule is not, and its average is dominated by the trials that saw a small effect and demanded an enormous n.

The other refusal is the identity. Blinded re-estimation is required to land on √(1 + Δ²/4σ²) at three effect sizes and the unblinded one on σ at the largest of them — so a future change that quietly used within-arm variances in the blinded branch would be caught by the estimate coming out right, which is the only symptom it would have.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designBlindingError rateInternal pilotNuisance parameterOptional stoppingPlug in estimateSample sizeStatistical power