Designs that change while they run

Randomising towards the winner

Allocating more patients to the arm that is doing better is the humane thing to want and it buys nothing statistically: at a fixed total it costs thirty points of power. And because the allocation is a function of the outcomes, the ordinary test on it rejects a true null 7.8% of the time before any time trend is applied — and 58% after one.

Worth reading first: Randomisation is not balance · When the looking happens.

A trial that is going well is a trial in which one arm is better, and every patient assigned to the other one is being given something worse on purpose. Response-adaptive randomisation is the obvious response: update the allocation as the results arrive, so that later patients are more likely to receive whatever is winning.

It works, in the sense that it does what it says. What it costs is measurable in two currencies, and the second one is not the one anybody expects.

20 adaptive trials, 45% against 25%Each line is one trial allocating patients one at a time by the arm's own posterior. The average final share on the better arm is 84.7%, with a standard deviation of 10.3 points across these 20 trials. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.00.2500.5000.7501050100150200patients enrolledshare of them assigned to the better arman even splitmean 85% by the end20 trials of 200 patients, Thompson sampling84.7% ± 10.3 on the better arm
Fig. 1 Twenty adaptive trials of two hundred patients, allocated one at a time by each arm’s own posterior. The average trial ends with 87% of its patients on the better arm, and some of them end with most of theirs on the worse one.

What it delivers

Two arms with success rates of 45% and 25%, two hundred patients, allocated by drawing from each arm’s Beta posterior and assigning to whichever draw is higher.

86.6% of patients reach the better arm on average — 173 of the two hundred, against a hundred under a fair coin. That is the whole point of the method and it delivers it.

The number that is not in the figure is the one worth adding: 173 on the better arm means 27 on the worse one, and 27 is not zero. A design that is being defended on the grounds that it stops giving people the worse treatment still gives it to twenty-seven of two hundred, and would have given it to a hundred under a fair coin. The gain is real and it is 73 patients rather than the whole of the ethical objection.

The spread around that average matters as much as the average, which is why the figure draws twenty trials rather than one. The rule commits early on the basis of a handful of outcomes, and a run of bad luck on the genuinely better arm sends most of a trial to the worse one. That is not a rare failure; it is a tail of the same distribution the 86.6% is the mean of.

20 randomised trials, 45% against 25%. Each line is one trial allocating patients one at a time by a fair coin. The average final share on the better arm is 50.9%, with a standard deviation of 2.9 points across these 20 trials, and 7 of them finished with most of the units on the worse arm. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.
Fig. 2 The same twenty trials under a fair coin, for comparison. Every path converges on a half and stays there, and the arms are balanced across arrival time in every trial rather than on average.

What it costs in power

The same trials, tested for a difference between the arms at the end:

allocation patients on the better arm trials detecting the difference
a fair coin 100 85.0%
adaptive 173 55.9%

Thirty points of power, at the same total. The reason is the allocation field’s first result read backwards: a comparison’s variance is σ₁²/n₁ + σ₂²/n₂, and it is dominated by whichever arm has fewer observations. Putting 173 patients on one arm leaves 27 on the other, and 27 is what the comparison is worth.

So the trade is real and it is not subtle: the design buys better treatment for the patients inside the trial and pays for it with the trial’s ability to establish anything for the patients outside it. A design that ends at 55.9% power is one that will fail to demonstrate a difference it was built on top of, nearly half the time.

20 adaptive trials, 45% against 5%. Each line is one trial allocating patients one at a time by the arm's own posterior. The average final share on the better arm is 95.3%, with a standard deviation of 1.8 points across these 20 trials. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.
Fig. 3 The same rule against a much larger difference — 45% against 5%. The allocation commits harder and faster, which is what the rule is for, and leaves even less on the arm the comparison depends on.

And it breaks the test before any trend arrives

Here is the part that is not a trade.

Run the same design with both arms at 30% — no difference at all — and test it in the ordinary way. The fair coin rejects 4.98% of the time, which is the nominal level and is what randomisation buys. The adaptive rule rejects 7.8%.

There is no time trend here, no drift, nothing changing. The inflation is structural: the number of patients on each arm is a function of the outcomes on each arm, so the counts and the successes are not independent, and the two-proportion test assumes they are. An arm that got lucky early receives more patients, which makes its observed rate a mixture of its lucky start and its subsequent regression — and the test has no way to know that the denominator was chosen by the numerator.

What a time trend does

Now add the thing every real trial has. Patients recruited later differ from patients recruited earlier — healthier, better managed, recruited from different sites — so let the success rate rise over the course of the trial in both arms equally. There is still no difference between the arms, and a correct analysis must find none.

drift over the trial a fair coin adaptive adaptive, blocked by arrival time
0 4.98% 7.8% 7.0%
0.4 4.98% 58.2% 2.55%

The fair coin is untouched. That is what randomisation is for and it is worth stating plainly: a trend that moves both arms alike cannot bias a comparison in which the two arms are balanced across time, and a fair coin balances them across time by construction.

The adaptive rule rejects a true null 58.2% of the time. By the time the response has drifted upward, the allocation is heavily skewed towards whichever arm looked good early — so the early, low-response patients are mostly on one arm and the late, high-response patients are mostly on the other, and the difference between the arms is the drift with a label on it.

A trend that is in both arms, read as an effect in one. Both arms have a success rate of 30% and it rises by the same amount over the course of the trial, so there is no difference anywhere to find. A fair coin finds none: 5.4% at no drift and 5.6% at 0.5, the nominal rate throughout, and that is what randomisation buys. The adaptive rule starts at 8.4% with no drift at all — the allocation is a function of the outcomes, so the two are not independent — and reaches 75% at a drift of 0.5, because the early patients went to whichever arm looked good early and the late ones went to the other. Comparing within 8 blocks of arrival time removes the part the drift caused, bringing 75% down to 1.3%, and leaves the part the adaptation caused: 7.6% at no drift.
Fig. 4 Error rates against the size of a drift that moves both arms alike. Three curves: a fair coin, the adaptive rule, and the adaptive rule analysed within blocks of arrival time.

The repair is the field’s oldest idea

Compare patients only with the patients who arrived around the same time — divide the trial into blocks of arrival order, compute the difference within each block, and combine — and the drift cancels inside every block whatever the allocation did.

It works: 58.2% falls to 2.55%.

That is blocking, which this site introduced as a way of removing variance from a field trial, arriving in a place that looks nothing like a field. The nuisance factor is time rather than soil, the blocks are arrival windows rather than plots, and the arithmetic is identical: a contrast taken within a block cannot contain anything the block shares.

Two things are worth noticing about that repair, and neither is a footnote.

It over-corrects at large drifts. 2.55% against a nominal 5% is conservative, not exact. The blocks have very unequal allocation inside them by then — that is what the adaptive rule did — and the weighted combination handles that cautiously.

And it does not fix the other problem. At zero drift the blocked analysis rejects 7.0% against the unadjusted 7.8%. Blocking removes the part the trend caused and leaves the part the adaptation caused, because the dependence between the allocation and the outcomes is inside each block too.

So the honest summary is that blocking repairs the confounding and not the adaptation, and the adaptation needs a different repair — a randomisation test that conditions on the realised allocation, which is what randomisation actually buys and is exact here for the same reason it is exact there.

A trend that is in both arms, read as an effect in one. Both arms have a success rate of 30% and it rises by the same amount over the course of the trial, so there is no difference anywhere to find. A fair coin finds none: 5.4% at no drift and 5.6% at 0.5, the nominal rate throughout, and that is what randomisation buys. The adaptive rule starts at 8.4% with no drift at all — the allocation is a function of the outcomes, so the two are not independent — and reaches 75% at a drift of 0.5, because the early patients went to whichever arm looked good early and the late ones went to the other. Comparing within 1 blocks of arrival time removes the part the drift caused, bringing 75% down to 74.7%, and leaves the part the adaptation caused: 8.4% at no drift.
Fig. 5 The same measurement with the blocking switched off — one block is no blocking — so the middle curve and the unadjusted one coincide. The check that the blocked analysis is doing something is that it stops doing it when there is nothing to block on.

Where this sits among the site’s other adaptations

Three designs on this site now change something after looking at the data, and setting them in order makes the principle visible.

Re-estimating the sample size looks at a nuisance parameter and holds its error rate exactly. The interim cannot see the comparison.

Randomising towards the winner looks at the comparison itself — that is what “which arm is winning” means — and uses it to decide the allocation. Its error rate is 7.8% before any complication is added.

Dropping the losers also looks at the comparison, and holds its rate anyway, because the boundary is computed for the design that was actually run.

So the dividing line is not whether an adaptation looks at the comparison. It is whether the analysis knows that it did. The internal pilot is safe because the interim never touched the comparison; the arm-dropping design is safe because the critical value was solved for a procedure that selects; and response-adaptive randomisation is unsafe as usually analysed because the test applied at the end is the one that would have been right for a fair coin.

That reading makes the repair obvious in principle and awkward in practice: the analysis must condition on the rule that produced the allocation. It is exactly the randomisation test’s argument — the reference distribution is the set of outcomes the assignment mechanism could have produced — applied to a mechanism far more complicated than a coin.

Why the rule commits so hard

The allocation curves in the figures rise fast and flatten, and the shape is worth explaining because it decides everything above.

Thompson sampling assigns each patient by drawing once from each arm’s posterior and taking the higher draw. Early on the posteriors are wide and the draws overlap, so the allocation is near even. As observations accumulate the posteriors separate, the overlap shrinks, and the probability of assigning to the trailing arm falls towards zero — which is the rule working as designed, since the whole point is to stop giving people the worse treatment.

The consequence is that the trailing arm’s sample size stops growing, so the evidence that it is worse stops accumulating. The rule reaches a conclusion and then defends it by collecting no more data that could overturn it, and it does this whether the conclusion is right or wrong.

That is visible in the twenty paths: the ones that commit to the wrong arm early do not come back. There is no mechanism in the rule that recovers, because recovery requires observations from the arm the rule has stopped assigning to.

Fixes exist and they all amount to putting a floor under the trailing arm — never allocate below 20%, say — and every one of them gives back some of the patient benefit to buy back some of the power. The design space between “fair coin” and “fully adaptive” is continuous, and both endpoints are worse than the middle for almost any weighting of the two objectives.

The power loss is the allocation, to two points

The thirty points of power are attributed to the allocation arithmetic, and the attribution can be completed rather than asserted, using nothing but the two rates and the two splits.

At 45% and 25% the comparison’s variance is p(1 − p)/n summed over the arms. A fair coin gives 0.2475/100 + 0.1875/100 = 0.004350. The adaptive split of 173 and 27 gives 0.2475/173 + 0.1875/27 = 0.008375 — a ratio of 1.93, so a standard error 1.39 times larger. Inverting the coin’s 85.0% power gives a signal of 3.00 standard errors; dividing by 1.39 gives 2.16, which is a power of 57.9%.

The counted figure is 55.9%. Two points, and the two points are the one thing the calculation leaves out: 173 is a mean, the split varies from trial to trial, and power is concave in the smaller arm’s size, so a variable allocation is worth slightly less than a fixed one at the same average.

The same arithmetic says where the uncertainty actually lives. Of the adaptive design’s variance of 0.008375, the twenty-seven-patient arm contributes 0.00694483% of it. Five sixths of the trial’s uncertainty sits on 13.5% of its patients, which is the sharpest form of what an unequal split does to a difference and is why the rule’s success at treating people well is the same fact as its failure to establish anything.

What a floor on the trailing arm would buy

The fixes are described as putting a floor under the trailing arm and giving back some benefit to buy back some power. The same variance formula prices that exchange without any further simulation.

With a fraction f of the two hundred patients guaranteed to the worse arm, the variance is 0.2475/(200(1 − f)) + 0.1875/(200f), and the power follows from the ratio to the coin’s 0.004350:

floor on the better arm power
none — 13.5% as it falls out 173 58%
20% 160 71%
30% 140 81%
50% — a fair coin 100 85%

Those are computed at the mean split rather than over the distribution of splits, so each is a couple of points optimistic in the way the 57.9 against 55.9 above shows. The shape is not in doubt.

A floor of 30% gives back nine tenths of the coin’s power and keeps 40 of the 73 extra patients the unconstrained rule delivers. A floor of 20% keeps 60 of the 73 and three quarters of the power gap. Neither endpoint of the design space is on the frontier: the unconstrained rule spends 27 points of power on its last thirteen patients, and the coin spends seventy-three patients to buy the last four points.

That is the arithmetic behind the essay’s own claim that both endpoints are worse than the middle, and it says where the middle is. It also leaves the validity problem exactly where it was — a floor changes the split and not the fact that the split was chosen by the outcomes, so the 7.8% at zero drift survives every row of that table.

What is actually being traded

Set the three quantities beside each other, because the argument for these designs is usually made with one of them and the case has to be made with all three.

Patients inside the trial. 173 of 200 on the better arm against 100. That is the gain, it is real, and it is what the method is for.

The trial’s own conclusion. 55.9% power against 85.0%. A trial that fails to establish the difference leaves every future patient with the same uncertainty the trial started with.

And the trial’s validity. 7.8% against 5% with nothing wrong, and 58.2% under a drift that no trial is free of. The first two are a trade a reasonable person could take. The third is not a trade — it is a design whose error rate is not the one it states, and no amount of benefit to the patients inside buys that.

The ordering matters. If the validity problem were fixed — by a randomisation test, by a blocked analysis, by pre-specified weights — the remaining argument would be an honest one about whose interests the trial serves. Made without fixing it, the argument is about a benefit that is real and a cost that has been left out.

There is one more asymmetry worth putting on the record. The patient benefit is realised inside a single trial and is bounded by its size — seventy-three extra patients on the better arm, here. The validity cost is not bounded by anything: a trial whose error rate is 58% instead of 5% contributes a conclusion to a literature that other trials, guidelines and patients then rely on, and the harm has no natural limit because it is not confined to the people who were in the room.

20 adaptive trials, 30% against 30%. Each line is one trial allocating patients one at a time by the arm's own posterior. The average final share on the first arm — neither is better here is 47.5%, with a standard deviation of 24.1 points across these 20 trials, and 10 of them finished with most of the units on the worse arm. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.
Fig. 6 Twenty trials with no difference between the arms at all. The rule still commits — it has no way to know there is nothing to find — and the trials that commit hardest are the ones whose comparison the test then believes most.

What the drift number does not say

The 58.2% is measured with a drift of 0.4 — a success rate rising from about 10% to about 50% over the course of the trial — which is a large trend, and it is fair to ask whether real trials drift that far.

Some do and most do not. The useful question is not whether the number is realistic but how fast it grows, and the curve answers that: the error rate is already 25.6% at a drift of 0.2, and it is above the nominal level at every drift measured including zero. There is no threshold below which the design is safe, because the mechanism has no threshold in it — any correlation between arrival time and outcome turns an allocation that depends on early outcomes into a comparison confounded with time.

That is the general form of the argument and it is the reason the repair belongs in the design rather than in a judgement about how much drift is plausible. An experimenter who blocks the analysis by arrival time does not need to know how large the trend is, or whether there is one at all; a fair coin does not need to know either. Only the unadjusted adaptive analysis needs the assumption, and it needs it to be exactly true.

A trend that is in both arms, read as an effect in one. Both arms have a success rate of 30% and it rises by the same amount over the course of the trial, so there is no difference anywhere to find. A fair coin finds none: 5.4% at no drift and 5.6% at 0.5, the nominal rate throughout, and that is what randomisation buys. The adaptive rule starts at 8.4% with no drift at all — the allocation is a function of the outcomes, so the two are not independent — and reaches 75% at a drift of 0.5, because the early patients went to whichever arm looked good early and the late ones went to the other. Comparing within 16 blocks of arrival time removes the part the drift caused, bringing 75% down to 1.9%, and leaves the part the adaptation caused: 6.8% at no drift.
Fig. 7 Sixteen blocks rather than eight. More blocks make each one more internally comparable and leave fewer patients in each to compare, which is the same trade blocking always makes and is why the repair has a best size too.
What blocking is worth, 40 units. Each point is 4,000 randomisations of 40 units. The line is 1 − share, computed from the model with no simulation in it. Blocking on a factor carrying 70% of the variance leaves 32% of it, which is the same precision as 126 unblocked units.
Fig. 8 Blocking as the design field introduced it: the variance a blocking factor removes is the variance the blocks carry, exactly. The repair above is the same contrast taken within a block, with arrival time as the nuisance factor.

The refusal

The check requires the fair coin to be unharmed at every drift measured, and that is the load-bearing half.

If randomisation did not hold its level under a drift, the whole essay would be about the drift rather than about the adaptation, and the repair would be a different one. It holds — 4.98% at no drift and 4.98% at the largest — so what is being measured is the adaptation and nothing else.

The second refusal is the one that had to be added after the first run. The essay’s claim was going to be that a time trend breaks response-adaptive randomisation, and the measurement said the design was already broken at zero drift: 7.8%, with nothing to confound anything. The check now asserts that too, and in that order, because “adaptation is dangerous under a trend” and “adaptation is dangerous, and a trend makes it catastrophic” are different claims and only the second one is true.

The measurement that forced it is the cheapest kind there is: the same call with the drift set to zero. This site’s standing gotcha is an assertion conditioned on a parameter that turns out to be a fact about the default, and the general defence against it is to run the measurement at the boundary of its own range before writing the sentence it supports.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designAllocationBlockingConfoundingError rateRandomisationStatistical powerThompson samplingTime trend