Randomising towards the winner
Worth reading first: Randomisation is not balance · When the looking happens.
A trial that is going well is a trial in which one arm is better, and every patient assigned to the other one is being given something worse on purpose. Response-adaptive randomisation is the obvious response: update the allocation as the results arrive, so that later patients are more likely to receive whatever is winning.
It works, in the sense that it does what it says. What it costs is measurable in two currencies, and the second one is not the one anybody expects.
What it delivers
Two arms with success rates of 45% and 25%, two hundred patients, allocated by drawing from each arm’s Beta posterior and assigning to whichever draw is higher.
86.6% of patients reach the better arm on average — 173 of the two hundred, against a hundred under a fair coin. That is the whole point of the method and it delivers it.
The number that is not in the figure is the one worth adding: 173 on the better arm means 27 on the worse one, and 27 is not zero. A design that is being defended on the grounds that it stops giving people the worse treatment still gives it to twenty-seven of two hundred, and would have given it to a hundred under a fair coin. The gain is real and it is 73 patients rather than the whole of the ethical objection.
The spread around that average matters as much as the average, which is why the figure draws twenty trials rather than one. The rule commits early on the basis of a handful of outcomes, and a run of bad luck on the genuinely better arm sends most of a trial to the worse one. That is not a rare failure; it is a tail of the same distribution the 86.6% is the mean of.
What it costs in power
The same trials, tested for a difference between the arms at the end:
| allocation | patients on the better arm | trials detecting the difference |
|---|---|---|
| a fair coin | 100 | 85.0% |
| adaptive | 173 | 55.9% |
Thirty points of power, at the same total. The reason is the allocation field’s first result read backwards: a comparison’s variance is σ₁²/n₁ + σ₂²/n₂, and it is dominated by whichever arm has fewer observations. Putting 173 patients on one arm leaves 27 on the other, and 27 is what the comparison is worth.
So the trade is real and it is not subtle: the design buys better treatment for the patients inside the trial and pays for it with the trial’s ability to establish anything for the patients outside it. A design that ends at 55.9% power is one that will fail to demonstrate a difference it was built on top of, nearly half the time.
And it breaks the test before any trend arrives
Here is the part that is not a trade.
Run the same design with both arms at 30% — no difference at all — and test it in the ordinary way. The fair coin rejects 4.98% of the time, which is the nominal level and is what randomisation buys. The adaptive rule rejects 7.8%.
There is no time trend here, no drift, nothing changing. The inflation is structural: the number of patients on each arm is a function of the outcomes on each arm, so the counts and the successes are not independent, and the two-proportion test assumes they are. An arm that got lucky early receives more patients, which makes its observed rate a mixture of its lucky start and its subsequent regression — and the test has no way to know that the denominator was chosen by the numerator.
What a time trend does
Now add the thing every real trial has. Patients recruited later differ from patients recruited earlier — healthier, better managed, recruited from different sites — so let the success rate rise over the course of the trial in both arms equally. There is still no difference between the arms, and a correct analysis must find none.
| drift over the trial | a fair coin | adaptive | adaptive, blocked by arrival time |
|---|---|---|---|
| 0 | 4.98% | 7.8% | 7.0% |
| 0.4 | 4.98% | 58.2% | 2.55% |
The fair coin is untouched. That is what randomisation is for and it is worth stating plainly: a trend that moves both arms alike cannot bias a comparison in which the two arms are balanced across time, and a fair coin balances them across time by construction.
The adaptive rule rejects a true null 58.2% of the time. By the time the response has drifted upward, the allocation is heavily skewed towards whichever arm looked good early — so the early, low-response patients are mostly on one arm and the late, high-response patients are mostly on the other, and the difference between the arms is the drift with a label on it.
The repair is the field’s oldest idea
Compare patients only with the patients who arrived around the same time — divide the trial into blocks of arrival order, compute the difference within each block, and combine — and the drift cancels inside every block whatever the allocation did.
It works: 58.2% falls to 2.55%.
That is blocking, which this site introduced as a way of removing variance from a field trial, arriving in a place that looks nothing like a field. The nuisance factor is time rather than soil, the blocks are arrival windows rather than plots, and the arithmetic is identical: a contrast taken within a block cannot contain anything the block shares.
Two things are worth noticing about that repair, and neither is a footnote.
It over-corrects at large drifts. 2.55% against a nominal 5% is conservative, not exact. The blocks have very unequal allocation inside them by then — that is what the adaptive rule did — and the weighted combination handles that cautiously.
And it does not fix the other problem. At zero drift the blocked analysis rejects 7.0% against the unadjusted 7.8%. Blocking removes the part the trend caused and leaves the part the adaptation caused, because the dependence between the allocation and the outcomes is inside each block too.
So the honest summary is that blocking repairs the confounding and not the adaptation, and the adaptation needs a different repair — a randomisation test that conditions on the realised allocation, which is what randomisation actually buys and is exact here for the same reason it is exact there.
Where this sits among the site’s other adaptations
Three designs on this site now change something after looking at the data, and setting them in order makes the principle visible.
Re-estimating the sample size looks at a nuisance parameter and holds its error rate exactly. The interim cannot see the comparison.
Randomising towards the winner looks at the comparison itself — that is what “which arm is winning” means — and uses it to decide the allocation. Its error rate is 7.8% before any complication is added.
Dropping the losers also looks at the comparison, and holds its rate anyway, because the boundary is computed for the design that was actually run.
So the dividing line is not whether an adaptation looks at the comparison. It is whether the analysis knows that it did. The internal pilot is safe because the interim never touched the comparison; the arm-dropping design is safe because the critical value was solved for a procedure that selects; and response-adaptive randomisation is unsafe as usually analysed because the test applied at the end is the one that would have been right for a fair coin.
That reading makes the repair obvious in principle and awkward in practice: the analysis must condition on the rule that produced the allocation. It is exactly the randomisation test’s argument — the reference distribution is the set of outcomes the assignment mechanism could have produced — applied to a mechanism far more complicated than a coin.
Why the rule commits so hard
The allocation curves in the figures rise fast and flatten, and the shape is worth explaining because it decides everything above.
Thompson sampling assigns each patient by drawing once from each arm’s posterior and taking the higher draw. Early on the posteriors are wide and the draws overlap, so the allocation is near even. As observations accumulate the posteriors separate, the overlap shrinks, and the probability of assigning to the trailing arm falls towards zero — which is the rule working as designed, since the whole point is to stop giving people the worse treatment.
The consequence is that the trailing arm’s sample size stops growing, so the evidence that it is worse stops accumulating. The rule reaches a conclusion and then defends it by collecting no more data that could overturn it, and it does this whether the conclusion is right or wrong.
That is visible in the twenty paths: the ones that commit to the wrong arm early do not come back. There is no mechanism in the rule that recovers, because recovery requires observations from the arm the rule has stopped assigning to.
Fixes exist and they all amount to putting a floor under the trailing arm — never allocate below 20%, say — and every one of them gives back some of the patient benefit to buy back some of the power. The design space between “fair coin” and “fully adaptive” is continuous, and both endpoints are worse than the middle for almost any weighting of the two objectives.
The power loss is the allocation, to two points
The thirty points of power are attributed to the allocation arithmetic, and the attribution can be completed rather than asserted, using nothing but the two rates and the two splits.
At 45% and 25% the comparison’s variance is p(1 − p)/n summed over the arms. A fair coin gives 0.2475/100 + 0.1875/100 = 0.004350. The adaptive split of 173 and 27 gives 0.2475/173 + 0.1875/27 = 0.008375 — a ratio of 1.93, so a standard error 1.39 times larger. Inverting the coin’s 85.0% power gives a signal of 3.00 standard errors; dividing by 1.39 gives 2.16, which is a power of 57.9%.
The counted figure is 55.9%. Two points, and the two points are the one thing the calculation leaves out: 173 is a mean, the split varies from trial to trial, and power is concave in the smaller arm’s size, so a variable allocation is worth slightly less than a fixed one at the same average.
The same arithmetic says where the uncertainty actually lives. Of the adaptive design’s variance of 0.008375, the twenty-seven-patient arm contributes 0.006944 — 83% of it. Five sixths of the trial’s uncertainty sits on 13.5% of its patients, which is the sharpest form of what an unequal split does to a difference and is why the rule’s success at treating people well is the same fact as its failure to establish anything.
What a floor on the trailing arm would buy
The fixes are described as putting a floor under the trailing arm and giving back some benefit to buy back some power. The same variance formula prices that exchange without any further simulation.
With a fraction f of the two hundred patients guaranteed to the worse arm, the variance is 0.2475/(200(1 − f)) + 0.1875/(200f), and the power follows from the ratio to the coin’s 0.004350:
| floor | on the better arm | power |
|---|---|---|
| none — 13.5% as it falls out | 173 | 58% |
| 20% | 160 | 71% |
| 30% | 140 | 81% |
| 50% — a fair coin | 100 | 85% |
Those are computed at the mean split rather than over the distribution of splits, so each is a couple of points optimistic in the way the 57.9 against 55.9 above shows. The shape is not in doubt.
A floor of 30% gives back nine tenths of the coin’s power and keeps 40 of the 73 extra patients the unconstrained rule delivers. A floor of 20% keeps 60 of the 73 and three quarters of the power gap. Neither endpoint of the design space is on the frontier: the unconstrained rule spends 27 points of power on its last thirteen patients, and the coin spends seventy-three patients to buy the last four points.
That is the arithmetic behind the essay’s own claim that both endpoints are worse than the middle, and it says where the middle is. It also leaves the validity problem exactly where it was — a floor changes the split and not the fact that the split was chosen by the outcomes, so the 7.8% at zero drift survives every row of that table.
What is actually being traded
Set the three quantities beside each other, because the argument for these designs is usually made with one of them and the case has to be made with all three.
Patients inside the trial. 173 of 200 on the better arm against 100. That is the gain, it is real, and it is what the method is for.
The trial’s own conclusion. 55.9% power against 85.0%. A trial that fails to establish the difference leaves every future patient with the same uncertainty the trial started with.
And the trial’s validity. 7.8% against 5% with nothing wrong, and 58.2% under a drift that no trial is free of. The first two are a trade a reasonable person could take. The third is not a trade — it is a design whose error rate is not the one it states, and no amount of benefit to the patients inside buys that.
The ordering matters. If the validity problem were fixed — by a randomisation test, by a blocked analysis, by pre-specified weights — the remaining argument would be an honest one about whose interests the trial serves. Made without fixing it, the argument is about a benefit that is real and a cost that has been left out.
There is one more asymmetry worth putting on the record. The patient benefit is realised inside a single trial and is bounded by its size — seventy-three extra patients on the better arm, here. The validity cost is not bounded by anything: a trial whose error rate is 58% instead of 5% contributes a conclusion to a literature that other trials, guidelines and patients then rely on, and the harm has no natural limit because it is not confined to the people who were in the room.
What the drift number does not say
The 58.2% is measured with a drift of 0.4 — a success rate rising from about 10% to about 50% over the course of the trial — which is a large trend, and it is fair to ask whether real trials drift that far.
Some do and most do not. The useful question is not whether the number is realistic but how fast it grows, and the curve answers that: the error rate is already 25.6% at a drift of 0.2, and it is above the nominal level at every drift measured including zero. There is no threshold below which the design is safe, because the mechanism has no threshold in it — any correlation between arrival time and outcome turns an allocation that depends on early outcomes into a comparison confounded with time.
That is the general form of the argument and it is the reason the repair belongs in the design rather than in a judgement about how much drift is plausible. An experimenter who blocks the analysis by arrival time does not need to know how large the trend is, or whether there is one at all; a fair coin does not need to know either. Only the unadjusted adaptive analysis needs the assumption, and it needs it to be exactly true.
The refusal
The check requires the fair coin to be unharmed at every drift measured, and that is the load-bearing half.
If randomisation did not hold its level under a drift, the whole essay would be about the drift rather than about the adaptation, and the repair would be a different one. It holds — 4.98% at no drift and 4.98% at the largest — so what is being measured is the adaptation and nothing else.
The second refusal is the one that had to be added after the first run. The essay’s claim was going to be that a time trend breaks response-adaptive randomisation, and the measurement said the design was already broken at zero drift: 7.8%, with nothing to confound anything. The check now asserts that too, and in that order, because “adaptation is dangerous under a trend” and “adaptation is dangerous, and a trend makes it catastrophic” are different claims and only the second one is true.
The measurement that forced it is the cheapest kind there is: the same call with the drift set to zero. This site’s standing gotcha is an assertion conditioned on a parameter that turns out to be a fact about the default, and the general defence against it is to run the measurement at the boundary of its own range before writing the sentence it supports.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A count that has to be estimated — both name allocation, blocking, randomisation
- A dictionary that is a product — both name allocation, blocking, randomisation
- Guessing one arm in three — both name allocation, error rate, randomisation
- How many subjects — both name allocation, blocking, statistical power
- Simpson's reversal is a region, not a table — both name allocation, confounding, randomisation
- The reversal a coin cannot prevent — both name allocation, blocking, randomisation
Named objects
A flat tag is an object no other essay names yet.
Adaptive designAllocationBlockingConfoundingError rateRandomisationStatistical powerThompson samplingTime trend