Concept

Adaptive design — where it appears

An experiment that changes what it does next from what it has already seen — the arm split, the sample size, which arms survive. Every such change makes the assignment a function of the data, so the analysis has to condition on the rule that produced it rather than on a fixed design.

Named by 10 essays across 3 fields — each of them below, with the objects they name alongside it.

What the interim sees, at an effect of 1. The same 60 observations, estimated two ways. Keeping the arms separate gives 0.995, which is σ. Pooling them without separating the arms — the price of staying blind to the comparison — gives 1.114, against the identity √(1 + Δ²/4σ²) = 1.118. The sample size is proportional to the variance, so a blinded design at this effect asks for 25% more units than it needs, and it does so systematically rather than by chance.

Choosing n after looking

Re-estimating the sample size from an interim is the one adaptation with a defence, and the defence is exactly what it costs: an analyst kept blind to the arms measures a spread that contains the effect, so the design overshoots by 1 + Δ²/4σ². Re-estimating the effect instead breaks the error rate.

adaptive · Stopping
The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 258 of the 999 re-randomisations reach it, and the p-value is (1 + 258)/(1 + 999) = 0.2590. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 2.101.

The experiments that could have happened

An adaptive trial's allocation is a function of the outcomes it will later be compared against, so the ordinary analysis rejects a true null 9.2% of the time. Hold the outcomes fixed, re-run the rule that assigned them, and count — the same statistic against a reference distribution the trial could actually have drawn from is back at 4.0%.

exact · Reference
20 adaptive trials, 45% against 25%. Each line is one trial allocating patients one at a time by the arm's own posterior. The average final share on the better arm is 84.7%, with a standard deviation of 10.3 points across these 20 trials. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.

Randomising towards the winner

Allocating more patients to the arm that is doing better is the humane thing to want and it buys nothing statistically: at a fixed total it costs thirty points of power. And because the allocation is a function of the outcomes, the ordinary test on it rejects a true null 7.8% of the time before any time trend is applied — and 58% after one.

adaptive · Randomisation
The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs a fair coin rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 155 of the 999 re-randomisations reach it, and the p-value is (1 + 155)/(1 + 999) = 0.1560. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 1.946.

The test that needs the rule

A randomisation test assumes almost nothing about the data and one thing about the experiment. Tell it a fair coin produced an allocation that an adaptive rule produced — which is what every off-the-shelf permutation routine does — and it rejects 8.0% of true nulls where knowing the rule gives 4.0%.

exact · Reference
The best of 8 arms, tested as though it were the only one. 8,000 trials with no effect in any arm. Stage one runs 8 arms at 60 each, the best is carried forward, and stage two adds 30 more to it and to the control. The histogram is where the final statistic lands and the curve is the standard normal it is being read against — shifted right, because the arm was chosen for being ahead. 10.5% of these trials clear 1.96 against a claimed 5%, and the value that actually holds the rate for this design is 2.313.

Dropping the losers

Carrying the best of eight arms forward and testing it at 1.96 rejects a true null 10.3% of the time — the hypothesis was chosen by looking at the data, so the statistic is a maximum wearing a single comparison's clothes. The value that holds the rate is 2.313, and it has to be solved for.

adaptive · Stopping
One experiment finding out where to look. A single run of the fully sequential design: 40 runs, the first 8 placed at the guess K = 1, then the model refitted and the design revised after every 2. The marks are the settings the runs were made at. The horizontal lines are where a design built at the truth K = 3 would have put them — 1.875 and 10.00 — and the rule walks onto them without being told: its estimate of K after the first eight runs was 2.694, and by the end 2.765 against a truth of 3. The whole experiment is 96.5% as efficient as the design that knew the answer, where running all 40 at the guess would have been 81.1%.

The design that stops guessing

Every repair so far protects a guess. The alternative is to run part of the experiment, estimate the parameter from it, and design the rest at the estimate — which recovers most of what a threefold wrong guess costs, and has a best moment to stop guessing that is earlier than anyone expects.

robust · Local design
Where the +1 matters, and why nobody has noticed that it does. The true size of the two rules at every B, computed rather than simulated: under the null the count of re-randomisations reaching the observed statistic is uniform over {0 … B}, so both sizes are integer arithmetic. With the +1 the size is (⌊α(B+1)⌋)/(B+1), which never exceeds 5%. Without it the size is (⌊αB⌋+1)/(B+1), which is larger except at B = 19, 39, 59 — the values with B + 1 a multiple of 1/α, and the values everybody uses. At B = 19 the two rules are the same rule; at B = 20 the uncorrected one is a 9.5% test. The marks are simulated on 500 trials of 120 patients, as the second route to the same numbers.

The plus one and the round number

A sampled randomisation test counts the observed allocation as one of its own reference draws, and the correction is invisible at B = 19, 39, 59 and 999 — every value anybody uses. At B = 20 the version without it is an 8.00% test where the corrected one is 3.80%, and the convention protecting everybody is a preference for round numbers minus one.

exact · Reference
Three estimates of the same effect, and the honest one is the worst. 8,000 trials in which every arm has an effect of exactly 0. The arm carried forward is reported by its first stage at +0.187 above the truth — it was chosen for being ahead — and by its second stage at -0.0032, which is unbiased by construction because the selection could not see it. The combination that gets published is at 0.123, exactly the share of the first stage's bias the first stage contributes. Root mean squared error: 0.240, 0.261, 0.182 — the unbiased estimate is the least accurate of the three.

The estimate after the choice

An arm chosen for being ahead is ahead by more than it should be, and the trial then publishes the average of the stage that chose it and the stage that did not. The unbiased estimate is the one built from a third of the data — and it is the least accurate of the three.

adaptive · Curse
What the interval covers, after a design that read the data. 800 experiments of 12 runs, all at the same truth. The first pair is a design fixed in advance; the second is one whose settings were chosen from the first stage's own outcomes. If choosing the design from the data broke the inference, the second pair would sit below the first, and it does not — 91.3% against 92.6%. What does move the coverage is the shape of the interval rather than the design: the Wald interval assumes the estimate is normal around its own standard error and falls short under both designs, and the profile interval, computed from the same residual sums of squares with no derivative in it, covers 94.3% and 94.9%.

What a design chosen from the data costs

Two fields on this site measured what happens when a rule reads the data, and the error rate broke both times. A design that reads the data to decide where to put its runs breaks nothing — and the control that proves it also finds what the real shortfall is.

robust · Local design
The cheap repair needs a number nobody has. The obvious alternative to re-randomising is to simulate the design under its null once and use the critical value that comes out — which is what the arm-dropping design does, where the critical value has to be solved for and is 2.313. It does not transfer here. The rule chases outcomes, so how imbalanced the allocation gets depends on how often anything succeeds, and the critical value moves from 1.668 at a success rate of 0.05 to 2.718 at 0.8. Calibrated at 0.3 and used at 0.8 the test's real size is 12.4%; used at 0.05 it is 0.12%. The randomisation test needs none of this, because it conditions on the outcomes that happened rather than on a rate they were supposed to come from.

What the exactness buys

Against a z test calibrated to reject exactly 5% of true nulls on this design, the randomisation test loses nineteen points of power. What it buys is that the calibration needs the success rate — which moves the critical value from 1.668 to 2.718 and is the quantity the trial was run to find out.

exact · Nuisance

Named alongside it

The objects these essays reach for when they reach for this one.

Error rateExperimental designp-valueRandomisation testReference distributionCritical valueMonte CarloOptional stoppingResponse-adaptive randomisationSelection biasStatistical powerTwo-stage design

All concepts