The experiments that could have happened
Worth reading first: Randomisation is not balance · When the looking happens.
The adaptive-design field ends with a repair it names and does not build.
Response-adaptive randomisation — send each new patient to whichever arm currently looks better — rejects a true null 9.2% of the time with no time trend anywhere in the trial, and catastrophically more with one. That essay says what would fix it: the analysis has to condition on the rule that produced the allocation. It does not do it, because doing it needs a whole piece of machinery that field does not have.
This one does. And the machinery is not new — it is the randomisation test from the design field, whose whole argument is that a p-value can come from the assignment mechanism rather than from a population, applied to a mechanism considerably more complicated than a coin.
What goes wrong, restated in one paragraph
The allocation rule looks at the outcomes so far and sends the next patient to the arm that is ahead. So an arm that gets a lucky run of successes early receives more patients, and those extra patients are allocated because of the successes that will be used to judge that arm.
The consequence is that the two groups’ sizes are correlated with their outcomes. The pooled z statistic assumes they are not — its standard error is computed as though the group sizes were fixed in advance — and a statistic whose standard error is wrong is not a statistic with the distribution its table says.
Nothing about the printed output records any of this. The trial reports two proportions, two group sizes and a p-value, exactly as a fair-coin trial would.
The construction
Three steps, and each of them is doing work.
Fix the outcomes. The null hypothesis being tested is the sharp one: this treatment changed nothing for anybody. Under it, each patient’s outcome is whatever it was, whichever arm they were sent to — the arm assignment did not matter to them, which is exactly what the null asserts. So the observed sequence of outcomes stops being a sample. It becomes a fixed list, and the only random thing left in the whole experiment is who got what.
Re-run the rule. Not a coin: the same adaptive rule, fed the same fixed outcomes in the same arrival order, with a fresh stream of random numbers. Patient 1’s outcome is whatever it was; the rule assigns them to an arm; that outcome updates the rule’s state; patient 2 arrives. What comes out is another allocation the trial genuinely could have made, drawn from exactly the distribution the real one came from.
Count. Do that B times, count how many of the re-randomised statistics are at least as extreme as the observed one, and
p = (1 + number at least as extreme) / (1 + B)
The +1 on both halves is not a rounding nicety. It is the observed allocation counting as one of its own reference draws, and it is what makes a sampled test exact at every B rather than only in the limit.
Why re-running the rule is legitimate
The step that deserves scrutiny is the second one, because it looks circular: the rule uses the outcomes, and the outcomes are what is being held fixed, so is the re-randomisation really drawing from the same distribution as the trial?
It is, and the argument is worth writing out.
Under the sharp null the outcome sequence y₁ … yₙ does not depend on the allocation in any way. So the joint distribution of (outcomes, allocation) factorises: the outcomes arrive however they arrive, and given them, the allocation is generated by the rule with probability the rule specifies. The trial’s actual allocation is one draw from that conditional distribution.
The re-randomisation draws from the same conditional distribution — same rule, same outcomes, same order, different random numbers. So the observed allocation is exchangeable with the B re-randomised ones, and a p-value formed by counting where the observed sits among them is exact by construction.
The circularity that seems to threaten this is real and it is on the other side. The rule genuinely does depend on the outcomes, which is precisely why the ordinary analysis fails: the z statistic’s null distribution is derived under fixed group sizes, and the group sizes here are outcome-dependent. The dependence is not a problem for the randomisation test because the randomisation test never assumed it away. It reproduces the dependence in every re-randomisation instead.
That is the general lesson and it is worth carrying past this field. A method that models a complication has to get the model right. A method that reproduces the complication has only to know what the complication was.
Where this construction comes from
None of the above is new machinery and it is worth saying whose it is, because the lineage is what makes the extension obvious in retrospect.
The design field’s randomisation test does exactly this for a coin: hold the outcomes fixed, enumerate every possible assignment, count how many give a difference as large as the observed one. For twenty units split ten and ten there are 184,756 assignments and the enumeration is exhaustive, so the test is exact with no B in it at all — a p-value computed rather than estimated.
Two things stop that working here.
The assignment space is too large to enumerate. Two hundred patients allocated one at a time gives 2²⁰⁰ possible allocations, which is not a number anything can walk. So the reference set has to be sampled, and sampling is what makes the +1 load-bearing rather than decorative.
And the assignments are not equally likely. With a coin, every assignment has the same probability and enumerating them is enumerating the distribution. With an adaptive rule they emphatically do not: an allocation that sends everybody to the arm that happened to do badly is astronomically improbable. So the reference set cannot be built by listing assignments — it has to be built by running the rule, which samples each allocation with exactly the probability the rule gives it.
That second point is the one that makes this an extension rather than an application. The coin’s version can be described as “shuffle the labels”; this one cannot, because shuffling labels would generate allocations the trial could never have produced and would weight them all equally.
What the reference distribution looks like
The figure at the top of this essay is the whole argument in one picture, and it is worth reading carefully because it is not a sampling distribution.
Every bar in it comes from the same two hundred patients with the same two hundred outcomes. What varies across the histogram is only the assignment — which patients ended up in which arm — and the rule that generated each assignment is the rule the trial actually used.
The observed statistic is 1.417. Two hundred and fifty-eight of the nine hundred and ninety-nine re-randomisations reach it, so the p-value is (1 + 258)/(1 + 999) = 0.2590.
The grey curve is the normal distribution the ordinary analysis would use. It is visibly narrower. This particular reference distribution’s own 5% point is 2.101, not 1.96 — so a statistic between those two values is one the ordinary analysis calls significant and this one does not, and the gap between the two curves’ tails is exactly the 9.2% that should have been 5%.
The trial in the figure is a mild one
The reference distribution drawn here has a 5% point of 2.101, and turning that into a rejection rate says how representative the picture is.
A statistic from a distribution whose own 5% point is 2.101, read against 1.96, exceeds the threshold about 6.8% of the time. The counted rate for the ordinary analysis across six hundred trials is 9.2%. So the trial drawn for this essay’s hero is a less distorted trial than the average one: to produce 9.2% the typical adaptive trial’s own 5% point has to sit near 2.28.
That is worth saying because a reader takes the width of the histogram as the size of the problem, and here it understates it. The gap between the grey curve and the histogram in the figure is about two-thirds of the gap in a typical trial — a fact about which draw was plotted, and one no single picture can carry.
The fair-coin figure is the control for that reading and it behaves: its 5% point of 1.997 implies a rejection rate of 5.4%, which is 5% to within the noise of six hundred trials, and the histogram and the curve agree because there is nothing for them to disagree about.
Two reasons enumeration fails, and only one is about adaptation
The construction is defended by two obstacles to enumerating the reference set — the space is too large, and the assignments are not equally likely — and only the second is specific to this rule.
A fair coin over twenty units gives 184,756 assignments and enumerates comfortably. Over thirty it gives 155 million, which is a long afternoon; over forty, 1.4 × 10¹¹, which is not a computation anybody runs inside a simulation study. So enumeration dies at about thirty units whatever the rule is, and a coin-allocated trial of two hundred patients has to be sampled for exactly the same reason an adaptive one does.
What is genuinely particular to adaptation is the second obstacle. A coin’s reference set can be sampled by shuffling labels, because every assignment carries the same probability and a uniform draw over assignments is a draw from the right distribution. An adaptive rule’s cannot: its assignments differ in probability by many orders of magnitude, so the only way to sample each one at its own probability is to run the rule.
Sampling is forced by the size and running the rule is forced by the weighting, and separating the two says what a practitioner has to get right. A trial that is large but coin-allocated may use any off-the-shelf permutation routine; a trial that is small but adaptive may not, even where its space could be walked.
One arithmetic consequence of sampling is worth stating in the same breath. With B = 999 the p-value of 0.2590 carries a Monte Carlo standard error of 0.014, so it has two significant figures and the last two of its four decimals are the grid rather than the number. At the 5% boundary the same B gives ±0.007, which is enough to decide a rejection and not enough to report a p-value to three places.
Three analyses, counted
The test is now a thing that can be measured like any other on this site: run six hundred trials with the same success rate in both arms, so every rejection is a false one, and count.
With no time trend at all:
- the ordinary z against 1.96 rejects 9.2%,
- blocking by arrival time rejects 7.00%,
- the randomisation test rejects 4.0%.
The blocked analysis is the repair the adaptive field offers, and it is a partial one: it fixes the confounding between arrival time and allocation, and it does not fix the adaptation, because the adaptation is a problem even when every patient is identical.
The gap between 7.00% and 4.0% is the part of the failure that has nothing to do with time. Blocking compares like with like within a stretch of the recruitment, which removes any drift; what it cannot remove is that within each block the allocation is still a function of the outcomes in the blocks before it. Nothing that stratifies on a covariate can fix that, because the offending variable is not a covariate — it is the outcome itself, arriving earlier.
And the case that made the repair necessary
Now add a drift — a trend in the success rate across the recruitment, moving both arms equally, so it is not an effect and a correct analysis must not find one.
At a drift of 0.3 the ordinary z rejects 44.2% of true nulls. At 0.6 it rejects 87.7%. Those are not tests.
The blocked analysis repairs it and over-repairs: 4.50% at drift 0.3 and 0.83% at 0.6. It is conservative to the point of being nearly useless at the larger drift, because stratifying finely enough to absorb the trend throws away most of the comparison.
The randomisation test gives 4.3% and 4.2%.
It does not move. And the reason it does not move is the most satisfying thing in this field: a drift changes the outcomes, and the outcomes are held fixed. The reference distribution is built from the observed outcome sequence, whatever shape that sequence has, so a trend in it is not something the test has to model — it is already inside every one of the nine hundred and ninety-nine comparisons.
What the test assumes
It is worth being precise, because the list is unusually short and its shortness is the claim.
It does not assume the outcomes are normal, or identically distributed, or independent of arrival time. It does not assume the allocation is balanced, or that the groups are comparable, or that n is large. It does not need a variance estimate to be right, or a standard error to have the right form.
It assumes the sharp null: that under the null, each patient’s outcome would have been the same in either arm. That is stronger than “the average effect is zero” — it rules out an effect that helps some patients and hurts others by an equal amount — and it is the assumption under which the outcomes can be held fixed.
And it assumes the rule is known. That one is the field’s real content and it gets its own essay, because a test told the wrong rule fails as badly as the thing it was fixing.
The sharp null deserves one more paragraph, because it is the assumption most likely to be waved through and it is genuinely stronger than the null most analyses have in mind. “No average effect” permits a treatment that helps half the patients and harms the other half equally; the sharp null does not. So a rejection here is evidence against “the treatment did nothing to anyone”, which is a different and larger claim than “the treatment’s average effect is non-zero”, and a treatment with a genuinely offsetting effect would be caught by this test and missed by a test of means.
Whether that is a feature depends entirely on the question. For a safety analysis it is plainly the right null — “nobody was harmed” is what needs testing. For an effectiveness comparison where individual variation is expected and uninteresting, it is stronger than what anybody wants to assume, and the rejection has to be read carefully. The assumption is cheap here in the sense that nothing else has to be assumed alongside it, and cheap is not the same as innocuous.
That last figure is the check that makes the rest of the essay mean something. A test that always rejects less often than the alternatives is not necessarily correct — it might just be timid. Run on a design where the ordinary analysis is already right, the ordinary analysis gives 4.5% and the randomisation test gives 5.8%. They agree, and neither is systematically below the other.
What it costs
The exactness is not free and this site prices things in counts rather than in seconds.
One closed-form p-value is a division. One randomisation p-value is B re-randomisations of an n-patient trial, each of which walks every patient and draws two Beta variates for the adaptive rule. At B = 999 on two hundred patients that is 199,800 allocation steps and 399,600 Beta draws, for one number.
Two hundred thousand operations is nothing on a modern computer and it is not nothing in every context — a simulation study evaluating the design across a thousand scenarios pays it a thousand times over. The number is here so the trade is visible rather than assumed away, and the more interesting half of the trade is not the arithmetic at all. It is what the exactness is bought with, which turns out to be power.
The check, and the refusal that makes it mean something
The claim gated in this field’s library is the two-part one this essay is built on: the ordinary analysis rejects more than it claims on an adaptive trial, and conditioning on the rule puts it back — asserted on the same trials and the same seeds, at drift 0 and again at a drift where the ordinary analysis reaches 30%.
The refusal is the one that stops “conservative” being confused with “correct”. On a fair-coin trial with no drift the ordinary z is already right, so the randomisation test is required to agree with it rather than to reject less often. Had the assertion only ever required the new test to be nearer 5%, a test that rejected nothing at all would have passed it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
- A taper and a critical value
- Draws that repeat each other
- Errors generated from a fitted model
- Half a reference distribution
- How long a block a multiplier shares
- The analysis after three arms
- The corner the test is calibrated at
- The plus one and the round number
- The reference the covariates supply
- The test that needs the rule
- The triangle that was not the multiplier's
- Walking the admissible set
- What a multiplier cannot keep
- What a reference distribution costs to sample
- What the exactness buys
- The null the exactness is for
- A statistic that is exact twice
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Before the trial and after — both name p-value, randomisation test, reference distribution, sharp null
- Half a reference distribution — both name p-value, randomisation test, reference distribution, sharp null
- What a reference distribution costs to sample — both name error rate, p-value, randomisation test, reference distribution
- A probe chosen from the design — both name experimental design, randomisation test, reference distribution
- A probe nobody chose — both name randomisation test, reference distribution, sharp null
- A proposal that moves more than two units — both name experimental design, randomisation test, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Adaptive designBlockingError rateExperimental designp-valueRandomisation testReference distributionResponse-adaptive randomisationSharp null