The reference distribution the design supplies

What the exactness buys

Against a z test calibrated to reject exactly 5% of true nulls on this design, the randomisation test loses nineteen points of power. What it buys is that the calibration needs the success rate — which moves the critical value from 1.668 to 2.718 and is the quantity the trial was run to find out.

Worth reading first: The experiments that could have happened · What a p-value does not say.

Everything so far in this field has been about a test that works. This essay is about what it costs and what it is bought with, because a repair with no price is a repair nobody had to think about, and this one has a price.

The price is power. What it buys is freedom from a number nobody has.

What the exactness costs, against a competitor held to the same size. Power against a real difference of 0.45 to 0.25, on 400 trials of 200 patients. The first row is not a 5% test — it rejects 9.0% of true nulls on this design — so its 57.0% is not power that anybody has. The second is the honest competitor: the same statistic against a critical value simulated for this design under its null, which is a 5% test at the success rate it was calibrated at and at no other. Against that, the randomisation test loses 19.2 points. That is the price, and what it buys is the row's second line.
Fig. 1 Three analyses of the same four hundred trials with a real difference in them. The first is not a 5% test and its power is not power anybody has; the second and third are, and they differ by nineteen points.

The comparison has to be fair before it is interesting

The obvious comparison is the randomisation test against the ordinary z, and it is not a comparison at all.

On this design the ordinary z rejects 9.05% of true nulls at 1.96. Against a real difference — 0.45 against 0.25, on two hundred patients — it detects 57.0%. That 57.0% is not power. Power is the rejection rate of a test at a stated level, and this is not a test at a stated level; it is a procedure that rejects nine percent of nothings and fifty-seven percent of somethings, and comparing it to a 5% test on the second number alone would be scoring one competitor against a rule the other is obeying.

So the honest competitor has to be constructed. Simulate the design under its own null, find the value the statistic exceeds 5% of the time, and use that instead of 1.96.

That value is 2.112.

Against it, the same z statistic on the same four hundred trials detects 42.3%. That is a real 5% test with real power, and it is what the randomisation test has to be measured against.

The randomisation test detects 23.0%.

What the 9.05% says, and what it does to the 57.0%

The unfair comparison can be made fair arithmetically, which is worth doing because it puts a size on the unfairness rather than merely naming it.

A two-sided test rejecting 9.05% of true nulls at 1.96 is a test whose reported standard error is too small by a factor r, with 2[1Φ(1.96/r)]=0.09052\left[1 - \Phi(1.96/r)\right] = 0.0905. That gives 1.96/r=1.6921.96/r = 1.692, so

r=1.16.r = 1.16 .

The ordinary z is understating its own standard error by sixteen per cent.

Correct it and the 57.0% moves. In honest units the test is rejecting whenever the properly scaled statistic exceeds 1.692, and a 57.0% chance of that puts the effect at 1.869 standard errors. A test that rejects at 1.96 instead detects Φ(1.8691.96)\Phi(1.869 - 1.96), which is 46.4%.

So the z’s advertised 57.0% is worth 46.4% once it is made a 5% test — ten and a half points of the apparent advantage being nothing but the level it was not holding. That is the correction the comparison needs before any question of what exactness costs can be asked.

Nineteen points

The gap is large and it is worth being blunt about it before explaining it away, because there is a temptation in a field like this to present the exact test as a free lunch.

It is not. On this design, against this alternative, conditioning on the rule costs about nineteen points of power relative to a critical value calibrated for the same design. Half of that experiment’s ability to detect a real difference is gone.

Some of that is recoverable. The randomisation test at B = 39 has a coarse reference set — the p-value can only take the values 1/40, 2/40, and so on — and increasing B buys resolution back. Some of it is not: the reference distribution is genuinely wider than the calibrated normal’s, because it accounts for allocation variability the calibrated approach smooths over by assuming a single nuisance value.

More re-randomisations buy power. Nothing buys back the rest. That is the trade, stated at its worst.

Where the +1 matters, and why nobody has noticed that it does. The true size of the two rules at every B, computed rather than simulated: under the null the count of re-randomisations reaching the observed statistic is uniform over {0 … B}, so both sizes are integer arithmetic. With the +1 the size is (⌊α(B+1)⌋)/(B+1), which never exceeds 5%. Without it the size is (⌊αB⌋+1)/(B+1), which is larger except at B = 19, 39, 59 — the values with B + 1 a multiple of 1/α, and the values everybody uses. At B = 19 the two rules are the same rule; at B = 20 the uncorrected one is a 9.5% test. The marks are simulated on 500 trials of 200 patients, as the second route to the same numbers.
Fig. 2 And a floor on how far B can be reduced to save arithmetic: below about twenty re-randomisations the test’s attainable sizes are so coarse that the level being claimed may not be among them.

Where the power actually goes

The nineteen points have two sources and separating them matters, because one is a choice and the other is not.

The reference set is coarse. At B = 39 the p-value can only take forty values, so the test rejects only when the observed statistic is among the two most extreme of the forty. A statistic that is genuinely unusual but not in the top two is not rejected, and at a larger B it would be. This part shrinks as B grows and it is bought back with arithmetic.

The reference distribution is genuinely wider. This is the part that does not shrink. The calibrated z uses one number — 2.112 — for every trial, which amounts to asserting that every trial of this design has the same null distribution. It does not: a trial in which the rule happened to concentrate heavily has a different null distribution from one in which it stayed near balance, and the calibrated approach averages over that variation while the randomisation test conditions on it.

Conditioning is what costs the power. The randomisation test asks “given an allocation as concentrated as the one that happened, is this statistic surprising?” and a concentrated allocation makes large statistics unsurprising, so the bar rises for exactly the trials where the ordinary analysis was most wrong. That is the repair working, and the power it costs is not waste — it is the rejections the calibrated test makes on the trials where it should not.

Which reframes the nineteen points. They are not a tax paid for nothing; they are the rejections that were wrong in the first place, plus the resolution cost of a finite B. The first group is only visible as lost power if the trials where they occurred are counted as successes, and under the null they are not.

What the calibrated test is standing on

Now the other side, and it is where the field’s argument actually is.

The calibrated critical value of 2.112 was obtained by simulating the design under its null. To do that, the simulation needed a success rate — a value of p to generate outcomes from. It used 0.3.

That number is not a design parameter. It is not the level, or the sample size, or anything in the protocol. It is the rate at which patients succeed, which is the quantity the trial exists to learn.

And the critical value depends on it. Substantially.

The cheap repair needs a number nobody has. The obvious alternative to re-randomising is to simulate the design under its null once and use the critical value that comes out — which is what the arm-dropping design does, where the critical value has to be solved for and is 2.313. It does not transfer here. The rule chases outcomes, so how imbalanced the allocation gets depends on how often anything succeeds, and the critical value moves from 1.668 at a success rate of 0.05 to 2.718 at 0.8. Calibrated at 0.3 and used at 0.8 the test's real size is 12.4%; used at 0.05 it is 0.12%. The randomisation test needs none of this, because it conditions on the outcomes that happened rather than on a rate they were supposed to come from.
Fig. 3 The critical value the design needs, at five different success rates, with the true size of a test using the p = 0.3 value written above each point. It is a curve, not a constant.

Simulated at each of five rates, the value the statistic exceeds 5% of the time is:

  • 1.668 at a success rate of 0.05,
  • 1.820 at 0.1,
  • 2.106 at 0.3,
  • 2.302 at 0.5,
  • 2.718 at 0.8.

From below 1.96 to well above 2.7, across a range of rates any real trial might sit anywhere in.

The value at 0.3 comes out as 2.106 here and 2.112 in the power comparison above, because the two are independent simulations of the same quantity at different trial counts. The six-thousandths between them is simulation error and it is worth leaving visible: it is the scale at which these numbers are known, and it is two orders of magnitude smaller than the differences the rest of this section is about.

What using the wrong one costs

Calibrate at 0.3 and apply it wherever the trial actually lands:

  • at a true rate of 0.8, the test’s real size is 12.40%;
  • at 0.5, it is 8.40%;
  • at 0.3, it is 5.04% — correct, because that is where it was calibrated;
  • at 0.1, it is 0.60%;
  • at 0.05, it is 0.12%.

A 5% test at one rate and nowhere else. And the failure is two-sided: too large a critical value at a low rate makes the test find nothing, too small a value at a high rate makes it find things that are not there, and the experimenter has no way of knowing which side of their calibration point the trial landed on until it is over — at which point the calibration has already been used.

The mechanism is the same one that broke the ordinary analysis in the first place. The rule chases outcomes, so how imbalanced the allocation becomes depends on how often anything succeeds. A trial where almost nothing works produces little to chase and stays near balance; one where most things work produces long runs and heavy imbalance. The null distribution of the statistic moves with the imbalance, so it moves with the rate.

The randomisation test needs none of it

Here is the payoff, and it is one sentence.

The randomisation test conditions on the outcomes that actually happened, rather than on a rate they were supposed to come from. Whatever the success rate turned out to be, the re-randomisations use that sequence, so the reference distribution is built at the trial’s own rate without anyone having to name it.

The nineteen points of power are the price of not having to know a number nobody knows.

Whether that is a good trade depends on how confidently the rate can be pinned down in advance, and this is a place where the honest answer is “sometimes very confidently”. A trial in a well-studied area with a control arm whose rate is known from a decade of prior work is one where calibrating at the right value is plausible, and where nineteen points of power is a lot to give up for insurance against something unlikely. A trial in a new area, or one where the control rate itself is uncertain, is the opposite.

What the field supplies is the price, not the decision.

Three analyses of the same trials, with no time trend at all600 trials of 200 patients with the same success rate in both arms — every rejection below is a false one. Allocation is response-adaptive, so the assignment is a function of the outcomes it is later compared with. The ordinary z rejects 9.2%, blocking by arrival time gives 7.00%, and conditioning on the rule that produced the allocation gives 4.0%.the ordinary z against 1.969.2%blocked by arrival time7.0%the randomisation test4.0%5%, which is what all three claim600 true nulls of 200 patients, 39 re-randomisations each, one set of seeds9.2% · 7.00% · 4.0%, all at a nominal 5%
Fig. 4 And the reason the decision is not simply “calibrate carefully”. The comparison above holds the design fixed; a drift across the recruitment moves the null distribution again, and a calibration done for a flat trial is wrong for a drifted one in exactly the same way it is wrong for a different rate.

The nuisance parameter, as a general shape

This field’s argument has a name in the wider subject, and naming it says how common the situation is.

A nuisance parameter is a quantity that is not the object of interest and that the analysis nevertheless depends on. The success rate here is one. So is σ in a t test, which is the canonical example and the one where the problem was solved so completely that it stopped looking like a problem: Student’s t exists because the z test’s critical value depends on a σ nobody has, and the t distribution is the exact answer that removes the dependence.

Three ways of dealing with a nuisance parameter, and this site now has an example of each.

Find a pivot. Construct a statistic whose distribution does not depend on the nuisance at all. That is what t does with σ, exactly and once. It is the best answer available and it is rarely available.

Plug in an estimate. Use σ̂, or p̂, and proceed as though it were the truth. Cheap, and its cost is measurable: the plug-in interval covers 78.8% where the version that carries the uncertainty covers 95.2%, on the same data.

Condition on something that makes it irrelevant. Hold fixed a quantity the nuisance parameter governs, and ask the question within that. That is what this field does — the outcome sequence is held fixed, and the success rate is a property of that sequence rather than an input to the analysis.

The third route is the one with the fewest assumptions and the least power, and that ordering is not a coincidence. Every assumption an analysis makes is information it is being given for free; refusing the information means paying for it out of the data.

A design where the calibration does work

It is worth naming a case where the cheaper repair is the right one, because the field’s own sibling supplies it.

Dropping the losers has the same structure — a design whose statistic is not what its table thinks — and it is repaired by exactly the calibration this essay is warning about: the critical value is solved for as 2.313 rather than 1.96, and the resulting test holds its level.

The difference is what the value depends on. For that design the critical value is a function of how many arms there were and when the interim was — both of which are design parameters, written in the protocol before any data exists. For this one it is a function of the success rate, which is not.

So “simulate the design under its null and use the value that comes out” is a good repair whose applicability has a precise condition: the null distribution must depend only on things the protocol fixes. Where it does, calibration is cheaper and more powerful than re-randomising. Where it depends on a nuisance parameter, calibration is a test at one value of that parameter, and the randomisation test is the one that does not care.

The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 258 of the 999 re-randomisations reach it, and the p-value is (1 + 258)/(1 + 999) = 0.2590. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 2.101.
Fig. 5 The construction that does not care. Its reference distribution is built from the outcome sequence the trial produced, so a trial at a success rate of 0.05 and one at 0.8 each get their own — without anybody choosing.

And where it costs nothing at all

One more measurement, because a price should be quoted against the cases where it is not charged.

On a fair-coin trial with no drift, the ordinary z is already a 5% test — it rejects 4.5% of true nulls — and there is no calibration to do and no nuisance parameter to worry about. The randomisation test run on the same trials gives 5.8%. Neither is systematically better and the difference is inside simulation error.

Three analyses of the same trials, with no time trend at all. 600 trials of 200 patients with the same success rate in both arms — every rejection below is a false one. Allocation is by fair coin. The ordinary z rejects 4.5%, blocking by arrival time gives 5.17%, and conditioning on the rule that produced the allocation gives 5.8%.
Fig. 6 The comparison where nothing is broken. All three analyses land near 5%, which is what stops everything above from being a story about a conservative test dressed as a correct one.

So the price attaches to the adaptation rather than to the method. A trial that did not adapt pays nothing for using this analysis, and gains nothing either. Everything in this field — the nineteen points, the documentation requirement, the two hundred thousand allocation steps — is the bill for a design decision taken earlier, and it belongs on the same page as that decision’s benefits rather than in the analysis section.

What the whole field costs, added up

Three prices, and they are different kinds of thing.

Arithmetic. B re-randomisations of an n-patient trial: 199,800 allocation steps and 399,600 Beta draws at B = 999 on two hundred patients, against one division for a closed form. Cheap in absolute terms, and it multiplies in a simulation study.

Power. Nineteen points against a fairly calibrated competitor, of which some is recoverable by increasing B and some is not.

Documentation. The allocation rule has to travel with the data in executable detail, which is the previous essay’s subject and the only one of the three that requires anybody to change what they do.

Against which: a test whose level is 5% at every success rate, under any drift, at any allocation balance, without a distributional assumption, without a large-sample argument, and without a variance estimate that has to be right.

The three prices are also not paid by the same person, which is worth noticing because it decides whether any of this gets adopted. The arithmetic is paid by whoever runs the analysis and is trivial. The power is paid by the trial, in the form of a larger sample or a missed effect, and is not trivial. The documentation is paid by whoever wrote the protocol, years earlier, and is free at the time and impossible to pay retrospectively.

That last asymmetry is the one that matters in practice. A trial that adapted and did not record its rule in executable form cannot be given this analysis afterwards at any price, and the only moment at which the cost was near zero has passed. Every other repair on this site can be applied to data that already exists; this one has a deadline.

What one p-value costs, counted in allocation steps. A closed-form p-value is one division. A randomisation p-value is B re-randomisations of an 200-patient trial, each of which walks every patient and draws two Beta variates for the adaptive rule — 199,800 allocation steps and 399,600 Beta draws at B = 999. Counted rather than timed, because a duration is a fact about a machine and a count is a fact about the procedure. It is also the only cost the test has: no assumption about the outcome distribution, no large-sample argument, no nuisance parameter.
Fig. 7 The first of the three prices, counted rather than timed — a fact about the procedure rather than about any machine.

The check, and the refusal that makes it mean something

Two claims are gated in this field’s library and they are the two sides of this essay. The exactness costs power — asserted against the calibrated competitor rather than the naive one, with the naive one’s true size checked in the same assertion so that a version comparing against an invalid test would fail. And the calibration does not transfer: the critical value moves by more than half again across the range of rates, and one calibration is a 5% test at one rate and a 12% test at another.

The refusal is the assertion this essay was written expecting to make and could not. The plan was that the randomisation test would detect nearly what the z test does, and it does not — it loses nearly half the power against a fair competitor. Writing that down rather than finding a setting where the gap was smaller is what the check enforces: assertTheExactnessCostsPower requires the loss to be more than ten points, so a future change that quietly narrowed it would fail the gate and have to be explained rather than absorbed.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designCritical valueError rateExperimental designNuisance parameterp-valueRandomisation testReference distributionResponse-adaptive randomisationStatistical power