What the exactness buys
Worth reading first: The experiments that could have happened · What a p-value does not say.
Everything so far in this field has been about a test that works. This essay is about what it costs and what it is bought with, because a repair with no price is a repair nobody had to think about, and this one has a price.
The price is power. What it buys is freedom from a number nobody has.
The comparison has to be fair before it is interesting
The obvious comparison is the randomisation test against the ordinary z, and it is not a comparison at all.
On this design the ordinary z rejects 9.05% of true nulls at 1.96. Against a real difference — 0.45 against 0.25, on two hundred patients — it detects 57.0%. That 57.0% is not power. Power is the rejection rate of a test at a stated level, and this is not a test at a stated level; it is a procedure that rejects nine percent of nothings and fifty-seven percent of somethings, and comparing it to a 5% test on the second number alone would be scoring one competitor against a rule the other is obeying.
So the honest competitor has to be constructed. Simulate the design under its own null, find the value the statistic exceeds 5% of the time, and use that instead of 1.96.
That value is 2.112.
Against it, the same z statistic on the same four hundred trials detects 42.3%. That is a real 5% test with real power, and it is what the randomisation test has to be measured against.
The randomisation test detects 23.0%.
What the 9.05% says, and what it does to the 57.0%
The unfair comparison can be made fair arithmetically, which is worth doing because it puts a size on the unfairness rather than merely naming it.
A two-sided test rejecting 9.05% of true nulls at 1.96 is a test whose reported standard error is too small by a factor r, with . That gives , so
The ordinary z is understating its own standard error by sixteen per cent.
Correct it and the 57.0% moves. In honest units the test is rejecting whenever the properly scaled statistic exceeds 1.692, and a 57.0% chance of that puts the effect at 1.869 standard errors. A test that rejects at 1.96 instead detects , which is 46.4%.
So the z’s advertised 57.0% is worth 46.4% once it is made a 5% test — ten and a half points of the apparent advantage being nothing but the level it was not holding. That is the correction the comparison needs before any question of what exactness costs can be asked.
Nineteen points
The gap is large and it is worth being blunt about it before explaining it away, because there is a temptation in a field like this to present the exact test as a free lunch.
It is not. On this design, against this alternative, conditioning on the rule costs about nineteen points of power relative to a critical value calibrated for the same design. Half of that experiment’s ability to detect a real difference is gone.
Some of that is recoverable. The randomisation test at B = 39 has a coarse reference set — the p-value can only take the values 1/40, 2/40, and so on — and increasing B buys resolution back. Some of it is not: the reference distribution is genuinely wider than the calibrated normal’s, because it accounts for allocation variability the calibrated approach smooths over by assuming a single nuisance value.
More re-randomisations buy power. Nothing buys back the rest. That is the trade, stated at its worst.
Where the power actually goes
The nineteen points have two sources and separating them matters, because one is a choice and the other is not.
The reference set is coarse. At B = 39 the p-value can only take forty values, so the test rejects only when the observed statistic is among the two most extreme of the forty. A statistic that is genuinely unusual but not in the top two is not rejected, and at a larger B it would be. This part shrinks as B grows and it is bought back with arithmetic.
The reference distribution is genuinely wider. This is the part that does not shrink. The calibrated z uses one number — 2.112 — for every trial, which amounts to asserting that every trial of this design has the same null distribution. It does not: a trial in which the rule happened to concentrate heavily has a different null distribution from one in which it stayed near balance, and the calibrated approach averages over that variation while the randomisation test conditions on it.
Conditioning is what costs the power. The randomisation test asks “given an allocation as concentrated as the one that happened, is this statistic surprising?” and a concentrated allocation makes large statistics unsurprising, so the bar rises for exactly the trials where the ordinary analysis was most wrong. That is the repair working, and the power it costs is not waste — it is the rejections the calibrated test makes on the trials where it should not.
Which reframes the nineteen points. They are not a tax paid for nothing; they are the rejections that were wrong in the first place, plus the resolution cost of a finite B. The first group is only visible as lost power if the trials where they occurred are counted as successes, and under the null they are not.
What the calibrated test is standing on
Now the other side, and it is where the field’s argument actually is.
The calibrated critical value of 2.112 was obtained by simulating the design under its null. To do that, the simulation needed a success rate — a value of p to generate outcomes from. It used 0.3.
That number is not a design parameter. It is not the level, or the sample size, or anything in the protocol. It is the rate at which patients succeed, which is the quantity the trial exists to learn.
And the critical value depends on it. Substantially.
Simulated at each of five rates, the value the statistic exceeds 5% of the time is:
- 1.668 at a success rate of 0.05,
- 1.820 at 0.1,
- 2.106 at 0.3,
- 2.302 at 0.5,
- 2.718 at 0.8.
From below 1.96 to well above 2.7, across a range of rates any real trial might sit anywhere in.
The value at 0.3 comes out as 2.106 here and 2.112 in the power comparison above, because the two are independent simulations of the same quantity at different trial counts. The six-thousandths between them is simulation error and it is worth leaving visible: it is the scale at which these numbers are known, and it is two orders of magnitude smaller than the differences the rest of this section is about.
What using the wrong one costs
Calibrate at 0.3 and apply it wherever the trial actually lands:
- at a true rate of 0.8, the test’s real size is 12.40%;
- at 0.5, it is 8.40%;
- at 0.3, it is 5.04% — correct, because that is where it was calibrated;
- at 0.1, it is 0.60%;
- at 0.05, it is 0.12%.
A 5% test at one rate and nowhere else. And the failure is two-sided: too large a critical value at a low rate makes the test find nothing, too small a value at a high rate makes it find things that are not there, and the experimenter has no way of knowing which side of their calibration point the trial landed on until it is over — at which point the calibration has already been used.
The mechanism is the same one that broke the ordinary analysis in the first place. The rule chases outcomes, so how imbalanced the allocation becomes depends on how often anything succeeds. A trial where almost nothing works produces little to chase and stays near balance; one where most things work produces long runs and heavy imbalance. The null distribution of the statistic moves with the imbalance, so it moves with the rate.
The randomisation test needs none of it
Here is the payoff, and it is one sentence.
The randomisation test conditions on the outcomes that actually happened, rather than on a rate they were supposed to come from. Whatever the success rate turned out to be, the re-randomisations use that sequence, so the reference distribution is built at the trial’s own rate without anyone having to name it.
The nineteen points of power are the price of not having to know a number nobody knows.
Whether that is a good trade depends on how confidently the rate can be pinned down in advance, and this is a place where the honest answer is “sometimes very confidently”. A trial in a well-studied area with a control arm whose rate is known from a decade of prior work is one where calibrating at the right value is plausible, and where nineteen points of power is a lot to give up for insurance against something unlikely. A trial in a new area, or one where the control rate itself is uncertain, is the opposite.
What the field supplies is the price, not the decision.
The nuisance parameter, as a general shape
This field’s argument has a name in the wider subject, and naming it says how common the situation is.
A nuisance parameter is a quantity that is not the object of interest and that the analysis nevertheless depends on. The success rate here is one. So is σ in a t test, which is the canonical example and the one where the problem was solved so completely that it stopped looking like a problem: Student’s t exists because the z test’s critical value depends on a σ nobody has, and the t distribution is the exact answer that removes the dependence.
Three ways of dealing with a nuisance parameter, and this site now has an example of each.
Find a pivot. Construct a statistic whose distribution does not depend on the nuisance at all. That is what t does with σ, exactly and once. It is the best answer available and it is rarely available.
Plug in an estimate. Use σ̂, or p̂, and proceed as though it were the truth. Cheap, and its cost is measurable: the plug-in interval covers 78.8% where the version that carries the uncertainty covers 95.2%, on the same data.
Condition on something that makes it irrelevant. Hold fixed a quantity the nuisance parameter governs, and ask the question within that. That is what this field does — the outcome sequence is held fixed, and the success rate is a property of that sequence rather than an input to the analysis.
The third route is the one with the fewest assumptions and the least power, and that ordering is not a coincidence. Every assumption an analysis makes is information it is being given for free; refusing the information means paying for it out of the data.
A design where the calibration does work
It is worth naming a case where the cheaper repair is the right one, because the field’s own sibling supplies it.
Dropping the losers has the same structure — a design whose statistic is not what its table thinks — and it is repaired by exactly the calibration this essay is warning about: the critical value is solved for as 2.313 rather than 1.96, and the resulting test holds its level.
The difference is what the value depends on. For that design the critical value is a function of how many arms there were and when the interim was — both of which are design parameters, written in the protocol before any data exists. For this one it is a function of the success rate, which is not.
So “simulate the design under its null and use the value that comes out” is a good repair whose applicability has a precise condition: the null distribution must depend only on things the protocol fixes. Where it does, calibration is cheaper and more powerful than re-randomising. Where it depends on a nuisance parameter, calibration is a test at one value of that parameter, and the randomisation test is the one that does not care.
And where it costs nothing at all
One more measurement, because a price should be quoted against the cases where it is not charged.
On a fair-coin trial with no drift, the ordinary z is already a 5% test — it rejects 4.5% of true nulls — and there is no calibration to do and no nuisance parameter to worry about. The randomisation test run on the same trials gives 5.8%. Neither is systematically better and the difference is inside simulation error.
So the price attaches to the adaptation rather than to the method. A trial that did not adapt pays nothing for using this analysis, and gains nothing either. Everything in this field — the nineteen points, the documentation requirement, the two hundred thousand allocation steps — is the bill for a design decision taken earlier, and it belongs on the same page as that decision’s benefits rather than in the analysis section.
What the whole field costs, added up
Three prices, and they are different kinds of thing.
Arithmetic. B re-randomisations of an n-patient trial: 199,800 allocation steps and 399,600 Beta draws at B = 999 on two hundred patients, against one division for a closed form. Cheap in absolute terms, and it multiplies in a simulation study.
Power. Nineteen points against a fairly calibrated competitor, of which some is recoverable by increasing B and some is not.
Documentation. The allocation rule has to travel with the data in executable detail, which is the previous essay’s subject and the only one of the three that requires anybody to change what they do.
Against which: a test whose level is 5% at every success rate, under any drift, at any allocation balance, without a distributional assumption, without a large-sample argument, and without a variance estimate that has to be right.
The three prices are also not paid by the same person, which is worth noticing because it decides whether any of this gets adopted. The arithmetic is paid by whoever runs the analysis and is trivial. The power is paid by the trial, in the form of a larger sample or a missed effect, and is not trivial. The documentation is paid by whoever wrote the protocol, years earlier, and is free at the time and impossible to pay retrospectively.
That last asymmetry is the one that matters in practice. A trial that adapted and did not record its rule in executable form cannot be given this analysis afterwards at any price, and the only moment at which the cost was near zero has passed. Every other repair on this site can be applied to data that already exists; this one has a deadline.
The check, and the refusal that makes it mean something
Two claims are gated in this field’s library and they are the two sides of this essay. The exactness costs power — asserted against the calibrated competitor rather than the naive one, with the naive one’s true size checked in the same assertion so that a version comparing against an invalid test would fail. And the calibration does not transfer: the critical value moves by more than half again across the range of rates, and one calibration is a 5% test at one rate and a 12% test at another.
The refusal is the assertion this essay was written expecting to make and could not. The plan was that the
randomisation test would detect nearly what the z test does, and it does not — it loses nearly half the
power against a fair competitor. Writing that down rather than finding a setting where the gap was
smaller
is what the check enforces: assertTheExactnessCostsPower requires the loss to be more than ten points,
so a future change that quietly narrowed it would fail the gate and have to be explained rather than
absorbed.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The null the exactness is for — both name error rate, experimental design, nuisance parameter, randomisation test, reference distribution
- What a reference distribution costs to sample — both name critical value, error rate, p-value, randomisation test, reference distribution
- When the constraints run out — both name experimental design, p-value, randomisation test, reference distribution, statistical power
- A taper and a critical value — both name critical value, error rate, reference distribution, statistical power
- Choosing n after looking — both name adaptive design, error rate, nuisance parameter, statistical power
- Draws that repeat each other — both name critical value, p-value, randomisation test, reference distribution
Named objects
A flat tag is an object no other essay names yet.
Adaptive designCritical valueError rateExperimental designNuisance parameterp-valueRandomisation testReference distributionResponse-adaptive randomisationStatistical power