The allocations the rule could have made, from these exact patients
One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 258 of the 999 re-randomisations reach it, and the p-value is (1 + 258)/(1 + 999) = 0.2590. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 2.101.
The reference distribution the design supplieswide6 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
600 trials of 200 patients with the same success rate in both arms — every rejection below is a false one. Allocation is response-adaptive, so the assignment is a function of the outcomes it is later compared with. The ordinary z rejects 9.2%, blocking by arrival time gives 7.00%, and conditioning on the rule that produced the allocation gives 4.0%.
The true size of the two rules at every B, computed rather than simulated: under the null the count of re-randomisations reaching the observed statistic is uniform over {0 … B}, so both sizes are integer arithmetic. With the +1 the size is (⌊α(B+1)⌋)/(B+1), which never exceeds 5%. Without it the size is (⌊αB⌋+1)/(B+1), which is larger except at B = 19, 39, 59 — the values with B + 1 a multiple of 1/α, and the values everybody uses. At B = 19 the two rules are the same rule; at B = 20 the uncorrected one is a 9.5% test. The marks are simulated on 500 trials of 120 patients, as the second route to the same numbers.
Power against a real difference of 0.45 to 0.25, on 400 trials of 200 patients. The first row is not a 5% test — it rejects 9.0% of true nulls on this design — so its 57.0% is not power that anybody has. The second is the honest competitor: the same statistic against a critical value simulated for this design under its null, which is a 5% test at the success rate it was calibrated at and at no other. Against that, the randomisation test loses 19.2 points. That is the price, and what it buys is the row's second line.
How often each analysis reports an effect when the average treatment effect is exactly zero and the effect varies between units, at 150 units with 25% treated. The permutation test on the difference in means reads 4.20% where the effect is constant — where the two nulls coincide and its exactness applies — and 22.93% where the effect varies with a standard deviation of 3. The same test on the studentised difference reads 6.27% there, and the ordinary large-sample t, which makes no exactness claim at all, reads 6.60%.
Rejection rates for both statistics under both nulls, at 25% of 150 units treated, with the weak-null readings taken at an effect spread of 3. The difference in means is exact under the sharp null and rejects 22.93% of true weak nulls. The studentised difference is exact under the sharp null — 4.07% — and reads 6.27% under the weak one. The repair is a change of statistic inside the same construction: the same re-randomisations, the same fixed outcomes, a different number compared across them.
Where it is used
8 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 8 different questions.
- The experiments that could have happened The reference distribution the design supplies
- The test that needs the rule The reference distribution the design supplies
- The plus one and the round number The reference distribution the design supplies
- The reference the covariates supply Balancing on what was recorded first
- What the exactness buys The reference distribution the design supplies
- The score is the modelling Coverage without a distribution
- The null the exactness is for The reference distribution the design supplies
- A statistic that is exact twice The reference distribution the design supplies