the-decisions-before-the-data

The same 40 units, arranged two ways

Both designs estimate the same effect of 0.5 and both are unbiased — 0.488 and 0.497. The blocked design's estimate has standard deviation 0.318 against 0.692, a variance ratio of 0.21 where the model predicts 0.20.

Decided before the dataslider: how much the pairs differ, 5 positionswide13 views

What else it draws

The same object, drawn to answer the other questions the essays put to it.

All 12,870 assignments, enumerated. The spread of the standardised imbalance is exactly 2/√n = 0.500, whatever the covariate's own distribution, and 33.3% of assignments differ by more than 0.5 standard deviations. Randomisation does not deliver balance; it delivers a known distribution of imbalance.

Under the null that the treatment does nothing, every unit's value is what it is whichever arm it landed in — so the 252 possible assignments give 252 possible differences, and that set IS the null distribution. Rejecting the most extreme 5% of them gives a test whose size is 4.76%, counted rather than assumed.

Both designs spend the same number of runs. The factorial estimates every main effect from every run; one-at-a-time estimates each from two conditions. The ratio of variances is (k+1)/2 — 2.0 at 3 factors — measured here at 2.00 over 2,000 experiments.

The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.

Four thousand pilots of 10 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 55.9% of the trials have less than 80% power and 11.1% less than 50%; the median trial has 76.8%.

From the pilot's own standard deviation, 55.9% of trials fall short of 80% power at an average size ×0.99 of the known-spread size. From its 80% upper limit, 19.8% fall short at ×1.65; from its 90% limit, 10.0% at ×2.12.

With the effect known, 80% is reached at 63 per arm. With the effect uncertain by 0.25 standard deviations it takes 113; by 0.5, 1268; by 0.75, no sample size at all, because the chance can never exceed the 74.8% prior probability that the effect is positive.

For an expected effect of half a standard deviation: 63 per arm with the effect known, 113 at an uncertainty of 0.25, 1268 at 0.5, and no finite number past 0.594, where the prior probability of a positive effect falls to 80%.

For a small shift, a cut at the mean keeps 63.7% of the information, so the trial needs 1.57 times the sample. A cut at the top tenth keeps 34.2% and needs 2.92 times; a cut two standard deviations out keeps 13.1%.

The outcome as measured reaches 80% power at 63 per arm; cut at the control mean, at 102, 1.62 times as many.

Sixty-four per arm, no difference between the arms, candidate cuts spread between one standard deviation below and above the mean, and the most significant comparison reported. With one cut fixed in advance the false-positive rate is 4.2%; choosing among five it is 17.7%, and among seventeen 27.4%.

The treated arm is the control arm moved half a standard deviation, for every patient. Above the threshold lie 15.9% of control patients and 30.9% of treated ones — a difference of 15.0% that reads as a share of patients who respond, when every patient responded by the same amount.

Where it is used

15 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 15 different questions.

All 80 figures

FieldsThreadsSeriesConceptsAll essaysSearch