Concept

Allocation ratio — where it appears

How many units go to one arm for each one that goes to the other. It costs (1 + k)²/k observations per unit of effective size, so a ratio of two to one is 12.5% more expensive than balance and a ratio of ten to one is three times.

Named by 13 essays across 6 fields — each of them below, with the objects they name alongside it.

The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%.

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

stop · Stopping
Every split of 100 units, σ = 1 against 3. Each point is one integer split, with its variance computed exactly rather than simulated. The minimum is at 25:75, which is the ratio of the spreads 25:75, and equal allocation costs 25% more variance — the same as throwing away 20 of the 100 units. The shaded band is every split within 5% of the best, and it runs from 17% to 35%: sharp to state, flat to sit on.

Not half and half

The same units, the same measurements, the same analysis — and a different variance, decided before anything is measured. When the two arms have different spreads the best split is σ₁ : σ₂, equal allocation costs 2(σ₁²+σ₂²)/(σ₁+σ₂)², and at three to one that is a quarter of the experiment.

allocation · Allocation
Exact in the corner, where nothing was. Coverage of a nominal 95% interval on five designs, at a required half-width of 0.3. The first four are the two-arm field's own and the fifth is its corner — two variances, block sizes that swing by eight, and an allocation that alternates between five to one and one to five — where neither of that field's two conditions holds. The effective-size weights over-cover there at 98.40%; the weights h_b(λ) = (1/m_A + λ/m_B)⁻¹ cover at 94.84%, and at 94.84% when λ is estimated from the within-arm contrasts rather than known. Nothing here is supposed to move.

Weights that need only a ratio

A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.

corner · Nuisance
A trial designed 2:1:1, and what two scores deliver. 500 trials of 180 patients, three arms, a target of 2:1:1. The shaded bars are a minimisation score that divides each arm's count by the share that arm is supposed to receive before measuring the spread; it delivers 49.9% : 25.1% : 25.1%. The others are the same rule with the counts left raw, which delivers 33.4% : 33.3% : 33.3% — the balance it enforces inside every factor level is equality, and equality is what it gets. The marks are the shares that were asked for.

Balancing towards unequal targets

A three-arm trial allocating two to one to one is the ordinary case, and a balancing rule built from raw counts does not know it. It balances the arms towards equality inside every factor level, delivers a third to each arm, and reports that it minimised imbalance.

multiarm · Allocation
A budget of 4,000, at 1 and 20 a unit. Every affordable pair, enumerated. The best is 280 cheap units and 186 expensive ones — a ratio of 1.51, against the σᵢ/√cᵢ rule's 1.49. The unit rule, which says buy in the ratio of the spreads, lands at 66:197 and costs 17% more variance for the same money. Both rules are right about their own constraint; only one of them was asked.

The cost of a unit

Change the constraint from units to money and the allocation rule changes with it — from σᵢ to σᵢ/√cᵢ, which can point the other way. An arm that is noisy and expensive gets fewer units than the same arm would if the money were not the thing running out.

allocation · Allocation
3 arms against one control, 360 units in all. Every control size, enumerated. The best is 132 on the control and 76 on each arm — a ratio of 1.74, against √3 = 1.73. Splitting the units evenly over all 4 groups costs 7.2%, which is small; what the larger control also does is lower the correlation between the comparisons, from 0.50 to 0.37, and that changes which multiplicity correction is right.

One control, many arms

The control appears in every comparison, so it is worth √k treatment arms — and the same sharing makes the k tests correlated at n/(n+n₀), which is the quantity Bonferroni ignores. Both facts come out of one design decision, and it is the size of the control.

allocation · Multiplicity
A wrong weight costs width; a random weight costs level. Five weightings on a trial whose variance ratio drifts by a factor of 20.1 between the first block and the last, over 4000 runs. The rule that knows every λ_b covers at 95.1% and sets the width. One ratio for the whole trial is wrong for every block and costs nothing in level — 94.8% — while being 20% wider; equal weights are calibrated by an identity and 22% wider. The ratio estimated inside each block is the only rule aimed at the quantity that actually varies, and it is the only one that misses the level, at 92.0%: a weight computed from a handful of degrees of freedom is mostly noise, and noise in a weight is not a wrong weight. Modelling the drift across blocks recovers the oracle's width at 94.8%.

A ratio that changes between blocks

A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.

blocks · Nuisance
What a pilot buys, σ = 1 against 3. Each point is 6,000 two-stage experiments of 100 units: a pilot of m per arm, then the rest split by the pilot's own estimate of the two spreads. Above the line the pilot has made the experiment worse than not bothering. The best pilot here is 8 per arm at 0.809, against 0.800 for a designer who knew the spreads — so the rule recovers 96% of what knowing them is worth. A larger pilot estimates the ratio better and has less left to apply it to, which is why the curve turns.

Allocating on a guess

Every allocation rule in this field is a function of quantities the experiment is being run to find out. Fed a pilot's estimate of them, the rule that minimises the variance makes the experiment worse than not bothering — until the arms differ by about a factor of two, which is further than anyone would guess.

allocation · Allocation
Where the bias lands. The drift in the log variance ratio, fitted across 12 blocks over 4000 trials. E[log λ̂_b] is log λ_b plus ψ(k_B/2) − log(k_B/2) − ψ(k_A/2) + log(k_A/2), which depends on nothing but the degrees of freedom — so the tempting sentence is that it goes into the intercept and leaves the slope alone. It does not, because the blocks alternate between allocations and the alternation is correlated with the covariate being fitted: the lopsided blocks carry 0.5383 of bias and the even ones carry none. Uncorrected the slope reads 1.5597 against a truth of 1.5, which is 8.0 standard errors. Subtracting the two digammas block by block leaves 1.4976.

The bias that lands in the slope

The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.

blocks · Width
An exact test rejecting a true hypothesis a fifth of the time. How often each analysis reports an effect when the average treatment effect is exactly zero and the effect varies between units, at 150 units with 25% treated. The permutation test on the difference in means reads 4.20% where the effect is constant — where the two nulls coincide and its exactness applies — and 22.93% where the effect varies with a standard deviation of 3. The same test on the studentised difference reads 6.27% there, and the ordinary large-sample t, which makes no exactness claim at all, reads 6.60%.

The null the exactness is for

A permutation test is exact under the hypothesis that the treatment changed nothing for anybody. Under the hypothesis it changed nothing on average, with a quarter of the units treated and the effect varying between them, it rejects a true null 22.93% of the time.

exact · Nuisance
What the guess is worth, when it is worth anything. The variance cost of an even split relative to the variance-minimising one for a risk difference, against the first arm's proportion, with the second at 0.3. The cost is a pure number: it does not depend on the trial's size. It is exactly zero at 0.3 and at 0.70, where the two arms have the same p(1 − p); it is 0.19% at a half and 4.36% at a tenth. Across the whole range from a tenth to nine tenths it never exceeds 4.36%, which is what the variance-minimising rule is worth here — and what it is worth is the reason it is safe to use with a guess.

The arm whose variance is its answer

With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.

allocation · Allocation
One statistic that is right under both hypotheses. Rejection rates for both statistics under both nulls, at 25% of 150 units treated, with the weak-null readings taken at an effect spread of 3. The difference in means is exact under the sharp null and rejects 22.93% of true weak nulls. The studentised difference is exact under the sharp null — 4.07% — and reads 6.27% under the weak one. The repair is a change of statistic inside the same construction: the same re-randomisations, the same fixed outcomes, a different number compared across them.

A statistic that is exact twice

Dividing the difference in means by its own separate-variance standard error before permuting takes the rejection rate under a true weak null from 20.47% to 6.07%, keeps the exactness under the sharp null at 4.07%, and costs 0.8 points of power against a real effect. At an even split it changes nothing at all, in every draw.

exact · Nuisance
Three contrasts on one dataset, three different splits. The variance-minimising allocation for each of three ways of reporting the same two-arm comparison, against the first arm's proportion, with the second at 0.1. A risk difference wants the arm with the larger p(1 − p) to get more units; a log odds ratio wants it to get fewer, and the two curves are exact reflections of each other in the half line. A log risk ratio wants something else again. At a first-arm proportion of 0.6 they ask for 62.0%, 21.4% and 38.0% of the units. A trial reporting more than one of them cannot be optimal for either.

Two contrasts, one split

A risk difference wants 62.0% of the units in the first arm, a log risk ratio wants 21.4% and a log odds ratio wants 38.0% — on one dataset, with one pair of proportions. The difference's rule and the odds ratio's are exact reflections of each other, so no split can be near-optimal for both.

allocation · Allocation

Named alongside it

The objects these essays reach for when they reach for this one.

Experimental designAllocationNeyman allocationNuisance parameterVariance reductionBlockingCoverageEfficiencyFixed-width intervalMonte CarloSample sizeVariance ratio

All concepts