Concept

Variance reduction — where it appears

Making an estimate less variable without changing what it estimates, which is what a balancing rule is for. How much it is worth depends on the shape the outcome enters by, and against a shape the rule cannot see it is worth nothing at all.

Named by 18 essays across 9 fields — each of them below, with the objects they name alongside it.

The six best bases of 2 functions, and what each protects. Every cell is R²(g | span B) — the share of the imbalance in that shape a rule balancing that basis removes — computed from exact inner products between Hermite functions and indicators, with nothing simulated. The rows are ordered by their worst cell, which is the number an experimenter who does not know the shape is exposed to. The best row here guarantees 26.8% against every shape in the list, and the worst of the six guarantees 15.1%: the difference between them is entirely which subspace was picked, at the same cost per arrival.

A basis is a subspace

A balancing rule cannot tell one basis from another with the same span, so choosing what to hand it is choosing a subspace — and then what it removes of any outcome shape is a projection, computable exactly, with no trial anywhere in it.

basis · Criterion
One factor moves and the other does not. The two factors of the same average, each drawn against its own largest value so that they share an axis. The rate at which the five candidates disagree about the tuning parameter rises from 28.6% at 4 values on the list to 43.3% at 8, a factor of 1.52. What a disagreement costs, given that there was one, is 0.00975 ± 0.00224 and 0.00848 ± 0.00113 at the same two points — 0.5 standard errors apart, and the paired comparison on the draws that disagree under both lists puts it the other way. The guess this field was written to test was that a longer list makes disagreements commoner and each one smaller. The first half is right and there is no second half.

A rate times a size

A sweep reported what it costs to let every candidate choose its own tuning parameter and found it flat across the list. It was reporting a product, and the two things multiplied together do not behave the same way at all.

apiece · Order-selection
Four rules, three shapes, and no ordering that survives. The variance of the unadjusted treatment estimate under each rule, as a fraction of the variance a coin gives, over 450 trials of 200 units each. Against a covariate that enters linearly the rule that reads the number nearly halves it. Against a threshold at 1 it removes about a fifth. Against a quadratic every rule here is at or worse than a coin — they are all optimising a criterion that is one over the variance of an estimate in a model this outcome does not obey, and a constraint that helps nothing still costs something. Nothing in a trial says which column it is in.

Balanced on the wrong function

A rule that reads a covariate's numbers halves the variance of the treatment estimate, if the covariate enters the outcome as a straight line. If it enters as a threshold the rule is worth a fifth of that, and if it enters as a curve every rule here is worse than a coin.

shape · Blocking
Every split of 100 units, σ = 1 against 3. Each point is one integer split, with its variance computed exactly rather than simulated. The minimum is at 25:75, which is the ratio of the spreads 25:75, and equal allocation costs 25% more variance — the same as throwing away 20 of the 100 units. The shaded band is every split within 5% of the best, and it runs from 17% to 35%: sharp to state, flat to sit on.

Not half and half

The same units, the same measurements, the same analysis — and a different variance, decided before anything is measured. When the two arms have different spreads the best split is σ₁ : σ₂, equal allocation costs 2(σ₁²+σ₂²)/(σ₁+σ₂)², and at three to one that is a quarter of the experiment.

allocation · Allocation
The same 40 units, arranged two ways. Both designs estimate the same effect of 0.5 and both are unbiased — 0.488 and 0.497. The blocked design's estimate has standard deviation 0.318 against 0.692, a variance ratio of 0.21 where the model predicts 0.20.

The variance removed before the data

Arranging forty units in pairs rather than assigning them at random cuts the variance of the estimated effect to a fifth — and the fifth is knowable in advance, because it is exactly the share of the variance the pairs do not carry.

design · Blocking
A threshold in the tail is a threshold nothing balances. The share of a coin's imbalance in an indicator 1{x > c} that survives a rule which balances the covariate itself. The smooth curve is 1 − ρ² with ρ = φ(c)/√(p(1−p)), a closed form with no trial in it; the points are counted over 500 trials of 200 units at each threshold. At the median the two agree that about a third survives — the removed share is exactly 2/π — and by two standard deviations 86.9% survives. The closed form is exact in the limit and optimistic by a few points at this many units, because the rule balances the sample's mean rather than the population's.

A threshold in the tail

How much of a threshold's imbalance a balanced covariate removes is a correlation, and the correlation is a closed form. At the median it is exactly 2/π — the same 2/π a median split throws away — and two standard deviations out it is an eighth.

shape · Blocking
Stationary is not the same as convergent. How far each k-swap walk is from uniform after t steps, started at the least balanced admissible assignment of 410. Every one of these chains has a symmetric proposal and rejects by standing still, so every one of them is doubly stochastic and every one preserves the uniform distribution exactly. Only five of the six get there. Exchanging all six units of each arm is a single proposal — the complement — and the admissible set is closed under complement, so the walk takes it every time and oscillates between two assignments for ever: after 160 steps it has visited 1 state and sits 0.9976 from uniform. Its stationary distribution is a fact about the matrix; its limit does not exist.

Stationary is not convergent

A walk that exchanges every unit in each arm preserves the uniform distribution exactly and never gets near it. Every doubly stochastic matrix has the same stationary distribution; only some of them have a limit.

blocks · Randomisation
A budget of 4,000, at 1 and 20 a unit. Every affordable pair, enumerated. The best is 280 cheap units and 186 expensive ones — a ratio of 1.51, against the σᵢ/√cᵢ rule's 1.49. The unit rule, which says buy in the ratio of the spreads, lands at 66:197 and costs 17% more variance for the same money. Both rules are right about their own constraint; only one of them was asked.

The cost of a unit

Change the constraint from units to money and the allocation rule changes with it — from σᵢ to σᵢ/√cᵢ, which can point the other way. An arm that is noisy and expensive gets fewer units than the same arm would if the money were not the thing running out.

allocation · Allocation
The guarantee, as the basis is allowed more functions. The lower line is the best worst case over the six named shapes for a basis of each size, found by scoring every subset of the dictionary — an exact answer, since the problem is finite. One function guarantees 2.3%, which is nearly nothing; three guarantee 59.0% and the basis that does it is the covariate, its square and its cube, with no indicator in it. The upper line is the same problem with the basis drawn rather than fixed, which is worth 2.09 times as much at two functions and 1.32 at three. The two lines converge because a basis large enough to protect everything has nothing left to randomise over.

Which shapes are worth protecting

Choosing a basis by its worst case is a finite problem with an exact answer. The answer has no tie in it, which a maximin optimum is supposed to have — and the tie comes back, along with twice the guarantee, when the basis is drawn rather than chosen.

basis · Blocking
What a rule reads, against what the outcome uses. The variance of the treatment estimate relative to a coin's, for four things a rule might balance against three shapes the outcome might have, over 350 trials of 200 units. The diagonal is the easy part — a rule that reads the function the outcome uses removes about half the variance. What the table is for is the off-diagonal: reading the covariate alone is worth nothing against a quadratic (0.755), and reading all three is worth nearly as much against every shape as the matching rule is against its own (0.532, 0.493, 0.514).

Three functions of one number

A rule that balances the covariate is exposed to every shape the outcome might have. A rule that balances three functions of it costs two points of variance against the shape the first was built for and takes the worst case from a coin's to about half of it.

shape · Criterion
Two structures in three are made worse. What adjusting for every covariate measured does to the bias in the treatment's estimated effect, against adjusting for none, over 4000 randomly drawn structures of 6 covariates each. Each covariate is independently a common cause with probability 0.25, a cause of the treatment only, a cause of the outcome only, a cause of neither, a step on the causal path, or a common effect. The rule leaves a larger bias on 65.5% of structures, a smaller one on 33.8%, and the same on 0.7%. The share is a property of that population of structures rather than of adjustment, which is why the weights are stated; what does not depend on them is that the rule has no direction — it is not a conservative default that occasionally overcorrects, it is a rule whose error is whatever the structure happens to be.

Adjusting for everything

"Control for every covariate that was measured" leaves a larger bias than controlling for nothing on 65.5% of four thousand randomly drawn structures and a smaller one on 33.8%. Its squared error is 4.110 times that of using no covariate at all, and half of it sits in its worst tenth of structures.

collider · Conditioning
What a pilot buys, σ = 1 against 3. Each point is 6,000 two-stage experiments of 100 units: a pilot of m per arm, then the rest split by the pilot's own estimate of the two spreads. Above the line the pilot has made the experiment worse than not bothering. The best pilot here is 8 per arm at 0.809, against 0.800 for a designer who knew the spreads — so the rule recovers 96% of what knowing them is worth. A larger pilot estimates the ratio better and has less left to apply it to, which is why the curve turns.

Allocating on a guess

Every allocation rule in this field is a function of quantities the experiment is being run to find out. Fed a pilot's estimate of them, the rule that minimises the variance makes the experiment worse than not bothering — until the arms differ by about a factor of two, which is further than anyone would guess.

allocation · Allocation
Two groups, a baseline and a follow-up, and nothing happening in between — baseline reliability 0.6. 600 units in two pre-existing groups whose true means are 1.00 apart, read once at baseline and once at follow-up, with no change for anybody. The two groups' mean changes are −0.075 and −0.032, so the change-score analysis reports a group difference of 0.043. The regression of follow-up on baseline and group reports 0.409, against a closed form of (1 − λ) × 1.00 = 0.400: at any one baseline reading the two groups' lines sit that far apart, because each group's units regress towards their own group's mean. The pooled slope in this sample is 0.614, the baseline's reliability.

Two analyses of one baseline

Two groups read at baseline and again at follow-up, with no change for anybody. Subtracting the baseline reports a group difference of −0.0014 and adjusting for it reports 0.4008 — and each analysis is exactly right about one reason the groups started apart and wrong by 0.40 about the other.

paradox · Rtm
Which samples Wilson and Clopper–Pearson each cover, n = 50, p = 0.2. Each bar is the probability of one count, shaded by which interval built on that count contains 0.2. Both cover 95.1% of samples, only Wilson 0.0%, only Clopper–Pearson 1.6%, neither 3.3%. The correlation between their hits is 0.810, so on shared draws the variance of their difference is 4.891 times smaller than on independent ones.

The same draws for both methods

Two intervals computed on the same simulated datasets give a difference in coverage whose variance can be 4.891 times smaller than on separate datasets — or, for a pair that covers different samples, 1.164 times larger. Which one a comparison gets is an exact sum over the counts each interval covers, and a standard error that ignores the sharing covers 100.00% for one pair and 93.07% for the other.

method · Seeds
What the guess is worth, when it is worth anything. The variance cost of an even split relative to the variance-minimising one for a risk difference, against the first arm's proportion, with the second at 0.3. The cost is a pure number: it does not depend on the trial's size. It is exactly zero at 0.3 and at 0.70, where the two arms have the same p(1 − p); it is 0.19% at a half and 4.36% at a tenth. Across the whole range from a tenth to nine tenths it never exceeds 4.36%, which is what the variance-minimising rule is worth here — and what it is worth is the reason it is safe to use with a guess.

The arm whose variance is its answer

With a binary outcome the allocation rule is a function of the proportions the trial exists to estimate. It costs at most 4.36% of variance to ignore it anywhere between a tenth and nine tenths, because √(p(1−p)) stays within a factor of two of its peak across 98% of the unit interval.

allocation · Allocation
Estimating P(Z > 5) = 2.8665×10⁻⁷ with plain draws and with four proposals. One seed each. Plain simulation draws nothing past 5 in 100,000 and estimates zero throughout. At 100,000 draws the proposal N(5, 1) reads 1.009 of the truth, N(9, 1) 1.041, N(4.5, 0.25²) 1.039 and N(5, 0.3²) 1.008. Values above 2.2 are drawn at the top edge.

The draws aimed at the tail

The chance a standard normal exceeds 5 is 2.8665×10⁻⁷, and a plain simulation needs 349 million draws to estimate it to within ten per cent. Draws aimed at the tail and weighted back need 565. Aimed slightly too narrowly, the same method has an infinite variance, an interval that covers 86.0% and gets worse with more draws, and an effective sample size that reads healthier than a proposal that works.

method · Seeds
Three contrasts on one dataset, three different splits. The variance-minimising allocation for each of three ways of reporting the same two-arm comparison, against the first arm's proportion, with the second at 0.1. A risk difference wants the arm with the larger p(1 − p) to get more units; a log odds ratio wants it to get fewer, and the two curves are exact reflections of each other in the half line. A log risk ratio wants something else again. At a first-arm proportion of 0.6 they ask for 62.0%, 21.4% and 38.0% of the units. A trial reporting more than one of them cannot be optimal for either.

Two contrasts, one split

A risk difference wants 62.0% of the units in the first arm, a log risk ratio wants 21.4% and a log odds ratio wants 38.0% — on one dataset, with one pair of proportions. The difference's rule and the odds ratio's are exact reflections of each other, so no split can be near-optimal for both.

allocation · Allocation
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

method · Routes

Named alongside it

The objects these essays reach for when they reach for this one.

Monte CarloClosed formAllocation ruleEfficiencyExperimental designCovariate balanceRandomisationTreatment effectAllocation ratioImbalanceModel misspecificationSample size

All concepts