Decided before the data

The variance removed before the data

Arranging forty units in pairs rather than assigning them at random cuts the variance of the estimated effect to a fifth — and the fifth is knowable in advance, because it is exactly the share of the variance the pairs do not carry.

Forty units, twenty to be treated and twenty not, and the effect estimated as the difference between the two averages. The units are not identical: they arrive in pairs that share something — two plots in the same field, two patients at the same hospital, two measurements taken on the same day — and whatever they share affects the outcome.

There are two ways to allocate the treatment, both correct, both unbiased, and they differ in how much the answer moves from experiment to experiment.

The same 40 units, arranged two waysBoth designs estimate the same effect of 0.5 and both are unbiased — 0.488 and 0.497. The blocked design's estimate has standard deviation 0.318 against 0.692, a variance ratio of 0.21 where the model predicts 0.20.0200400600-1012the estimated effectexperiments out of 4,000the true effect, 0.5unblocked, sd 0.69blocked, sd 0.324,000 randomisations, block spread 2variance ratio 0.21
Fig. 1 The same forty units, arranged two ways, across four thousand randomisations. Both distributions are centred on the true effect. Only one of them is narrow.

The two arrangements

Complete randomisation ignores the pairing. Choose twenty of the forty units at random, treat them, and compare the averages. It is what “randomised” means without further qualification, and it may put both members of a pair in the same arm, or neither.

Blocking uses the pairing. Within each of the twenty pairs, treat one member and not the other, chosen at random. Every pair contributes one treated and one control unit, and the effect is the average of the twenty within-pair differences.

Both are randomised. Both are unbiased — measured across four thousand randomisations, the completely randomised design estimates the true effect of 0.5 and so does the blocked one. Nothing about the allocation biases anything, and any argument for blocking that rests on “balance” is arguing for the wrong thing.

What differs is the spread. The blocked estimate has a standard deviation of 0.318; the completely randomised one has 0.692. The variance ratio is 0.212.

Why the difference is exact rather than approximate

The reason the blocked design is tighter is worth stating as arithmetic rather than as intuition, because the arithmetic gives the size of the gain in advance.

Write each unit’s outcome as its pair’s shared value, plus the unit’s own noise, plus the treatment effect if it was treated. In the blocked design the estimate is an average of within-pair differences, and in every one of those differences the shared value appears twice with opposite signs. It cancels. Not approximately, not on average: it is not in the estimate at all.

So the blocked estimate’s variance contains only the unit noise, while the completely randomised estimate’s variance contains the unit noise and the pair-to-pair variation, because a random split of forty units into two halves does not balance the pairs.

The prediction that follows is that the variance ratio is

σ² / (σ² + β²)

where β is the spread between pairs and σ the spread within them — the share of the total variance the blocks do not carry. With σ = 1 and β = 2, that is 1/5 = 0.200, against a measured 0.212. The agreement is the two-routes rule applied to an experimental design rather than to a distribution: a count over four thousand randomisations and a closed form derived from the model, sharing no arithmetic.

What blocking is worth, 40 units. Each point is 4,000 randomisations of 40 units. The line is 1 − share, computed from the model with no simulation in it. Blocking on a factor carrying 70% of the variance leaves 32% of it, which is the same precision as 126 unblocked units.
Fig. 2 The measured variance ratio against the share of the variance the blocks carry, with the closed form drawn over it. The line is not fitted to the points.

What the gain is worth in units

A variance ratio is not the currency an experiment is planned in. The currency is units, and the conversion is direct: a design with a fifth of the variance has the precision of an experiment five times the size.

Forty blocked units here buy the precision of 189 unblocked units. That is the number worth carrying, because it is what blocking is competing with: any argument that the pairing is inconvenient has to be weighed against nearly a hundred and fifty extra units.

It is also why blocking is the cheapest thing in this field. It costs no observations, no additional measurement, and no assumption that the completely randomised design does not also make. It is a change in the order in which the same units are allocated.

The variance ratio is one minus the pair correlation

The 0.212 is not a number that had to be simulated, and writing down what it equals turns the whole field into one parameter.

Let each unit’s outcome be its pair’s value, with variance σp2\sigma_p^2, plus its own noise, with variance σu2\sigma_u^2. The blocked estimate averages twenty within-pair differences, each with variance 2σu22\sigma_u^2, so its variance is 2σu2/202\sigma_u^2/20. The completely randomised estimate compares two means of twenty drawn from units whose variance is σu2+σp2\sigma_u^2 + \sigma_p^2, so its variance is 2(σu2+σp2)/202(\sigma_u^2+\sigma_p^2)/20.

The ratio is

σu2σu2+σp2=1ρ,\frac{\sigma_u^2}{\sigma_u^2 + \sigma_p^2} = 1 - \rho,

where ρ is the correlation between the two members of a pair. The gain from blocking is entirely the intraclass correlation and nothing else — not the number of pairs, not the size of the effect, not the noise level, only what share of the variation is between pairs rather than within them.

The figures here are drawn at a pair spread of 2 against a unit noise of 1, so ρ=4/5\rho = 4/5 and the ratio should be 0.200. Measured, 0.212, the difference being the finite-population correction that comes from splitting exactly forty units rather than sampling them.

Which prices the design in units

A variance ratio converts into sample size directly, since variance goes as 1/n: blocking multiplies the effective sample by 1/(1ρ)1/(1-\rho).

Here that is five. Forty units blocked into twenty pairs carry the precision of about 189 completely randomised units, measured, or 200 by the closed form.

And the formula says where that stops being worth having. At a pair correlation of 0.5 the multiplier is 2 — blocking is worth a doubling of the trial. At 0.2 it is 1.25, which on forty units is the difference between forty and fifty. At 0.05 it is 1.05, and the pairing is worth two units.

So the question to ask of any proposed blocking factor is not whether the units within a block are similar; it is how much of the outcome’s variance lies between blocks, because that number is the answer. A blocking variable explaining four fifths of the variation quintuples the experiment. One explaining a twentieth does nothing at all, and costs the analysis a constraint it now has to respect.

The case where it is worth nothing

A method whose advantage has been demonstrated once has been demonstrated against one configuration, and this site’s habit is to show the same machinery a case it must refuse.

Blocking on a factor that shares nothing is that case. When the two members of a pair have no common value — when the pairing is arbitrary — the variance ratio measured the same way is 1.03: no gain, and a whisker of loss.

The same 40 units, arranged two ways. Both designs estimate the same effect of 0.5 and both are unbiased — 0.488 and 0.497. The blocked design's estimate has standard deviation 0.318 against 0.314, a variance ratio of 1.03 where the model predicts 1.00.
Fig. 3 The same comparison when the pairs share nothing at all. The two distributions sit on top of each other, which is what the closed form requires.

This matters more than as a check on the arithmetic. It says the gain is not produced by blocking itself, but by blocking on something relevant, and the phrase “the study was blocked” carries no information without saying what it was blocked on and how much of the variation that factor accounts for.

The small loss is real and is worth naming honestly. A blocked analysis spends a degree of freedom per block, so where the blocks carry nothing the estimate is very slightly less precise than the unblocked one. At twenty pairs that is a 3% penalty against a potential fivefold gain, which is why the asymmetry makes blocking on a plausible factor an easy decision.

The same 40 units, arranged two ways. Both designs estimate the same effect of 0.5 and both are unbiased — 0.487 and 0.497. The blocked design's estimate has standard deviation 0.318 against 1.274, a variance ratio of 0.06 where the model predicts 0.06.
Fig. 4 And where the pairs differ enormously, which is the common case in field trials and multi-site studies: the unblocked design’s estimate is spread across a range in which the effect is invisible.

Where the blocks come from

Nothing above says how the pairs were formed, and the answer decides whether any of it applies.

The blocking factor has to be known before the outcome is measured and must not be affected by the treatment. Litter, batch, day, machine, site, ward, matched pairs by age and baseline severity — all fine, because all are fixed before treatment is allocated.

Anything measured after treatment is not a blocking factor, and using it as one produces a confident answer to a different question. This is the same rule that makes leverage a property of the x values: a quantity known before the outcome exists can be used to arrange the experiment; a quantity known only afterwards can only be adjusted for, with all the hazards that carries.

The practical guidance follows from the closed form: block on the factor that carries the most variance. Since the gain is 1 − share, a factor carrying half the variance halves the variance of the estimate, and a factor carrying a tenth is barely worth the paperwork.

The analysis has to match the design

The most common way to lose the gain is to arrange the experiment correctly and then analyse it as though it had not been.

A paired design analysed with an unpaired comparison gets the same point estimate — the difference in averages is the average of the differences when the arms are equal in size — and the wrong standard error. The unpaired formula computes the variability of the outcomes, which includes the pair-to-pair variation the design removed, so it reports the spread of the design that was not run.

The consequence is not a small inefficiency. On the numbers here it reports a standard error near 0.692 for an estimate whose true spread is 0.318, so every interval is more than twice as wide as it should be and the power calculation the study was planned on is wrong by a factor of five in units.

The direction is worth noting because it is the safe one: the mismatch makes the analysis conservative, not anti-conservative. A real effect goes undetected, which is a loss rather than a false claim. It remains a straightforward waste of an experiment that was designed properly.

Getting it right requires only that the analysis contain the blocking structure — differences within pairs, or the block as a term in the model. What makes the error common is that the design decision and the analysis decision are usually taken by different people at different times, and nothing in the dataset records that the pairs were pairs.

Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.
Fig. 5 What the wasted variance costs, in the currency an experiment is planned in: the chance of detecting an effect against the number of units per arm.

Blocking, matching and adjusting are three different things

Three operations get discussed interchangeably and only one of them is available before the data exists.

Blocking arranges the allocation so that the nuisance factor is balanced by construction. The gain is exact and the analysis is simple.

Matching pairs treated units with similar untreated ones after the fact, in an observational study where nobody allocated anything. It resembles blocking and is doing something much weaker: it balances the factors matched on, and can do nothing about the ones not measured.

Adjusting puts the nuisance factor in the model as a covariate. In a randomised experiment this recovers much of the same gain and is a reasonable substitute — with the difference that adjustment is a choice made after seeing the data, which is exactly the kind of choice that multiplies the analyses available.

The order is the point. Blocking is a decision taken when nothing is known, so it cannot be influenced by the results; adjusting is a decision taken when everything is known, so it can.

The same idea in other names

Blocking is one instance of a general move: remove a known source of variation by design rather than by arithmetic afterwards.

Paired comparisons are blocking with a block size of two, which is the case above. Crossover designs are blocking on the subject, giving every subject both treatments in sequence, which makes the block the tightest possible one at the cost of assuming no carry-over. Split-plot designs are blocking when one factor is hard to change and another is easy. Stratified sampling is the same operation in surveys, where the strata are the blocks.

In every one, the gain is the share of variance the arrangement removes, and every one of them can be worth nothing if the thing arranged on turns out to explain nothing.

Every way of splitting 16 units into two halves. All 12,870 assignments, enumerated. The spread of the standardised imbalance is exactly 2/√n = 0.500, whatever the covariate's own distribution, and 33.3% of assignments differ by more than 0.5 standard deviations. Randomisation does not deliver balance; it delivers a known distribution of imbalance.
Fig. 6 Why the unblocked design is so much wider: every one of the ways sixteen units can be split, and how far apart the two halves are on a covariate. Random assignment leaves that difference to chance.

Blocks larger than two

Everything above uses pairs, which is the clearest case and not the general one. A block is any set of units expected to be alike, and the arithmetic is the same for blocks of three, five or twelve.

The choice is a trade. Small blocks are tighter — the units inside them are more alike, so more of the nuisance variance cancels — and large blocks are more flexible, because a block has to be at least as large as the number of treatments being compared, and a block of two can only compare two things.

Comparing four treatments needs blocks of four, or an incomplete design where each block holds a subset and the comparisons are reassembled across blocks. That is where the classical design names live — balanced incomplete blocks, Latin squares, Youden squares — and all of them are solving the same problem: arrange the allocation so that every comparison of interest is made within a block rather than across them.

The quantity to keep an eye on is unchanged. Whatever the structure, the gain is the share of the variance that the blocking factor accounts for, and a design of great combinatorial elegance blocked on something irrelevant is worth what the figure above shows it is worth: nothing.

What blocking is worth, 80 units. Each point is 4,000 randomisations of 80 units. The line is 1 − share, computed from the model with no simulation in it. Blocking on a factor carrying 70% of the variance leaves 29% of it, which is the same precision as 273 unblocked units.
Fig. 7 The same relationship at twice as many units. The gain from blocking does not fade as the experiment grows — it is a property of the variance decomposition, not of the sample size.

That last point is worth separating out, because it is the opposite of most advice about small studies. Blocking is not a technique for rescuing an underpowered experiment. The proportional gain is the same at forty units and at four hundred, so a large study blocked on a factor carrying half the variance is getting the same halving that a small one would.

What it does not fix

Blocking removes variance from the estimate. It does not make the estimate mean anything more than it did, and three limits are worth stating.

It does not reduce the number of units needed to detect a small effect below what precision requires. It reduces the variance, which reduces the number needed — the fivefold gain above is exactly a fivefold reduction in the units required — but the arithmetic of how many subjects still governs, applied to the smaller variance.

It does not remove the need to randomise within blocks. A blocked design that allocates systematically — always the left plot, always the first patient — has built a second nuisance factor into the allocation, and it will not be visible.

It does not protect against factors that were not blocked on. Everything not in the blocking structure is still left to randomisation, which delivers a known distribution of imbalance rather than balance — the subject of the next essay.

What to check in a reported experiment

Three questions, all answerable from a methods section, and each of them corresponds to one of the measurements above.

Was there a blocking factor, and what was it? “Randomised” on its own describes the completely randomised design, whose estimate here is more than twice as variable. A study that blocked will say what it blocked on, because the analysis has to name it.

Does the analysis use it? A blocking factor mentioned in the design section and absent from the model is the mismatch above: the right estimate with a standard error from a different experiment.

Is the blocking factor plausibly a large share of the variance? Blocking on litter in an animal study or on centre in a multi-centre trial is almost always worth a great deal; blocking on alphabetical position is worth 1.03. The report should make it possible to tell which kind it was, and a study reporting the within- and between-block variation makes it possible directly.

None of the three requires the data, and the third has an answer in the fitted model of any study that blocked properly, since the block variance is a quantity the analysis had to estimate anyway. That quantity is the same one a hierarchical model reports as τ̂: the variation between groups, separated from the variation within them, which is the number both fields turn out to be about.

What this field is about

Four essays, and they share a shape: a decision taken before the data exists, whose consequences are computable in advance.

This one is the cheapest of them. Blocking costs nothing, its gain is exactly the share of variance the blocks carry, and the gain here is fivefold — forty units doing the work of a hundred and eighty-nine.

Randomisation is next, and the surprise there is what it does not buy: not balance, which it fails to deliver a third of the time by half a standard deviation, but a reference distribution that makes a test exact without assuming anything about the shape of the data.

Then factorial arrangement, where varying every factor at once estimates each of them from every run, and where changing one thing at a time recommends a setting it never tried. And last how many units — the one calculation in this field that everybody has heard of, and the one this site had been computing with a simulation and no closed form beside it until now.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationBlockingExperimental designRandomisationRandomised block designVariance reduction