The variance removed before the data
Forty units, twenty to be treated and twenty not, and the effect estimated as the difference between the two averages. The units are not identical: they arrive in pairs that share something — two plots in the same field, two patients at the same hospital, two measurements taken on the same day — and whatever they share affects the outcome.
There are two ways to allocate the treatment, both correct, both unbiased, and they differ in how much the answer moves from experiment to experiment.
The two arrangements
Complete randomisation ignores the pairing. Choose twenty of the forty units at random, treat them, and compare the averages. It is what “randomised” means without further qualification, and it may put both members of a pair in the same arm, or neither.
Blocking uses the pairing. Within each of the twenty pairs, treat one member and not the other, chosen at random. Every pair contributes one treated and one control unit, and the effect is the average of the twenty within-pair differences.
Both are randomised. Both are unbiased — measured across four thousand randomisations, the completely randomised design estimates the true effect of 0.5 and so does the blocked one. Nothing about the allocation biases anything, and any argument for blocking that rests on “balance” is arguing for the wrong thing.
What differs is the spread. The blocked estimate has a standard deviation of 0.318; the completely randomised one has 0.692. The variance ratio is 0.212.
Why the difference is exact rather than approximate
The reason the blocked design is tighter is worth stating as arithmetic rather than as intuition, because the arithmetic gives the size of the gain in advance.
Write each unit’s outcome as its pair’s shared value, plus the unit’s own noise, plus the treatment effect if it was treated. In the blocked design the estimate is an average of within-pair differences, and in every one of those differences the shared value appears twice with opposite signs. It cancels. Not approximately, not on average: it is not in the estimate at all.
So the blocked estimate’s variance contains only the unit noise, while the completely randomised estimate’s variance contains the unit noise and the pair-to-pair variation, because a random split of forty units into two halves does not balance the pairs.
The prediction that follows is that the variance ratio is
σ² / (σ² + β²)
where β is the spread between pairs and σ the spread within them — the share of the total variance the blocks do not carry. With σ = 1 and β = 2, that is 1/5 = 0.200, against a measured 0.212. The agreement is the two-routes rule applied to an experimental design rather than to a distribution: a count over four thousand randomisations and a closed form derived from the model, sharing no arithmetic.
What the gain is worth in units
A variance ratio is not the currency an experiment is planned in. The currency is units, and the conversion is direct: a design with a fifth of the variance has the precision of an experiment five times the size.
Forty blocked units here buy the precision of 189 unblocked units. That is the number worth carrying, because it is what blocking is competing with: any argument that the pairing is inconvenient has to be weighed against nearly a hundred and fifty extra units.
It is also why blocking is the cheapest thing in this field. It costs no observations, no additional measurement, and no assumption that the completely randomised design does not also make. It is a change in the order in which the same units are allocated.
The variance ratio is one minus the pair correlation
The 0.212 is not a number that had to be simulated, and writing down what it equals turns the whole field into one parameter.
Let each unit’s outcome be its pair’s value, with variance , plus its own noise, with variance . The blocked estimate averages twenty within-pair differences, each with variance , so its variance is . The completely randomised estimate compares two means of twenty drawn from units whose variance is , so its variance is .
The ratio is
where ρ is the correlation between the two members of a pair. The gain from blocking is entirely the intraclass correlation and nothing else — not the number of pairs, not the size of the effect, not the noise level, only what share of the variation is between pairs rather than within them.
The figures here are drawn at a pair spread of 2 against a unit noise of 1, so and the ratio should be 0.200. Measured, 0.212, the difference being the finite-population correction that comes from splitting exactly forty units rather than sampling them.
Which prices the design in units
A variance ratio converts into sample size directly, since variance goes as 1/n: blocking multiplies the effective sample by .
Here that is five. Forty units blocked into twenty pairs carry the precision of about 189 completely randomised units, measured, or 200 by the closed form.
And the formula says where that stops being worth having. At a pair correlation of 0.5 the multiplier is 2 — blocking is worth a doubling of the trial. At 0.2 it is 1.25, which on forty units is the difference between forty and fifty. At 0.05 it is 1.05, and the pairing is worth two units.
So the question to ask of any proposed blocking factor is not whether the units within a block are similar; it is how much of the outcome’s variance lies between blocks, because that number is the answer. A blocking variable explaining four fifths of the variation quintuples the experiment. One explaining a twentieth does nothing at all, and costs the analysis a constraint it now has to respect.
The case where it is worth nothing
A method whose advantage has been demonstrated once has been demonstrated against one configuration, and this site’s habit is to show the same machinery a case it must refuse.
Blocking on a factor that shares nothing is that case. When the two members of a pair have no common value — when the pairing is arbitrary — the variance ratio measured the same way is 1.03: no gain, and a whisker of loss.
This matters more than as a check on the arithmetic. It says the gain is not produced by blocking itself, but by blocking on something relevant, and the phrase “the study was blocked” carries no information without saying what it was blocked on and how much of the variation that factor accounts for.
The small loss is real and is worth naming honestly. A blocked analysis spends a degree of freedom per block, so where the blocks carry nothing the estimate is very slightly less precise than the unblocked one. At twenty pairs that is a 3% penalty against a potential fivefold gain, which is why the asymmetry makes blocking on a plausible factor an easy decision.
Where the blocks come from
Nothing above says how the pairs were formed, and the answer decides whether any of it applies.
The blocking factor has to be known before the outcome is measured and must not be affected by the treatment. Litter, batch, day, machine, site, ward, matched pairs by age and baseline severity — all fine, because all are fixed before treatment is allocated.
Anything measured after treatment is not a blocking factor, and using it as one produces a confident answer to a different question. This is the same rule that makes leverage a property of the x values: a quantity known before the outcome exists can be used to arrange the experiment; a quantity known only afterwards can only be adjusted for, with all the hazards that carries.
The practical guidance follows from the closed form: block on the factor that carries the most variance. Since the gain is 1 − share, a factor carrying half the variance halves the variance of the estimate, and a factor carrying a tenth is barely worth the paperwork.
The analysis has to match the design
The most common way to lose the gain is to arrange the experiment correctly and then analyse it as though it had not been.
A paired design analysed with an unpaired comparison gets the same point estimate — the difference in averages is the average of the differences when the arms are equal in size — and the wrong standard error. The unpaired formula computes the variability of the outcomes, which includes the pair-to-pair variation the design removed, so it reports the spread of the design that was not run.
The consequence is not a small inefficiency. On the numbers here it reports a standard error near 0.692 for an estimate whose true spread is 0.318, so every interval is more than twice as wide as it should be and the power calculation the study was planned on is wrong by a factor of five in units.
The direction is worth noting because it is the safe one: the mismatch makes the analysis conservative, not anti-conservative. A real effect goes undetected, which is a loss rather than a false claim. It remains a straightforward waste of an experiment that was designed properly.
Getting it right requires only that the analysis contain the blocking structure — differences within pairs, or the block as a term in the model. What makes the error common is that the design decision and the analysis decision are usually taken by different people at different times, and nothing in the dataset records that the pairs were pairs.
Blocking, matching and adjusting are three different things
Three operations get discussed interchangeably and only one of them is available before the data exists.
Blocking arranges the allocation so that the nuisance factor is balanced by construction. The gain is exact and the analysis is simple.
Matching pairs treated units with similar untreated ones after the fact, in an observational study where nobody allocated anything. It resembles blocking and is doing something much weaker: it balances the factors matched on, and can do nothing about the ones not measured.
Adjusting puts the nuisance factor in the model as a covariate. In a randomised experiment this recovers much of the same gain and is a reasonable substitute — with the difference that adjustment is a choice made after seeing the data, which is exactly the kind of choice that multiplies the analyses available.
The order is the point. Blocking is a decision taken when nothing is known, so it cannot be influenced by the results; adjusting is a decision taken when everything is known, so it can.
The same idea in other names
Blocking is one instance of a general move: remove a known source of variation by design rather than by arithmetic afterwards.
Paired comparisons are blocking with a block size of two, which is the case above. Crossover designs are blocking on the subject, giving every subject both treatments in sequence, which makes the block the tightest possible one at the cost of assuming no carry-over. Split-plot designs are blocking when one factor is hard to change and another is easy. Stratified sampling is the same operation in surveys, where the strata are the blocks.
In every one, the gain is the share of variance the arrangement removes, and every one of them can be worth nothing if the thing arranged on turns out to explain nothing.
Blocks larger than two
Everything above uses pairs, which is the clearest case and not the general one. A block is any set of units expected to be alike, and the arithmetic is the same for blocks of three, five or twelve.
The choice is a trade. Small blocks are tighter — the units inside them are more alike, so more of the nuisance variance cancels — and large blocks are more flexible, because a block has to be at least as large as the number of treatments being compared, and a block of two can only compare two things.
Comparing four treatments needs blocks of four, or an incomplete design where each block holds a subset and the comparisons are reassembled across blocks. That is where the classical design names live — balanced incomplete blocks, Latin squares, Youden squares — and all of them are solving the same problem: arrange the allocation so that every comparison of interest is made within a block rather than across them.
The quantity to keep an eye on is unchanged. Whatever the structure, the gain is the share of the variance that the blocking factor accounts for, and a design of great combinatorial elegance blocked on something irrelevant is worth what the figure above shows it is worth: nothing.
That last point is worth separating out, because it is the opposite of most advice about small studies. Blocking is not a technique for rescuing an underpowered experiment. The proportional gain is the same at forty units and at four hundred, so a large study blocked on a factor carrying half the variance is getting the same halving that a small one would.
What it does not fix
Blocking removes variance from the estimate. It does not make the estimate mean anything more than it did, and three limits are worth stating.
It does not reduce the number of units needed to detect a small effect below what precision requires. It reduces the variance, which reduces the number needed — the fivefold gain above is exactly a fivefold reduction in the units required — but the arithmetic of how many subjects still governs, applied to the smaller variance.
It does not remove the need to randomise within blocks. A blocked design that allocates systematically — always the left plot, always the first patient — has built a second nuisance factor into the allocation, and it will not be visible.
It does not protect against factors that were not blocked on. Everything not in the blocking structure is still left to randomisation, which delivers a known distribution of imbalance rather than balance — the subject of the next essay.
What to check in a reported experiment
Three questions, all answerable from a methods section, and each of them corresponds to one of the measurements above.
Was there a blocking factor, and what was it? “Randomised” on its own describes the completely randomised design, whose estimate here is more than twice as variable. A study that blocked will say what it blocked on, because the analysis has to name it.
Does the analysis use it? A blocking factor mentioned in the design section and absent from the model is the mismatch above: the right estimate with a standard error from a different experiment.
Is the blocking factor plausibly a large share of the variance? Blocking on litter in an animal study or on centre in a multi-centre trial is almost always worth a great deal; blocking on alphabetical position is worth 1.03. The report should make it possible to tell which kind it was, and a study reporting the within- and between-block variation makes it possible directly.
None of the three requires the data, and the third has an answer in the fitted model of any study that blocked properly, since the block variance is a quantity the analysis had to estimate anyway. That quantity is the same one a hierarchical model reports as τ̂: the variation between groups, separated from the variation within them, which is the number both fields turn out to be about.
What this field is about
Four essays, and they share a shape: a decision taken before the data exists, whose consequences are computable in advance.
This one is the cheapest of them. Blocking costs nothing, its gain is exactly the share of variance the blocks carry, and the gain here is fivefold — forty units doing the work of a hundred and eighty-nine.
Randomisation is next, and the surprise there is what it does not buy: not balance, which it fails to deliver a third of the time by half a standard deviation, but a reference distribution that makes a test exact without assuming anything about the shape of the data.
Then factorial arrangement, where varying every factor at once estimates each of them from every run, and where changing one thing at a time recommends a setting it never tried. And last how many units — the one calculation in this field that everybody has heard of, and the one this site had been computing with a simulation and no closed form beside it until now.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
- A basis is a subspace
- A covariate with no levels
- A cut is not a polynomial, and it does not have to be
- A dictionary that is a product
- A dictionary that is neither
- A threshold in the tail
- A zero that was an assumption
- Balanced on the wrong function
- Balancing more than one number
- Balancing what is known in advance
- How many subjects
- Not half and half
- One factor at a time
- Randomisation is not balance
- The arcsine that closes it, and the error that was overstated
- The degrees of freedom in the sums
- The design that cannot see a curve
- The fourth moment that was missing
- The zero that survives a cut
- Three functions of one number
- What the extra function buys
- Where the enumeration stops
- Which shapes are worth protecting
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A count that has to be estimated — both name allocation, blocking, randomisation
- A covariate with no levels — both name blocking, experimental design, randomisation
- A dictionary that is a product — both name allocation, blocking, randomisation
- Allocating on a guess — both name allocation, experimental design, variance reduction
- Guessing one arm in three — both name allocation, experimental design, randomisation
- Not half and half — both name allocation, experimental design, variance reduction
Named objects
A flat tag is an object no other essay names yet.
AllocationBlockingExperimental designRandomisationRandomised block designVariance reduction