Concept

Blocking — where it appears

Grouping units expected to be alike and comparing only within the groups, so that whatever differs between groups leaves the comparison. What it removes is the between-group variance exactly, and what it costs is the degrees of freedom the group means absorb.

Named by 22 essays across 11 fields — each of them below, with the objects they name alongside it.

One experiment, with the blocks getting smaller as the target comes into range. A single run at a requirement of 0.25, with the block sizes 5, 5, 11, 25, 11, 8, 3, 2 and a total of 70 observations in 8 blocks. The rule stops when the observations in hand reach z²σ̂²/d², with σ̂² pooled from the within-block contrasts — an estimate that moves as the run goes on, so the target moves too. Early blocks are large because the target is far away and cannot be overshot; late ones are small because a block is the granularity of the answer. The interval afterwards is built from the 8 block means and from nothing the rule looked at, and it has 7 degrees of freedom against the rule's 62.

A block size that changes

The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.

pace · Stopping
What a median split can see. A standard normal covariate with its median marked, and the mean of each category as a vertical rule: -0.7979, 0.7979. A rule that balances the categories is balancing those numbers and nothing else, so the part of the covariate it can act on is the variance between them — 0.6366 of the total, which at two categories is exactly 2/π because the two half-normal means are ±√(2/π). The rest, 0.3634, is variation inside the categories that the rule cannot see and does not touch: the assignment within a category is still a coin. Everything the next figure measures is a consequence of this one, and it is available before any unit has arrived.

A covariate with no levels

Every balancing rule on this site reads a level. Age and blood pressure have none, so somebody cuts them into categories — and a median split can see exactly 2/π of a normal covariate, whatever the rule does with the halves.

continuous · Assignment
Two covariates make the dictionary an outer product. Four functions of each covariate, and everything a balancing rule may be handed. The margins are the 8 main effects and the block between them is the 16 interactions, which are 66.7% of the dictionary. Every inner product in it is closed form — ⟨f₁g₁, f₂g₂⟩ = ⟨f₁,f₂⟩⟨g₁,g₂⟩ when the covariates are independent — so nothing about the geometry gets harder. What gets harder is the counting: choosing k of 24 is C(24, k), which is 10,626 at four and 735,471 at eight.

A dictionary that is a product

Two covariates make what a balancing rule may read an outer product — eight main effects and sixteen interactions — and every inner product in it is still closed form. What a rule holding all eight main effects removes of a pure interaction is not small. It is zero.

product · Blocking
The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%.

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

stop · Stopping
The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 258 of the 999 re-randomisations reach it, and the p-value is (1 + 258)/(1 + 999) = 0.2590. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 2.101.

The experiments that could have happened

An adaptive trial's allocation is a function of the outcomes it will later be compared against, so the ordinary analysis rejects a true null 9.2% of the time. Hold the outcomes fixed, re-run the rule that assigned them, and count — the same statistic against a reference distribution the trial could actually have drawn from is back at 4.0%.

exact · Reference
The same 40 units, arranged two ways. Both designs estimate the same effect of 0.5 and both are unbiased — 0.488 and 0.497. The blocked design's estimate has standard deviation 0.318 against 0.692, a variance ratio of 0.21 where the model predicts 0.20.

The variance removed before the data

Arranging forty units in pairs rather than assigning them at random cuts the variance of the estimated effect to a fifth — and the fifth is knowable in advance, because it is exactly the share of the variance the pairs do not carry.

design · Blocking
Exact in the corner, where nothing was. Coverage of a nominal 95% interval on five designs, at a required half-width of 0.3. The first four are the two-arm field's own and the fifth is its corner — two variances, block sizes that swing by eight, and an allocation that alternates between five to one and one to five — where neither of that field's two conditions holds. The effective-size weights over-cover there at 98.40%; the weights h_b(λ) = (1/m_A + λ/m_B)⁻¹ cover at 94.84%, and at 94.84% when λ is estimated from the within-arm contrasts rather than known. Nothing here is supposed to move.

Weights that need only a ratio

A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.

corner · Nuisance
How often a randomised trial reports the reversal, advantage 6 points. Simple randomisation against randomisation stratified by group, 4,000 trials at each size. The simple design reverses on 3.40% of trials at its worst size and 0.50% at 1280 units; the stratified design reverses on none of them, at any size.

The reversal a coin cannot prevent

Randomisation removes Simpson's reversal in expectation, which is not the same as removing it. A correctly randomised trial of eighty units, on a population where the treatment helps in both groups, reports it losing overall on 3.40% of trials — and stratifying the randomisation takes that to zero at every size.

reversal · Simpson
Every way of splitting 16 units into two halves. All 12,870 assignments, enumerated. The spread of the standardised imbalance is exactly 2/√n = 0.500, whatever the covariate's own distribution, and 33.3% of assignments differ by more than 0.5 standard deviations. Randomisation does not deliver balance; it delivers a known distribution of imbalance.

Randomisation is not balance

A third of all ways to split sixteen units leave the two halves more than half a standard deviation apart on a covariate. What randomisation delivers is not balance but a known reference distribution — and it makes a test exact with no assumption about the data's shape at all.

design · Randomisation
20 adaptive trials, 45% against 25%. Each line is one trial allocating patients one at a time by the arm's own posterior. The average final share on the better arm is 84.7%, with a standard deviation of 10.3 points across these 20 trials. The rule does not deliver a fixed advantage: it delivers one that depends on how the first few patients came out.

Randomising towards the winner

Allocating more patients to the arm that is doing better is the humane thing to want and it buys nothing statistically: at a fixed total it costs thirty points of power. And because the allocation is a function of the outcomes, the ordinary test on it rejects a true null 7.8% of the time before any time trend is applied — and 58% after one.

adaptive · Randomisation
Two promises, and no rule here keeps both. A fixed-width procedure promises two things: that the interval covers at its nominal rate, and that it is no wider than the width asked for. Over 1500 runs of the modelled weighting, a rule that stops when the interval it will report is short enough keeps the width — only 2.0% of runs come out wider than 0.34 — and covers at 91.13% against a nominal 95%. A rule that stops on a width predicted from the within-arm sums of squares covers at 94.80% and comes out wider than promised on 42.3% of runs. The two promises are in conflict because keeping the second one exactly requires conditioning on the very quantity that has to be independent of the stopping time for the first.

Stopping on the arms

The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.

stop · Width
No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.

Two degrees of freedom, one total

The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.

pace · Width
One curve is a binomial coefficient and the other is a line. The number of subsets a maximin over this dictionary would have to score, against the number the exchange algorithm actually scores. At three functions the walk is 2,024 subsets and is the honest answer; at eight it is 735,471 and the exchange algorithm has looked at 421. The warrant for the second curve is the four sizes where both exist and agree, which is a weak warrant — it says the algorithm has not yet been wrong, not that it cannot be — and it is the only one available past the point the first curve leaves the page.

Where the enumeration stops

A maximin over an eight-function dictionary is a walk over seventy subsets. Over twenty-four it is 735,471 at eight functions, and the exchange algorithm that replaces the walk scores 421. What licenses the second curve is four sizes where both exist and agree, which is a weaker warrant than it looks.

product · Optimum
A rate that does not know how large the trial is. The share of equal splits admitted by a tolerance of 1 coin-spreads on 3 functions, at six trial sizes. The first two are exact — 12,870 and 184,756 splits, walked, averaged over eight draws of the units — and the rest are sampled. From a hundred units on, the rate sits on (2Φ(1) − 1)^3 = 0.3182, which contains no n at all. The two small trials are 29.2% and 27.7% short of it, so the sixteen-unit measurement understates the rate rather than bracketing it. Meanwhile the admissible count — the rate times C(n, n/2) — goes from 2^11.5 to 2^393.7: the exhaustion a small trial runs into is a fact about small trials.

A count that has to be estimated

At sixteen units the admissible assignments can be counted by walking all 12,870 of them. At four hundred there are about 2^393.70, and the share admitted is 0.31885 against a closed form of 0.31818 that has no trial size in it at all. The exhaustion a small trial runs into is a fact about small trials.

product · Randomisation
The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

pace · Nuisance
A wrong weight costs width; a random weight costs level. Five weightings on a trial whose variance ratio drifts by a factor of 20.1 between the first block and the last, over 4000 runs. The rule that knows every λ_b covers at 95.1% and sets the width. One ratio for the whole trial is wrong for every block and costs nothing in level — 94.8% — while being 20% wider; equal weights are calibrated by an identity and 22% wider. The ratio estimated inside each block is the only rule aimed at the quantity that actually varies, and it is the only one that misses the level, at 92.0%: a weight computed from a handful of degrees of freedom is mostly noise, and noise in a weight is not a wrong weight. Modelling the drift across blocks recovers the oracle's width at 94.8%.

A ratio that changes between blocks

A wrong weight costs width and a random weight costs level. The rule aimed at the quantity that actually varies is the only one that misses its own coverage, and the rule that models it across blocks recovers the whole of what knowing it is worth.

blocks · Nuisance
What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

pace · Stopping
What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.

Balancing more than one number

The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.

continuous · Blocking
Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.

How many subjects

Sixty-four per arm for 80% power at half a standard deviation — a power figure that could only be simulated, with nothing to disagree with, until the non-central t was written. Two routes now, agreeing to within the simulation's own error.

design · Power
What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.

Two levels at once

A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.

multilevel · Levels
Where the bias lands. The drift in the log variance ratio, fitted across 12 blocks over 4000 trials. E[log λ̂_b] is log λ_b plus ψ(k_B/2) − log(k_B/2) − ψ(k_A/2) + log(k_A/2), which depends on nothing but the degrees of freedom — so the tempting sentence is that it goes into the intercept and leaves the slope alone. It does not, because the blocks alternate between allocations and the alternation is correlated with the covariate being fitted: the lopsided blocks carry 0.5383 of bias and the even ones carry none. Uncorrected the slope reads 1.5597 against a truth of 1.5, which is 8.0 standard errors. Subtracting the two digammas block by block leaves 1.4976.

The bias that lands in the slope

The bias in a log variance estimate depends on nothing but its degrees of freedom, so it goes into the intercept — unless the degrees of freedom alternate with the design, which is exactly what a block-randomised trial makes them do.

blocks · Width
What 240 observations are worth, by which question is asked. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1703 against a naive 0.0060, a design effect of 28.4 and 8.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 1.0845 against a naive 0.1200, a design effect of 9.0 and 26.6 effective observations.

Two groupings that cross

Pupils belong to a school and to a neighbourhood, and neither is nested in the other. There is then no design effect: the overall mean is worth 8.5 independent observations out of 240, a row difference 11.1 and a column difference 26.6, and which grouping matters depends on the question rather than on the study.

multilevel · Levels

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageFixed-width intervalAllocationDegrees of freedomRandomisationBlindingMonte CarloNuisance parameterSample sizeExperimental designStopping ruleWeighted least squares

All concepts