A promise about two arms

The degrees of freedom in the sums

One arm partitions N − 1 exactly. Two arms give the rule N − 2b and the interval b − 1, which is short by one per block — and the missing ones are in the block sums, which are correlated with the differences at −0.79 and are usable anyway.

Worth reading first: The variance removed before the data · The shortest interval is the one that misses.

The one-mean field’s identity is the reason its trade is a conserved one: the rule’s spread estimate has N − b degrees of freedom and the interval has b − 1, and (b − 1) + (N − b) = N − 1 on every run. Every schedule moves along that line rather than off it. A rule cannot make more of both.

Two arms do the arithmetic differently and the sum does not close. The rule pools within-arm within-block contrasts, which is Σ(m_A − 1) + (m_B − 1) = N − 2b, and the interval has b − 1. Those come to N − b − 1, and what two arms leave — after a grand mean and a treatment effect — is N − 2.

One degree of freedom per block is unaccounted for. This essay is about where it went, and the answer turns out to be usable, which is not what the correlation between the pieces suggests.

The shortfall, on every run

Two arms leave one degree of freedom per block unaccounted for. Each point is one run. The one-mean field's identity is (b − 1) + (N − b) = N − 1, and every schedule moves along that line rather than off it. Two arms give the rule N − 2b and the interval b − 1, which come to N − b − 1 — short of the N − 2 two arms leave by exactly one per block, since a block's arm counts absorb one degree of freedom each and only one of the two directions carries the difference. The hollow points add what the block sums are worth, b − 1 more, and land on the total. The missing degrees of freedom are not lost; they are in a place the interval has to be shown it may read.
Fig. 1 Each point is one run. The filled points are the two counts a two-arm rule uses; the hollow ones add what the block sums are worth, and land on the total.

The identity holds exactly and on every run, which is what makes it an identity rather than a tendency. Each block’s 2m observations carry 2m − 1 degrees of freedom about the variance after their own mean; the rule takes 2m − 2 of them, leaving one per block. Across b blocks that is b, of which one goes to the grand mean, leaving b − 1.

Those b − 1 are in the block sums. Each block yields two summaries — its difference ȳ_A − ȳ_B and its sum ȳ_A + ȳ_B — and a two-arm interval about the difference uses only the first. The sums are the other half, they are as numerous as the differences, and nobody looks at them.

The trade is a line and it is not symmetric

The identity says the two counts move along a straight line, and the t multiplier says the two ends of that line are worth very different amounts.

At two blocks the interval has one degree of freedom and the multiplier is 12.71. At three it has two and the multiplier is 4.30; at five, four and 2.78; at twelve, eleven and 2.20; at thirty, twenty-nine and 2.05.

So the first few blocks buy enormously and the rest buy almost nothing. Going from two blocks to five divides the multiplier by four and a half; going from twelve to thirty moves it by seven per cent.

The degrees of freedom trade one for one along the line; what they are worth does not.

The same variance, seen twice

The reason to look is arithmetic and it is a little surprising.

Var( ȳ_A + ȳ_B ) = σ_A²/m_A + σ_B²/m_B = Var( ȳ_A − ȳ_B )

The two expressions are identical, whatever the allocation and whatever the two spreads. So the block sums estimate exactly the quantity the interval needs, and there are b of them where there were b differences.

That is a free doubling of the interval’s degrees of freedom, and the immediate objection is obvious.

The same variance, seen twice, and strongly correlated with itself. Var(ȳ_A + ȳ_B) and Var(ȳ_A − ȳ_B) are the same expression — σ_A²/m_A + σ_B²/m_B — whatever the allocation and whatever the two spreads, so the block sums are a second, free estimate of exactly the quantity the interval needs. They are also strongly correlated with the block differences, at -0.794 even at balance once the two arms differ in spread, which is the objection everybody raises against using them. It does not bite: a variance estimate is a quadratic in the sums and the reported difference is linear in the differences, and a centred Gaussian has no third moment to couple them with. Pooling them doubles the interval's degrees of freedom for nothing.
Fig. 2 The two variances, on the plates, and the correlation between a block’s sum and its difference, on the curve. The variances agree at every allocation; the correlation is large at all of them.

The sums and the differences are strongly correlated. With spreads of 1 and 3 the correlation is −0.7938 at an even split and −0.9290 at three to one, and the closed form is σ_A²/m_A − σ_B²/m_B, which is zero only when the two variances and the two counts happen to balance each other. Using a quantity that correlated with the reported difference to build the interval about that difference looks like exactly the mistake this whole field exists to avoid.

Why the correlation does not bite

It does not bite, and the reason is a parity argument rather than a numerical accident.

The interval reports δ̂, which is linear in the block differences. The variance estimate is quadratic in the block sums, taken about their own weighted mean. For jointly normal quantities, a quadratic form in one set of centred variables and a linear function of another are uncorrelated whenever the covariance structure is constant across blocks, because the coupling term is a third moment and a centred Gaussian has none.

Working it through with weights h_b: Cov(s_b − s̄_w, δ̂) = (h_b/H)c_b − Σ_j (h_j/H)² c_j, where c_b = Cov(s_b, d_b). With a constant allocation every c_b is the same c, and the two terms are c/b and c/b. They cancel exactly.

So the block sums are independent of the reported difference at any allocation, however correlated their raw versions are, provided the allocation does not change between blocks — which is the same condition the weights need, arriving through a different argument.

The measurement agrees. Pooling the sums into the interval’s variance estimate covers 95.10% against the honest interval’s 95.15%, on an interval 2.16% narrower.

Two per cent of width for nothing. That is a small number and it is a real one: it is the whole value of the missing degrees of freedom, and it is what a two-arm rule leaves on the table by looking only at differences.

Two per cent is the right size, and here is why

A doubling of degrees of freedom sounds as though it should be worth more than two per cent of width, so the arithmetic is worth doing — because it says the number is right and not disappointing.

The interval is tb1S2/Ht_{b-1}\sqrt{S^2/H}. Doubling the degrees of freedom does not change H and does not change what S² estimates; it changes only the quantile and the spread of the estimate. At twenty-one blocks the t quantile on 20 degrees of freedom is 2.086 and on 40 it is 2.021 — a 3.1% reduction. The variance estimate also gets steadier, which pulls in the same direction by a smaller amount.

So 2.16% measured against 3.1% available from the quantile alone is about what should be expected once the pooled estimate’s own noise is accounted for. The gain would be larger with fewer blocks — at five blocks the quantiles are 2.776 and 2.306, a 17% reduction — and vanishing with many, since t converges.

That is the useful shape of it: the block sums matter most exactly where blocks are scarce, which is where a two-arm blinded interval is widest and where an experimenter most wants the width back. A trial with fifty blocks should not bother; one with six should.

What buying them costs

The free part has a condition, and the condition is not about the allocation. It is about whether the sums see what the differences see.

Free until the sums stop seeing what the differences see. Coverage with and without the block sums pooled into the interval's variance estimate. With one effect and one level they are free. With an effect that varies between blocks they are still free, because a block's sum picks that variation up exactly as its difference does. With a level that varies they make the interval 37% wider and conservative. And where the effect falls as the level rises — a ceiling, and not an exotic thing to suppose — the sums carry none of the between-block variation while the differences carry all of it, the pooled estimate is short, and the interval that uses it covers 88.75% on a width 20% narrower than the honest one.
Fig. 3 Coverage with and without the sums pooled, under four things a trial’s blocks might do. Three of them are safe and one is not.

One effect, one level. Honest 95.15%, pooled 95.10%, 2.16% narrower. Free.

An effect that differs between blocks. Honest 94.60%, pooled 94.40%, 2.69% narrower. Still free — and this is the case people expect to break it, so it is worth saying why it does not. A block whose effect is larger has a larger difference and a larger sum, because the treated arm’s mean moves and the sum contains it. Both estimates are inflated by the same amount, so pooling them is inflating a valid estimate by a valid amount.

A level that differs between blocks. Honest 94.50%, pooled 99.20%, and the interval is 37.35% wider. A period effect moves both arms together, so it cancels in the difference and arrives doubled in the sum. The pooled estimate is far too large and the interval is conservative — a real cost, since the promised quantity here is a width, but not a false statement.

An effect that falls where the level rises. Honest 94.95%, pooled 88.75%, on an interval 20.38% narrower. This is the failure.

The one configuration that costs

The last case is worth stating carefully because it is neither exotic nor obvious.

Let the block’s overall level vary — a period effect, a seasonal shift, a different recruiting site — and let the treatment’s effect fall where the level rises. A ceiling: when everybody is doing well the treatment adds little; when everybody is doing badly it adds a lot. That is one of the most commonly hypothesised patterns in applied work.

Arithmetically, with the effect regressing on the level at a slope of −2, the block sum 2·level + effect carries none of the variation while the difference carries all of it. So the sums’ estimate is of σ² alone and the differences’ estimate is of σ² plus the between-block variation. Pooling them halves the excess, the interval narrows by a fifth, and the coverage falls to 88.75%.

The block sums are a free second estimate for as long as they see what the differences see, and the one thing that makes them stop is an effect that moves against the level. Which is a hypothesis an experimenter usually has an opinion about, and which is therefore checkable in the only way that matters: by asking, before the trial, whether the effect is expected to depend on the baseline.

The three piles, and what each one is for

It helps to name the three piles rather than count them, because they are not interchangeable and the identity makes them look as though they might be.

N − 2b within-arm within-block contrasts. These belong to the stopping rule. They estimate σ² and they are what makes the whole construction exact, because they are independent of every block mean. They cannot be given to the interval — an interval built on them would be built on the quantity that decided when to stop, which is the dependence the field exists to break.

b − 1 between-block differences. These belong to the interval. They are the only pile that estimates the variance of the thing being reported without conditioning on anything the rule looked at.

b − 1 between-block sums. These belong to nobody by default. They estimate the same variance as the second pile, they are independent of the reported difference under a constant allocation, and they are the subject of this essay.

The first pile is large and unusable for the interval; the second is small and is what the interval has; the third is the same size as the second and is available under a condition. The scarcity is not in the total. It is in which pile is independent of what, and that is a fact about the construction rather than about the sample size — which is why a larger trial does not relieve it and why the ratio between the piles is set by the block size alone.

What the identity is worth knowing for

The identity itself is arithmetic and does not change any recommendation. What it does is make a question visible that a two-arm analysis does not otherwise raise.

The one-mean identity says a schedule cannot make more of both. The two-arm shortfall says something different: there is a third pile, and whether it may be used is a modelling question rather than an arithmetic one. That is a genuinely different kind of statement, and it is the sort of thing an identity is good for — not the number it produces, but the term it forces somebody to account for.

And it puts a bound on what is available. The interval can have b − 1 degrees of freedom or 2(b − 1), and no analysis of block differences can have more, because N − 2b of the total belong irreducibly to the within-arm contrasts the stopping rule is built on. A two-arm blinded interval is short of a one-arm one by exactly b degrees of freedom, and half of that is recoverable under a stated condition.

Why nobody looks at the sums

The last thing worth saying is why a quantity this available goes unused, because the reason is instructive rather than an oversight.

A block sum is not about the treatment. Its expectation is 2μ + δ, which contains the grand level, so it carries no information about δ that the difference does not already have — and an analyst reading a two-arm trial is reading it for δ. So the sums look like nuisance, and nuisance is what gets dropped.

But the interval needs two things: an estimate of δ and an estimate of its variance, and only the first has to come from a quantity that is about the treatment. The second is a variance, and a variance can be estimated from anything with the right variance, which the sums have. Nuisance for the estimate is not nuisance for the standard error, and conflating the two is what leaves the degrees of freedom on the floor.

This is the same observation that makes the whole blinded construction work, one level up. The stopping rule reads within-arm contrasts, which are nuisance for δ and are precisely what a rule may look at. The interval reads the sums, which are nuisance for δ and are precisely what a variance estimate may look at. Both are cases of a quantity being useless for the parameter and useful for the precision, and in a construction organised entirely around who may read what, that distinction is the one doing the work.

The two degrees of freedom are a partition of one total. Each point is one run of one schedule. Every point sits on a line of slope −1, because (b − 1) + (N − b) = N − 1 exactly: every observation after the first gives a degree of freedom to the rule's spread estimate or to the interval, and never to both. That identity is the whole reason there is no schedule that improves both ends of the block-size dial. The lines are the runs that ended with the same N, and a schedule moves along one of them rather than off it — what it can decide is only when in a run each degree of freedom is spent.
Fig. 4 The one-mean identity this is measured against: every run on a line of slope −1, because the two counts partition N − 1.
The same variance, seen twice, and strongly correlated with itself. Var(ȳ_A + ȳ_B) and Var(ȳ_A − ȳ_B) are the same expression — σ_A²/m_A + σ_B²/m_B — whatever the allocation and whatever the two spreads, so the block sums are a second, free estimate of exactly the quantity the interval needs. They are also strongly correlated with the block differences, at -0.794 even at balance once the two arms differ in spread, which is the objection everybody raises against using them. It does not bite: a variance estimate is a quadratic in the sums and the reported difference is linear in the differences, and a centred Gaussian has no third moment to couple them with. Pooling them doubles the interval's degrees of freedom for nothing.
Fig. 5 And the covariance the parity argument is about, at four allocations — large everywhere, and irrelevant everywhere the allocation is constant.

What a period effect actually does, and why it is the safe failure

The third row of that table is worth one more look, because it is the case most trials are actually in and because its failure runs the safe way.

A block-level shift in the overall level — a different month, a different site, a different batch of reagent — is what blocking exists for. It cancels in the difference exactly, which is the whole point of comparing within a block, and it arrives in the sum with a coefficient of two.

So a trial with period effects has block sums whose variance is genuinely larger than its block differences’ variance, and the two are no longer estimating the same thing. Pooling them produces a variance estimate that is too large, an interval 37.35% wider than the honest one, and coverage of 99.20%.

That is a real cost paid in the promised quantity — a fixed-width procedure delivering a third more width than it promised has not done the job — and it is not a false statement about coverage. Which is the asymmetry worth carrying: the block sums can only make the interval too wide, except in the one configuration where the effect moves against the level. Every other way the sums can be inflated inflates the variance estimate and is conservative.

So the decision rule is narrow. Do not pool the sums if the effect is expected to depend on the baseline. If period effects are expected but the effect is not expected to track them, pooling costs width rather than validity, and the honest interval is the better choice for a different reason than the one this essay is about.

What is claimed here, and what is not

This essay takes where the two-arm shortfall in degrees of freedom lives, and the claims are the identity holding on every run, the block sums having the same variance as the block differences at every allocation, the parity argument that makes them usable despite a correlation of −0.93, and the one configuration in which using them costs six points of coverage.

What stays out and is named as a decision: any repair for that configuration, which would need the between-block variation modelled rather than pooled and is a random-effects analysis rather than an exact one; the case of an allocation that changes between blocks, where the parity argument’s cancellation fails and which is not separately measured; and the width comparison in the pooled version, which is reported and is not the promised quantity the field is built around.

The boundary against the one-mean field is the second arm. What the second arm costs the promise itself is the field’s first essay. The identity for one mean, what each pile of degrees of freedom buys, and why a schedule moves along the line rather than off it are established there.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The two counts are required to come to N − b − 1 on every run and the three to come to N − 2, exactly and with no exception, which is what makes this an identity. And the block sums are required to have the same variance as the block differences at four allocations with unequal arm spreads, while being correlated with them above 0.5 at every one — which is the pair of facts the parity argument reconciles and which either alone would misrepresent.

The refusal is pooling the sums where the effect falls as the level rises: 88.75% against the honest interval’s 94.95%, on a width 20.4% narrower. The check refuses the pooling rather than the construction, because the construction is correct and the extra degrees of freedom are real — they are simply not degrees of freedom about the same quantity any more.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlindingBlockChi squaredConfidence intervalContrastCorrelationCoverageDegrees of freedomFixed-width intervalIndependenceOrthogonalityPeriod effectRandom-effectsTwo-sampleVariance