Which weights are the inverse variances
Worth reading first: What the exactness buys · The shortest interval is the one that misses.
The construction weights each block’s difference by its effective size h_b = (1/m_A + 1/m_B)⁻¹, and the reason it is exact is that h_b is the inverse variance of that block’s difference. That is true when the two arms share a variance. When they do not, the variance of d_b is σ_A²/m_A + σ_B²/m_B, and h_b is no longer one over it.
The question this raises is the one the previous field named and did not answer: whether the exactness survives a difference of two weighted means. It does, twice, under two different conditions — and the two conditions between them cover essentially every trial that is designed on purpose, which is why the failure took some effort to construct.
Three weightings, and what each one needs
Effective-size weights, ŵ_b = h_b. Exact whenever V_b ∝ 1/h_b. That holds when the arms share a variance, at any block sizes and any allocations. It also holds when the allocation ratio is the same in every block, whatever the two variances are, because then V_b = (1 + k)(σ_A² + kσ_B²)/(k·m_b) is proportional to 1/m_b, and so is 1/h_b, since h_b = k·m_b/(1 + k)². Either condition alone is enough.
Flat weights, ŵ_b = 1. Exact whenever every block has the same two arm counts, since then every d_b has the same variance and any equal weights will do. This is the unweighted mean of the block differences, which is what somebody writes down who has not noticed that the blocks are worth different amounts.
Estimated precision weights, ŵ_b = 1/V̂_b, with V̂_b built from the two within-arm spread estimates and the block’s own counts. Consistent under everything and exact under nothing, because the decomposition needs the weights to be the constants they are only estimating.
All three are functions of contrasts alone, so all three are things a blinded rule is allowed to use. What separates them is not permission; it is which of them is the inverse variance.
How wrong h is, as a factor
The mismatch has a size, and writing it down says which designs it matters on.
With and an even split , the true inverse variance is while is . So
which is exactly 1 at , 1.25 at a variance ratio of 1.5 and 1.5 at a ratio of 2.
So on an evenly split block the weight is wrong by half the departure of the ratio from one. On a lopsided block the factor involves and as well, which is why the corner where the construction fails had to be built rather than found.
Five designs
One variance, one block size. All three weightings coincide — the h_b are all equal — and all three cover 94.95%.
One variance, block sizes that alternate between four and thirty-two. The h_b vary eightfold. The effective-size weights and the precision weights are the same thing here and cover 94.95%; flat weights cover 97.10%, four standard errors high.
Two variances, one block size. The allocation is constant, so the effective-size weights are still the inverse variances. All three cover 94.40%.
Two variances, block sizes that alternate. The allocation ratio is still constant, so the second condition still holds. Effective-size weights 95.40%, precision 95.40%, flat 95.75%.
Two variances, block sizes that alternate, and an allocation that alternates with them. Now neither condition holds. Effective-size weights cover 98.45%, flat 95.65%, precision 95.05%. At a tighter requirement of 0.35 the same corner gives 98.85%, 95.15% and 94.55%.
Why either condition is enough, in one line each
Both conditions are one line and both are worth having, because they are what an experimenter checks rather than the coverage table.
One variance. If σ_A = σ_B = σ then V_b = σ²(1/m_A + 1/m_B) = σ²/h_b, so 1/V_b = h_b/σ² and the effective-size weights are the inverse variances up to the constant σ², which weights do not care about. Nothing about the block sizes or the allocations enters, so a rule may change both freely.
One allocation ratio. If m_A/m_B = k in every block, write m_b for the block’s total. Then V_b = (1 + k)(σ_A² + kσ_B²)/(k·m_b) and h_b = k·m_b/(1 + k)², so both are proportional to m_b in opposite directions and 1/V_b ∝ h_b once more. The two variances appear only in the constant.
The pattern is the same both times: the weighting is right whenever the shape of V_b across blocks matches the shape of 1/h_b, and it is only the ratio between blocks that matters, never the level. Two different things can make those shapes agree — a common σ, or a common ratio — which is why there are two conditions rather than one, and why breaking the theorem needs both broken at once.
The corner had to be built
It is worth saying how much work that last design took, because the difficulty is the result.
An allocation ratio that changes between blocks is not something a trial does by accident. Nearly every design either allocates evenly or allocates in a fixed ratio, and both of those satisfy the second condition. Neyman’s allocation, which a blinded rule may compute, does change the ratio — but it converges, so after the first few blocks it is nearly constant and the condition is nearly satisfied.
To break it, the design here alternates between five to one and one to five in successive blocks, with a fivefold difference in the arm spreads and a block size that swings by a factor of eight. That is a trial with a run-in period favouring one arm and a switch favouring the other, on outcomes whose spread differs fivefold, with the block size changing throughout. It is a describable trial. It is not a common one.
A theorem that only fails on a design nobody runs is a theorem in better shape than one that fails on a design somebody might. That is the honest summary, and it is why this field reports the corner rather than leading with it.
The one that holds its level has no proof
The most interesting line in the table is the precision weights, and it is a mildly uncomfortable one.
They are exact nowhere. Conditional on the estimated weights, δ̂ is a weighted mean of independent normals whose variances the weights are only approximating, so the weighted sum of squares is not σ² times a χ² and the t interval is not exact. Every other estimator in this field comes with an argument; this one comes with an asymptotic.
And it is the only one at its nominal level in all five designs — 94.95%, 94.95%, 94.40%, 95.40%, 95.05% — including the corner where both exact estimators miss.
The reason is not mysterious. Where a condition holds, the estimated weights converge to the exact ones and the estimation error is a second-order effect on a quantity that is exactly right to first order. Where no condition holds, the exact estimators are exactly right about the wrong variances, and being approximately right about the right ones is better.
Exactness is a property of a weighting under a condition, not a property of a weighting. An estimator that is exact under a condition that does not hold has no claim at all, and the word exact attached to it is actively misleading — which is the practical form of this whole essay.
Which way the errors go
Both failures in the table are over-coverage, and that is worth reading rather than passing over.
Flat weights on unequal blocks cover 97.10% where 95% was claimed. Effective-size weights in the corner cover 98.45%. Neither undercovers anywhere in this field.
The mechanism is the same in both cases: a weighted sum of squares with the wrong weights has a distribution that is a mixture of χ² terms rather than a scaled χ², and such a mixture is more spread out than the scaled χ² with the same mean. A more spread-out variance estimate makes the t interval wider more often than narrower, so coverage rises.
That is reassuring and it is not free. An interval covering 98.45% at a nominal 95% is an interval that is systematically too wide, on a trial that was run specifically to produce a stated width, and the width is the thing being promised. A fixed-width procedure that quietly delivers a wider interval than it promised has failed at its one job, even though the coverage statement it makes is true.
So the reading is not the failure is benign. It is that the failure shows up in the width rather than in the coverage, and a table of coverages is the wrong place to look for it.
What the corner costs in blocks, which is where it comes from
The corner’s failure has a size and the size is worth tracing, because it says the failure is not about the two variances at all.
The alternating design runs 22.6 blocks at a requirement of 0.5 and 45.3 at 0.35, and the over-coverage is 98.45% and 98.85% — larger at the tighter requirement, which is the opposite of what a small-sample artefact does. So this is not a t-interval with too few degrees of freedom; it is a variance estimate whose distribution is wrong in a way that does not average out.
What is wrong is a mixture. With weights ŵ_b and true variances V_b, the weighted sum of squares is Σ (ŵ_b V_b) χ²_1 rather than a common multiple of a , and a mixture of scaled χ² terms with unequal scales has more spread than a single scaled χ² with the same mean. More spread in the variance estimate means more very wide intervals, and coverage is dominated by them.
The size of the mismatch is the spread of ŵ_b V_b across blocks. In the corner the block sizes swing by eight and the allocation swings between five to one and one to five, so ŵ_b V_b alternates by a factor that stays the same however many blocks there are. A mismatch that alternates does not average away, which is why more blocks make it worse rather than better, and why the tighter requirement — which buys more blocks — is the one that misses further.
The practical form: the failure is a design property, visible from the block sizes and allocations alone, before any data. An experimenter can compute ŵ_b V_b for their planned design up to the unknown variances and see whether it is constant. If it is, the weighting is exact; if it alternates, it is not, and no sample size repairs it.
What an experimenter has to check
Two questions, both answerable before the trial and neither requiring any of this machinery.
Will the allocation ratio be the same in every block? If yes, use the effective-size weights and stop reading. This covers even, fixed-ratio, and any design where the ratio is set once.
Will the two arms have the same spread? If yes, use the effective-size weights whatever the block sizes and allocations do.
If both answers are no — a changing allocation and different spreads — the estimated precision weights are what the evidence here supports, with the caveat that they have no exactness argument and the coverage above is a measurement of five designs rather than a theorem.
And in no case use flat weights unless the block sizes are constant, which the third column of that table makes into a two-thousand-run refusal.
The same question in the one-mean field, where it does not arise
It is worth being explicit about what the one-mean construction was quietly getting for free, because it is a good illustration of a condition that is invisible until something breaks it.
There, the block mean ȳ_b has variance σ²/m_b, so the weights are the block sizes and there is exactly one σ in the problem. There is no second variance to be unequal, no allocation to be unbalanced, and therefore no condition to check: the weights are the inverse variances always, at every schedule, and the field’s whole discussion of schedules never has to mention weighting at all.
Two arms introduce a second variance and a second dial, and the two of them together are what create the possibility of a mismatch. So the one-mean field is not a special case of this one in which the corner happens not to occur; it is a case in which the corner cannot occur, because it needs two things the one-mean problem does not have.
That is the useful generalisation to carry out of this field: whenever a construction weights independent pieces, ask what makes those weights the inverse variances, and count how many things could make them not be. In one arm the answer is nothing. In two it is two, and they have to fail together.
What to do about a weighting with no theorem
The recommendation to use an estimator that is exact nowhere deserves more than one sentence, because it is the opposite of what the rest of this fleet argues for.
The case against it is real. Its warrant is five designs at two requirements, which is a measurement of where it has been tried, not a statement about where it works. An asymptotic argument says the estimation error in the weights is second order, which is true and says nothing about how large second order is at ten blocks. And a procedure whose validity rests on simulations of the designs somebody thought to run is exactly the kind of procedure that fails on the design somebody did not.
The case for it is that the alternatives in the corner are worse and are worse while claiming to be exact. An interval covering 98.45% under a proof that does not apply is not safer than one covering 95.05% under an asymptotic; it is wrong in a way that the word attached to it conceals.
So the honest recommendation is three-part and none of it is use the precision weights. It is: check the two conditions, which is arithmetic on the planned design; design so that one of them holds, which costs nothing at all since a constant allocation ratio is what nearly every trial has; and only if neither can be arranged, use the estimated weights and report them as the approximation they are.
Designing so that a theorem applies is almost always cheaper than finding a method that works without it, and in this case it is free.
What is claimed here, and what is not
This essay takes which weighting of the block differences is exact, and the claims are the two conditions, the five designs, the corner where neither holds and both exact estimators leave their level, and the estimated weights holding theirs everywhere.
What stays out and is named as a decision: any exactness argument for the estimated precision weights, which is not available and is not claimed; the width of the intervals in the corner, which is asserted above from the mechanism and is not separately measured; and the degrees of freedom, which come out one per block short of what the one-mean identity predicts and are the next essay.
The boundary against the first essay of this field is that this one asks when h_b is the inverse variance rather than assuming it. The construction, the factor of four, and Neyman’s allocation are established there.
The checks, and the refusals that make them mean something
Two claims are gated in this field’s library. The effective-size weighting is required to be at its level in each of the four designs where one of its two conditions holds, and off it by more than three standard errors in the corner where neither does — which fails if the corner were being described as a matter of degree rather than a boundary. And the estimated precision weights are required to hold their level in all five, which is the finding that the estimator with no theorem is the one to reach for when no theorem applies.
The refusal is flat weights on blocks worth different amounts: with a block size that alternates by a factor of eight, the unweighted mean of the block differences covers 97.10% against the weighted estimate’s 94.95%, and the check refuses it. The weights that make the decomposition work are the inverse variances, and every block having the same one is an assumption rather than a fact about blocks.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A block size that changes — both name blinding, confidence interval, coverage, degrees of freedom, fixed-width interval, weighted least squares
- A ratio that changes between blocks — both name blinding, coverage, degrees of freedom, fixed-width interval, two-sample, weighted least squares
- The condition that cannot be dropped — both name blinding, block, coverage, degrees of freedom, fixed-width interval, weighted least squares
- A schedule that reads the mean — both name blinding, confidence interval, coverage, degrees of freedom, fixed-width interval
- The bias that lands in the slope — both name blinding, coverage, degrees of freedom, fixed-width interval, weighted least squares
- A width the trial has to stop for — both name blinding, coverage, fixed-width interval, weighted least squares
Named objects
A flat tag is an object no other essay names yet.
AllocationBlindingBlockChi squaredConfidence intervalContrastCoverageDegrees of freedomEffective sample sizeFixed-width intervalHarmonic meanHeteroskedasticityInverse varianceTwo-sampleWeighted least squares