The weights the corner needs

Weights that need only a ratio

A fixed-width interval about a difference is exact under either of two conditions and under neither in the corner. It is exact there too, and the only thing it needs is how much larger one arm's variance is than the other's.

Worth reading first: What the exactness buys · The shortest interval is the one that misses.

The fixed-width promise about a difference of two arms is exact, and the essay that established it found exactly where it stops. Blocks arrive, each splitting its observations between two arms; the block difference is weighted by the block’s effective size h = (1/m_A + 1/m_B)⁻¹; the weighted mean and the between-block spread are independent, so the interval is exactly t on b − 1 degrees of freedom.

The proof needs the weights to be the inverse variances, and h is the inverse variance whenever the arms share a variance or the allocation ratio is the same in every block. Both conditions cover almost every trial anybody designs on purpose. Where neither holds — two variances, block sizes that change, and an allocation that changes with them — the interval over-covers at 98.40% against a nominal 95%, and the field ends by saying that whether an exact estimator exists there is not settled.

It exists, and the missing piece is one number.

What a block is worth when the arms differ

Var(d_b) = σ_A²/m_A + σ_B²/m_B. Write λ = σ_B²/σ_A² and factor:

var(db)=σA2(1mA+λmB)=σA2hb(λ),hb(λ)=(1mA+λmB)1\operatorname{var}(d_b) = \sigma_A^2\left(\frac{1}{m_A} + \frac{\lambda}{m_B}\right) = \frac{\sigma_A^2}{h_b(\lambda)}, \qquad h_b(\lambda) = \left(\frac{1}{m_A} + \frac{\lambda}{m_B}\right)^{-1}

So the inverse variances are h_b(λ), and everything the earlier proof needs goes through with h_b(λ) where h_b stood: the weighted mean has variance σ_A²/H, the weighted spread is σ_A² times a χ² on b − 1 independent of it, and the interval is exactly t. No condition on the block sizes and none on the allocation.

And only the ratio is needed. σ_A² is a common factor of every weight, and a weighted mean does not know about a common factor, so nothing in the construction requires either variance — only how much larger one is than the other. That is the sentence the rest of this field is built on, and its consequence for a blinded rule is the next essay.

Exact in the corner, where nothing wasCoverage of a nominal 95% interval on five designs, at a required half-width of 0.3. The first four are the two-arm field's own and the fifth is its corner — two variances, block sizes that swing by eight, and an allocation that alternates between five to one and one to five — where neither of that field's two conditions holds. The effective-size weights over-cover there at 98.40%; the weights h_b(λ) = (1/m_A + λ/m_B)⁻¹ cover at 94.84%, and at 94.84% when λ is estimated from the within-arm contrasts rather than known. Nothing here is supposed to move.0.9400.9600.980coverage of a nominal 95% intervalthe levelsame-samesame-varydiff-samediff-varycornereffective sizes · equal · h_b(λ) · h_b(λ̂), left to right in each column2500 runs per design, a required half-width of 0.3the corner covers at 94.84%
Fig. 1 Coverage of a nominal 95% interval on five designs under four weightings. The fifth design is the corner. The required half-width is on a slider.

On the corner design the h(λ) interval covers at 94.84% against a standard error of 0.44 points, and at 94.80% at a second requirement. On the four designs before it, where the earlier conditions hold, it is identical to the effective-size interval — as it must be, and as it is checked to be.

The two conditions were the two ways of not needing λ

The reason the earlier field found two conditions rather than one becomes obvious once λ is in the weights, and it is not two coincidences.

h_b(1) is the effective size. So the earlier weights are correct exactly when h_b(1) is proportional to h_b(λ) across blocks — proportional, not equal, because weights matter only up to scale.

The arms share a variance. Then λ = 1 and the two are the same function.

The allocation ratio is constant. Write m_A = c·m and m_B = (1 − c)m. Then h_b(λ) = m/(1/c + λ/(1 − c)), which is proportional to m; and h_b(1) is proportional to m as well. So the two agree up to a constant that does not depend on the block, whatever λ is.

There is no third way, because those are the only two ways a function of (m_A, m_B, λ) can lose its dependence on λ up to scale. The corner is not a pathological design; it is what is left when both escape routes are closed, and closing both requires only that the arm variances differ and that the allocation ratio changes between blocks — a run-in period favouring one arm, a switch to favour the other, which is an ordinary thing to do.

Why nobody found it

The construction is three lines and it sat unfound through a whole round of work on the same object, which is worth a paragraph because the reason is instructive rather than embarrassing.

The effective size h = (1/m_A + 1/m_B)⁻¹ arrives from a derivation that assumes one variance. Under that assumption it is the inverse variance, and it has an interpretation that makes it feel like a property of the block: it is the harmonic mean of the two arm counts, halved, and it is what the block is worth. Everything about the way it is written encourages reading it as a fact about the design.

It is a fact about the design and the two variances, with the two variances set equal. Once a symbol for their ratio exists the generalisation is immediate; without one there is nothing to generalise, because the object has no slot for it. The earlier field measured the corner failing, correctly diagnosed which two conditions the weights need, and asked whether an exact estimator exists — and the answer required writing down a quantity that its own notation had no room for.

A quantity that reads as a property of the design when it is a property of the design and a parameter is the shape of this. The parameter is invisible because it has been set to one, and one is the value at which it disappears.

The factor that says when an interval’s scale is right

There is a second quantity here, it is closed, and it explains more than the weights do.

For any fixed weights a and any block variances V, the interval’s own estimate of its own scale is off by exactly

E[S2](a)var(δ^a)=(a)(aV)/(a2V)1b1\frac{\operatorname{E}[S^2]}{\left(\sum a\right)\operatorname{var}(\hat\delta_a)} = \frac{\left(\sum a\right)\left(\sum aV\right)\big/\left(\sum a^2V\right) - 1}{b - 1}

which is arithmetic on the design with no data in it. Above one the interval is too wide and over-covers; below one it is too narrow.

It equals one in two cases and no others.

a ∝ 1/V, which is h_b(λ). Substituting gives ΣaV = b and Σa²V = Σa, so the bracket is Σa·b/Σa − 1 = b − 1 and the ratio is one.

a constant. Then Σa = b·a, ΣaV = aΣV, Σa²V = a²ΣV, and the bracket is b·a·aΣV/(a²ΣV) − 1 = b − 1 again — whatever the variances are.

The second root is the one nobody writes down, and it explains something the earlier field measured without remarking on. Its flat-weight column never misbehaves. On the corner design equal weights cover at 95.16%, which is at the level; on every other design they are at the level too. They are correctly scaled everywhere, and always have been, and the reason is a root of a factor that had not been computed.

An interval's estimate of its own width, and its two roots. For any fixed weights a and any block variances V, E[S²]/(Σa·Var(δ̂)) = [(Σa)(ΣaV)/(Σa²V) − 1]/(b − 1) — arithmetic on the design with no data in it. It is exactly 1 in two cases and no others: when a ∝ 1/V, which is h_b(λ); and when every a is equal, whatever the variances are. The second is why equal weights never misbehaved in the field before this one and is nowhere written down. The effective sizes h_b(1) are neither once the arms differ and the allocation changes, and they reach 1.581 — an interval a sixth too wide, which arrives as over-coverage.
Fig. 2 The factor as the variance ratio moves, on the corner design’s block compositions. The inverse-variance weights and the equal weights sit on one exactly; the effective sizes rise to 1.58.

At λ = 25 — the corner’s ratio, a fivefold difference in arm spreads — the effective sizes give a factor of 1.5808, so the interval’s variance estimate is fifty-eight per cent too large, so its half-width is 1.257 times what it should be, so it over-covers. That is not a small misspecification producing a small over-coverage; it is a large one producing a small one, because coverage is an insensitive function of width near 95%.

There is a third weighting worth putting through the same factor, because it is the one an experimenter would reach for without thinking. Weight each block by its size, m_A + m_B. On the corner’s compositions at λ = 25 that gives a calibration of 1.3953 and an efficiency of 1.0588 — less mis-scaled than the effective sizes and more efficient than them, which is an uncomfortable result for a quantity derived from a theorem against one chosen by intuition.

It is not an accident. The block sizes here alternate between six and forty-eight while the allocation alternates between five-to-one and one-to-five, so the effective sizes h_b(1) are pulled in one direction by the size and in the other by the imbalance, and end up further from the inverse variances than the sizes alone are. The derived quantity is worse than the naive one on this design, and would be better on almost any other. That is what a weighting justified by a condition does when the condition fails: it stops carrying any guarantee at all, and where it lands is a fact about the particular design.

Calibrated is not efficient

Equal weights are calibrated everywhere and they are not the answer, because there is a second thing weights do.

Var(δ̂_a)/Var(δ̂_opt) = (Σa²V/(Σa)²)·Σ(1/V), which is Cauchy–Schwarz and is one exactly when a ∝ 1/V. On the corner design’s compositions at λ = 25 it is 4.7647 for equal weights: the interval is 2.18 times wider than it needs to be, and the run that produced it spent more than four times the observations it needed to.

So the corner leaves three positions rather than a ranking:

Equal weights. Assume nothing about the ratio. Calibrated at every λ. 118% wider.

h_b(λ). Assume the ratio. Calibrated and efficient. The narrowest interval available.

The effective sizes. Assume one of two conditions that do not hold. Mis-scaled and inefficient — a factor of 1.5808 in calibration and 1.2353 in efficiency, which multiply to a width 1.397 times the best.

Only the first two are a trade. The third is not on the frontier at all, and it is the one the earlier field’s own recommendation lands on.

The widths are the arithmetic. How wide each weighting's interval is against the best available, counted over 2500 runs of the corner design and computed from the block compositions and the variance ratio alone. The closed factor is √(calibration × efficiency): the first says how wrong the interval's own scale estimate is and the second how much variance the weights waste. The two routes agree to 0.0040, and the runs were told neither number. Equal weights are correctly scaled and 118% wider; the effective sizes are 40% wider and mis-scaled.
Fig. 3 Each weighting’s half-width against the best available, counted over runs and computed from the block compositions. The closed factor is the square root of calibration times efficiency.

The calibration factor predicts the over-coverage

The closed factor is checked above against the widths, and it can be checked against the coverage column as well — which is the one number in the earlier field that started all of this.

A calibration of 1.5808 means the interval’s variance estimate is that much too large, so its half-width is √1.5808 = 1.257 times what it should be, so a nominal 95% interval covers about 2Φ(1.96 × 1.257) − 1 = 98.6%. The earlier field counts 98.40%, and the remaining fifth of a point is the t reference rather than a normal one on the small number of blocks these designs carry.

So the whole of the corner’s failure — the number that made the earlier field say the question was unsettled — comes out of an expression with no data in it, computed from the block compositions and one variance ratio. The over-coverage was never a property of the sample.

The same expression predicts a quantity nobody has counted. Weighting by block size gives a calibration of 1.3953, so a half-width factor of 1.181 and a coverage near 97.9%: still over-covering, by two points less than the effective sizes.

And the expression says the failure does not wash out. Its denominator is b − 1 and its numerator grows in proportion to b when the block compositions repeat, so the ratio tends to a constant rather than to one as blocks are added. A design that alternates two compositions is mis-calibrated by the same factor with a hundred blocks as with ten, which is worth knowing because “collect more blocks” is the reflex repair for a coverage that is off, and it is the one repair that does nothing here.

The derived weighting is dominated by the guessed one

The essay calls the size weighting’s showing an uncomfortable result and stops at two numbers. Put them through the identity and the discomfort is sharper than that.

Weighting by block size gives a width ratio of √(1.3953 × 1.0588) = 1.216 against the effective sizes’ 1.397. So on the corner design the naive weighting is not merely competitive: it is closer to calibrated — 1.3953 against 1.5808 — and more efficient — 1.0588 against 1.2353 — and therefore narrower on both counts at once.

The effective sizes are dominated, in the strict sense: there is no axis of this comparison on which they win. That is a stronger statement than “worse on this design”, because a dominated option cannot be recovered by re-weighting the two objectives or by deciding that calibration matters more than efficiency. Nothing anybody could want out of the pair points at it.

What survives is the four-position ladder rather than the three the essay ends on: h_b(λ) at a width of 1.000 and exact, needing the ratio; block size at 1.216, needing nothing and over-covering by about two points; the effective sizes at 1.397, needing a condition that fails and over-covering by three and a half; and equal weights at 2.183, needing nothing and exactly calibrated at every λ.

Read that way the choice is between the first and the last — assume the ratio and be exact, or assume nothing and pay 118% of width — with the two middle rows being what happens when a weighting is justified by a condition instead of by an identity. One of them was guessed and one was derived, and the derived one is behind.

What over-coverage is usually worth, and is not here

An interval that covers at 98% when it claims 95% is normally read as conservative — safe, and paid for in width. The reading is right and the accounting is usually left there, because the alternative is assumed to be an interval that covers at 95% and is narrower by exactly what the conservatism cost.

That is not the alternative here. The effective-size interval is over-covering and inefficient, and the two are separate defects with separate causes. The over-coverage comes from the calibration factor being 1.5808 — its own estimate of its own scale is too large. The inefficiency comes from the weights not being the inverse variances — a factor of 1.2353 in variance regardless of how the scale is estimated. Fixing one would not fix the other.

So the width penalty against the exact interval is 1.397 and the coverage is 3.4 points too high, and there is no sense in which the second is bought with the first. An experimenter accepting the over-coverage as the price of robustness is paying a price and not buying robustness: equal weights would give them the robustness — calibrated at any λ, assuming nothing — at a width penalty of 2.18, and h_b(λ) would give them an exact interval at a penalty of one.

Two routes to every width

The last claim is the one that makes the rest checkable rather than merely consistent.

A weighting’s half-width relative to the best available is √(calibration × efficiency) — the first factor being how wrong the interval’s own scale estimate is, the second how much variance the weights waste. Both come from the block compositions and the variance ratio, with no data anywhere in them.

Counted over twenty-five hundred runs of the corner design, the effective sizes give a width ratio of 1.3972 and the closed form gives 1.3974. Equal weights give 2.1788 counted and 2.1828 closed. The runs were never told either factor, and a mistake in what h_b(λ) is — a λ where a √λ should be, a harmonic mean where an arithmetic one should be — would show in the third decimal place of both.

An identity that reproduces four significant figures of a simulation it has no contact with is a different kind of evidence from a coverage rate near its level, and it is the reason the numbers in this field can be quoted at all.

The asymmetry between the two kinds is worth one more sentence. A coverage rate at 94.84% against a nominal 95% with a standard error of 0.44 points is consistent with exactness and consistent with a great many other things — a construction that covered at 94.6% would look identical on that evidence. A width ratio agreeing to four figures is consistent with almost nothing else.

What is claimed here, and what is not

This essay takes an exact interval for a difference where neither of the earlier conditions holds, and the claims are that the inverse variances are h_b(λ) and depend on the variance ratio alone, that the two earlier conditions are exactly the two ways h_b(1) can be proportional to h_b(λ), that the interval covers at 94.84% on the corner design at two requirements, that the calibration factor has exactly two roots — inverse-variance weights and equal ones — and that each weighting’s width against the best is √(calibration × efficiency) to four significant figures.

What stays out and is named as a decision: more than two arms, where the block difference becomes a vector of contrasts and the weights become a matrix; anything about the effect being unequal between blocks, which is a separate condition the earlier field measures; and the whole question of where λ comes from, which is the next essay.

The boundary against the field this one continues is that it measures the corner failing and this one supplies the weight that does not fail there. Everything it establishes about the two conditions is unchanged; what changes is that they turn out to be a special case.

The checks, and the refusals that make them mean something

Three claims are gated in this field’s library. h_b(1) is required to be the effective size and h_b(λ) proportional to it at a constant allocation ratio, which is the identity that makes this a generalisation. The interval is required to cover at its level on all five designs at two requirements. And the counted widths are required to reproduce √(calibration × efficiency), which is two routes to one number sharing no arithmetic.

The refusal for this essay is the effective sizes used in the corner and called conservative. They over-cover, which is the failure that gets forgiven — and they are wider than the interval that covers exactly, by a factor of 1.397. An over-covering interval is usually the price of not knowing something; here it is the price of using the wrong weights, and it is paid twice.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Allocation ratioBlockBlockingCalibrationContrastCoverageDegrees of freedomEffective sample sizeEfficiencyFixed-width intervalHarmonic meanInverse variance weightingNuisance parameterVariance ratioWeighted least squares