The weights the corner needs

The condition that cannot be dropped

The weights may not read the block they weight. Estimate the variance ratio inside each block rather than across the trial and the coverage falls to 83% — on an interval that is at the same time seventy per cent wider.

Worth reading first: Not half and half · The shortest interval is the one that misses.

The exact interval needs weights that are the inverse variances, and the inverse variances need one parameter. Everything about the construction so far says that parameter is cheap: it is a contrast, so a blinded rule may read it; it is estimated on twenty-five times the degrees of freedom the interval itself has; and misstating it by a factor of two costs half a point of coverage.

All of which invites an obvious economy. If the ratio is so cheap and so harmless, estimate it inside each block — where its own two arms are sitting, using its own observations — rather than pooling it across the trial. Each block then gets the weight its own data says it deserves.

That breaks the interval, and the way it breaks is worth more than the fact that it does.

What a decomposition needs

The exactness comes from a weighted least squares argument. With weights w and differences d, the weighted mean and the weighted spread about it are independent, and the spread is a χ² on b − 1 degrees of freedom, provided the weights are constants.

Not provided the weights are correct. Constants. A wrong constant weighting is still exact in the sense of producing an interval whose scale estimate has a computable calibration factor and whose coverage can be worked out; that is the whole content of the factor, and it is why equal weights — which are wrong for the corner and constant — are correctly calibrated everywhere.

A weight estimated inside the block it weights is not a constant. It is a function of the same few observations as the difference it multiplies, and the decomposition has nothing to say about that case. The failure is not that the estimate is imprecise; it is that its imprecision is correlated with the thing it is weighting.

The weights may not read the block they weight. A weighted least squares decomposition needs weights that are constants, or at least independent of the differences they multiply. One λ̂ pooled across the trial is estimated on hundreds of degrees of freedom and is effectively a constant; a λ̂ estimated inside each block is estimated on that block's own two or three, and is correlated with the difference it weights. Coverage falls from 94.68% to 82.76% — and the interval gets wider while doing it, 0.5163 against 0.3024, which is the signature of weights that are noise.
Fig. 1 Coverage on the corner design when the ratio is pooled across the trial and when it is estimated inside each block, with each interval’s half-width beside it.

Coverage falls from 94.68% to 82.76%.

And the interval gets wider while doing it: a half-width of 0.5163 against 0.3024. That is the signature the failure leaves and it is the useful part. An interval that is narrower and covers less is behaving like every over-fitted procedure; an interval that is wider and covers less is one whose weights have become noise, so that the weighted mean is being pulled around by the same randomness the spread is trying to measure.

There is a second reading of the same table that is worth having, because it is what makes the failure detectable in practice. A procedure whose weights have become noise produces intervals with a distinctive shape: the half-widths vary far more between runs than the pooled version’s do, and their average is pulled up by the runs where a badly estimated weight has landed on a block with an extreme difference. The mean half-width of 0.5163 is not a typical half-width; it is an average over a distribution with a long right tail.

So the diagnostic is available without knowing the truth. Run the same interval with weights pooled across the trial and with weights estimated per block, and compare the spread of the widths. If the second is much larger, the weights are carrying noise, and the coverage claim attached to them is not the claim being made. That costs one extra line of arithmetic and requires no simulation.

What the economy actually gives up

The argument for estimating the ratio inside each block rests on it being cheap, and the cheapness is the thing the economy destroys.

Pooled across the trial, the ratio is estimated on about twenty-five times the degrees of freedom the interval itself carries — on twelve blocks that is a few hundred, against the interval’s eleven. Estimated inside each block it has only that block’s own two arms behind it, which is of order twenty.

A variance ratio on a few hundred degrees of freedom is known to within about eight per cent. On twenty it is known to within about thirty.

So the weights the construction needs go from nearly fixed numbers to random variables with a third of their own size in noise — and, worse, noise correlated with the very block difference each weight multiplies. A weighted mean with fixed weights is a linear statistic with a t distribution; a weighted mean whose weights are functions of the same data is a ratio of two correlated random quantities and has no such thing.

That is what the argument for the economy leaves out. It prices the ratio’s accuracy, which is genuinely cheap, and says nothing about the ratio’s independence, which is what the blinded construction was protecting all along.

Where the counts went

The mechanism is a count and it is the mirror of the argument that made the pooled estimate cheap.

Pooled across the corner design, λ̂ is estimated on 1,533 within-arm degrees of freedom. Inside a block of six units split five-to-one, one arm contributes four and the other zero — so the block-level ratio is computed from four observations against one, and in the smallest blocks it is computed from as few as one degree of freedom in an arm.

A variance ratio estimated on that has a spread of the same order as itself. So h_b(λ̂_b) is not approximately the inverse variance; it is a random number roughly centred on it, with a heavy tail inherited from the F distribution’s, and it is correlated with the block difference through the shared observations.

The same estimator, on the same data, is essentially exact when pooled and destroys the coverage when localised, and the only thing that changed is how many observations went into it.

Two failures that look alike and are not

It is worth setting this beside the other thing the earlier field found a two-arm rule may not pool, because the two look like the same mistake and are opposite ones.

A within-block spread computed without the arm label carries δ²m/(2(2m − 1)) of the effect, so the stopping rule reads the thing the trial is measuring, and the trial runs 63% longer at an effect than at a null. That failure is about what the quantity contains: a pooled-over-arms spread is not a contrast, and the whole blinding argument turns on the rule reading contrasts only.

A variance ratio estimated inside each block is a contrast. It contains no mean, it is blind, and it is the right quantity. That failure is about how many observations it is computed from, and it would disappear entirely if the blocks were large.

So one is a failure that no amount of data fixes and the other is a failure that data fixes completely. The first is a claim about the estimator’s construction and the second about its variance, and distinguishing them matters because the repairs are different: the first needs a different quantity, the second needs the same quantity pooled.

Both produce a coverage that is wrong and an interval that is not obviously so, which is why both are refusals rather than warnings.

The abuse that does not work

There is a companion worry that turns out to be almost nothing, and it is worth reporting because it was the expected failure.

A new free parameter in a procedure is an invitation to choose it after the data. Sweep λ over a grid, take whichever setting makes the interval exclude zero, and report that one.

The dial an analyst would be tempted by is the wrong one. Choosing λ after the data so that the interval excludes zero is the obvious abuse of a new free parameter, and it is worth almost nothing: eight candidate ratios spanning a factor of three hundred reject 5.85% of true nulls against 5.05% for a ratio stated in advance. The reason is that the intervals they produce are nearly the same interval — λ moves the weights and the weights move a weighted mean of sixty numbers very little. Saying so is more useful than leaving it unmeasured, because the parameter that does break the coverage is the one estimated inside each block, which nobody would think of as shopping at all.
Fig. 2 The share of true nulls excluded by the interval, for a ratio stated in advance and for the best of eight ratios chosen afterwards.

Eight candidate ratios spanning a factor of three hundred — from 0.2 to 60 against a truth of 25 — reject 5.85% of true nulls against 5.05% for a ratio stated in advance. Eight-tenths of a point.

The reason is that the intervals they produce are nearly the same interval. λ moves the weights, and the weights move a weighted mean of sixty numbers by very little: the estimate is a weighted average and averages are insensitive to weights, which is the same second-order flatness that made the efficiency loss from a wrong λ negligible. The dial an analyst would be tempted by is not the dial that breaks anything.

Saying so is more useful than leaving it unmeasured, because it directs attention at the parameter that does break something — and that one is not shopping at all. Nobody estimating a variance ratio inside each block thinks of themselves as gaming a free parameter. They think of themselves as using the local information, which is normally the more careful thing to do.

What “exact” is doing in all of this

A word has been carrying a lot of weight and it is worth unpacking once, at the point where it stops being available.

Exact here means: the coverage of the interval is its nominal level, for every value of the parameters, at every sample size, conditional on the design that actually happened. It is not an asymptotic statement and it is not an average over designs. That is a strong property and it is why the whole construction is built the way it is — the stopping rule reads contrasts so that it is independent of the means, the weights are constants so that the decomposition holds, and the reference distribution is a t on the between-block degrees of freedom rather than a normal.

Each of the three positions relates to it differently. h_b(λ) is exact. Equal weights are calibrated — their scale estimate is right — which is not the same thing, because the weighted spread is not exactly a scaled χ² when the variances differ, so the t reference is an approximation and the coverage is at its level rather than provably at its level. The effective sizes are neither.

The gap between exact and calibrated is small enough here to be invisible in twenty-five hundred runs, and it is a real gap. A construction that covers at its level in every simulation anybody has run is in a different position from one that covers at its level because a theorem says so, and this whole field exists because the earlier one had the first and wanted the second.

The three positions

Putting all of it together, the corner leaves a choice rather than an answer, and the three positions are not on a line.

Three honest options, and they are not on a line. The corner leaves a choice rather than an answer. Equal weights assume nothing about the variance ratio, are correctly scaled at any ratio, and are 118% wider. The weights h_b(λ) assume the ratio, are exact, and are the narrowest interval available. The effective sizes assume one of two conditions that do not hold here, and are both mis-scaled and wide — over-covering at 98.40%, which reads as caution and is not. The first two are a trade; the third is not on the frontier at all.
Fig. 3 What each weighting assumes, what it covers at, and how wide it is, on the corner design at a required half-width of 0.3.

Equal weights. Assume nothing about the ratio. Calibrated at every λ, because equal weights are a root of the calibration factor. Cover at 95.16%. Half-width 0.6587, which is 118% wider than the best available.

h_b(λ̂). Assume the two arm variances are each constant across blocks, and estimate their ratio from within-arm contrasts. Cover at 94.84%. Half-width 0.3024, the narrowest available.

The effective sizes. Assume one of two conditions that do not hold here. Cover at 98.40%. Half-width 0.4224, which is 40% wider than the best and mis-scaled as well.

The first two are a genuine trade: robustness for width, and the exchange rate is 118%. The third is not on the frontier at all — it is beaten on both axes by the second and on calibration by the first — and it is the one the earlier field’s own weighting lands on.

There is a fourth position that is worth naming and dismissing, because it is what an experimenter usually does. Pick the weights from the design rather than from any variance at all — weight each block by its total size. On the corner design that gives a calibration factor of 1.3953 and an efficiency of 1.0588, so it is better than the effective sizes on both counts and worse than either of the two frontier positions. It is a fourth point inside the frontier rather than a fifth position on it.

The reason it beats the effective sizes is arithmetic rather than insight: the corner alternates block sizes by a factor of eight and allocation ratios by a factor of twenty-five, and the effective sizes respond to both while the total sizes respond to one. Where the true weights are dominated by the size, responding to fewer things happens to land closer.

That is not a recommendation. It is a demonstration that a weighting justified by a condition carries no guarantee once the condition fails, and where it then lands is a fact about the particular design rather than about the weighting.

Is anything exact with no assumption at all?

The honest answer is that equal weights are, and that they are the end of the line.

An exact interval needs weights proportional to the inverse variances, and the inverse variances depend on λ. So a weighting that ignores λ can be exact only by being one of the calibration factor’s two roots, and the λ-free root is the constant one. There is no third root, because a weighting that varies between blocks and does not depend on λ cannot be proportional to a quantity that does, except on designs where the λ-dependence happens to cancel — which is the constant-allocation-ratio condition the earlier field found.

So the position is not that nobody has looked hard enough. It is the two-sample problem’s usual shape: with two unknown variances and no assumption relating them, an exact procedure for a difference of means is not available, and what is available is a procedure that is exact under a stated relationship or one that is conservative under none. The corner is that problem with blocks in it, and the blocks do not change the answer.

What the blocks do change is the price. In the ordinary two-sample problem the conservative option is a small penalty. Here it is 118% in width, because the block compositions swing by a factor of eight and equal weights ignore all of it — so the choice between assuming a ratio and assuming nothing is a much sharper choice than it usually is, and the ratio is much easier to come by than it usually is, because a blinded rule may read it and has thousands of degrees of freedom to read it from.

Which is to say: the corner is the design where the assumption is cheapest to make and most expensive to avoid. That is a good place for it to be, and it is not a coincidence — both facts come from the same source, which is that the blocks differ a great deal in what they are worth.

What is claimed here, and what is not

This essay takes what the exact weighting is not allowed to do, and the claims are that a variance ratio estimated inside each block takes the coverage from 94.68% to 82.76% while making the interval seventy per cent wider, that shopping over eight ratios chosen after the data is worth eight-tenths of a point on a true null, that the corner leaves three positions of which only two are on the frontier, and that there is no λ-free weighting other than the constant one, because the calibration factor has two roots and one of them is the inverse variances.

What stays out and is named as a decision: a design where the arm variances differ between blocks as well as between arms, which would need a ratio per block and is precisely the case the refusal here shows cannot be estimated locally — whether it can be modelled rather than estimated is a real question and is not asked here; and more than two arms, where the weights become a matrix.

The boundary against what a two-arm rule may not pool is that it is about a spread computed without the arm label, which reads the effect; this one is about a ratio computed with the arm label and too few observations, which reads the noise.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The three positions are required to be what they are said to be — one calibrated and inefficient, one exact, one neither — which is a claim about a table rather than about a rate. And the block-level estimate is required to break the coverage by more than a point and a half, since a failure inside the noise would not be a failure.

The refusal for this essay is the variance ratio estimated inside each block. It is the most instructive of the round because it is not carelessness: using local information is normally the more careful choice, and here it substitutes a random number for a constant in the one place a weighted least squares decomposition requires a constant.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A width the trial has to stop for — both name blinding, coverage, estimated variance, fixed-width interval, interval width, inverse variance weighting, variance ratio, weighted least squares
  • Stopping on the arms — both name blinding, coverage, estimated variance, fixed-width interval, interval width, inverse variance weighting, variance ratio, weighted least squares
  • A width rule on skewed outcomes — both name blinding, coverage, estimated variance, fixed-width interval, inverse variance weighting, variance ratio
  • What the blindfold costs — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, interval width
  • Which weights are the inverse variances — both name blinding, block, coverage, degrees of freedom, fixed-width interval, weighted least squares
  • A block size that changes — both name blinding, coverage, degrees of freedom, fixed-width interval, weighted least squares

Named objects

A flat tag is an object no other essay names yet.

Behrens–FisherBlindingBlockCalibrationConservative intervalCoverageDegrees of freedomEfficiencyEstimated varianceFixed-width intervalInterval widthInverse variance weightingSelection effectVariance ratioWeighted least squares