A ratio that changes between blocks
Worth reading first: What the exactness buys · The shortest interval is the one that misses.
The weighting that makes a fixed-width interval about a difference exact needs one number: λ = σ_B²/σ_A², the ratio of the two arms’ variances. With h_b(λ) = (1/m_A + λ/m_B)⁻¹ the whole construction goes through — the estimate is minimum-variance unbiased, the between-block sum of squares is a scaled chi-square independent of it, the interval is exactly t, and no condition on the block sizes or the allocation is needed.
That field ends with a refusal: the ratio cannot be estimated inside the block it weights. Doing so substitutes a random number for a constant in the one place a weighted least squares decomposition requires a constant, and the coverage falls from 94.7% to 82.8% on an interval that is simultaneously seventy per cent wider.
The refusal is about a design where λ is genuinely one number for the whole trial. A multi-site trial does not have that. Sites differ in how variable their measurements are; a long trial changes instruments; one arm’s variability can drift while the other’s does not. Then the exact weights need one λ_b per block, the pooled estimate is wrong for every block, and the local estimate is the thing that was refused.
There is a third option and it works, and the measurement that says so also says something about what kind of mistake each of the others is.
A design where one ratio cannot be right
Twelve blocks. The even ones are twenty units split evenly; the odd ones are thirty units split three to twenty-seven. So both of the conditions the two-arm field needs fail at once — the arms differ in variance and the allocation ratio is not constant — which is the corner that needed the ratio weights in the first place.
On top of that, λ drifts. It runs from 1.785 in the first block to 35.854 in the last, a factor of 20.1, following an exponential in a block-level covariate. That is the shape a drift in one arm’s variability actually has: a site effect, a period, a change of instrument.
That reversal is the point of the design rather than a decoration. A single ratio for the whole trial cannot merely be imprecise about the weights: it has to get their order wrong over half the run. The pooled estimate lands at 12.835, which is above the true λ for the first seven blocks and below it for the last five, and under it every lopsided block is weighted more heavily than every even one — the arrangement that is correct at the end of the trial and backwards at the start.
Five rules, and two kinds of mistake
The rule that knows every λ_b covers at 95.13% and sets the width. Against that:
One ratio for the whole trial covers at 94.80% — inside a standard error of 0.345 points — and is 19.8% wider. It is wrong about every block and it costs nothing in level.
Equal weights cover at 95.65% and are 22.3% wider. That is the identity this field already had: an interval’s own scale estimate is exactly right when the weights are the inverse variances and when they are all equal, whatever the variances are, so flat weights are calibrated everywhere and efficient nowhere.
The ratio estimated inside each block covers at 92.05% — nine standard errors below the level — and is 0.25% narrower than the oracle. It is the only rule in the table aimed at the quantity that actually varies, and the only one that misses.
Modelling the drift across blocks covers at 94.85% and is 0.58% wider than the oracle. It recovers essentially all of what knowing the ratios is worth.
A wrong weight costs width; a random weight costs level
That is the sentence the table is, and both halves of it have a reason.
Weights that are wrong but fixed leave the estimator unbiased — any fixed weights do — and the interval’s own scale estimate remains a valid estimate of that estimator’s variance, up to a calibration factor that is a function of the design. The factor is computable in closed form and here it is 0.96815, so the interval is three per cent too narrow before efficiency is considered and the efficiency loss is 1.509 in variance. The two multiply to a width 1.209 times the best, and the level survives because the interval is measuring the variance of the estimator it actually computed.
Weights that are random break that. h_b(λ̂_b) is a function of the block’s own within-arm sums of squares, which are the same quantities the between-block spread is being compared against, so the weights and the deviations they weight are no longer independent. The decomposition that makes S² a scaled chi-square requires the weights to be constants; substitute estimates on a handful of degrees of freedom and the product is neither chi-square nor independent of the estimate.
Here the odd blocks give λ̂_b from three units against twenty-seven — two degrees of freedom in the numerator’s denominator — so the estimate is enormously variable, and the variability goes straight into the weights.
Nothing about aiming at the right quantity protects a rule from this. The local estimator is unbiased for λ_b in the sense that matters, and it is the worst rule in the table.
The local rule, calibrated to cover
The local estimate is the only rule in the table that misses its level, and the honest comparison puts it back at its level before ranking it.
An interval covering 92.05% where it claims 95% is short by a factor of 1.96/Φ⁻¹(0.96025) = 1.117. Widening it by that much would take its half-width from 0.9975 of the oracle’s to 1.114.
So a calibrated local rule would be 11.4% wider than the oracle — better than the pooled ratio’s 19.8% and far worse than the modelled drift’s 0.58%. That ordering is the useful one, because it separates two things the raw table runs together. The local rule’s information is real: it is aimed at the quantity that varies, and even after being widened enough to cover it beats the rule that pretends λ is one number. What it lacks is a way to be read at its own level, and the widening factor is not something a trial could compute — it depends on the degrees of freedom in every block’s own estimate.
So the local rule is not simply bad. It is a rule with a real advantage and no way to price its own error, which is the worse of the two failures because nothing in its output reports it.
One number, wrong by seven at one end and three at the other
The pooled ratio’s position in the range is worth locating, because “above for seven blocks and below for five” understates how uneven the error is.
λ runs from 1.785 to 35.854 and the pooled estimate is 12.835. On a logarithmic scale that is 66% of the way up the range rather than half — the pooled value is pulled towards the high-variance end. So it over-states the ratio by a factor of 7.19 in the first block and understates it by 2.79 in the last.
The design’s own reversal is the same asymmetry seen in the weights. An evenly split block is worth 1.44 times a lopsided one at the start and 0.45 times one at the end — a swing of a factor of 3.2 in relative worth across the run, which one number has to straddle.
That is why the pooled rule’s cost decomposes the way it does. Its calibration factor of 0.968 is mild, because a fixed wrong weighting is still a weighting; its efficiency factor of 1.509 is not, because the weights are in the wrong order over half the trial and an ordering error costs far more than a scaling one. A single ratio’s problem is not that it is imprecise. It is that a scalar cannot change sign, and the design is one where the right answer does.
Both routes to the widths
The width ratios are not only counted. Calibration × efficiency is closed arithmetic on the block compositions
and the true variances — no data in it at all — and the half-width ratio it predicts is the square root of the
product.
For one pooled ratio it predicts 1.209 and four thousand runs report 1.198. For equal weights it predicts 1.235 and the runs report 1.223. Both agree to about one per cent, which is what the difference between E[√S²] and √E[S²] costs, and neither route was told the other’s answer.
That check matters more than usual here, because the closed form is what says why. The counted width says the pooled rule is a fifth wider; the arithmetic says 1.509 of that is efficiency — weights that are simply not the inverse variances — and 0.968 is calibration, an interval mis-estimating its own scale. Two different repairs would be needed for the two halves, and only the counted number would not have said so.
What the model is, in four lines
The rule that works is small enough to write out, and writing it out is most of the argument for it.
Each block supplies a within-arm variance estimate for each arm, on m_A − 1 and m_B − 1 degrees of freedom. Take their ratio, λ̂_b. Take its logarithm. Regress that on a block-level covariate — the period, the site index, whatever the drift is thought to follow — across all twelve blocks. Exponentiate the fitted values and use them in h_b.
Three things about that are worth naming.
It is blinded. Every quantity in it is a within-arm contrast, and a within-arm contrast contains no treatment mean, so the whole model can be fitted by somebody who is not permitted to see the difference being estimated. That is the property the exact weighting turns on, and it survives the generalisation intact.
It is fitted before the weights are used, on quantities that are independent of the block differences under normality — the within-arm sums of squares and the between-arm difference are independent within each block. So the weights are not functions of the deviations they weight, which is the exact property the local rule destroys.
And it uses two numbers for twelve blocks rather than twelve for twelve. That is the whole of why it works, and it is the same arithmetic as the shared-nuisance argument everywhere else in this collection.
What the drift has to be
At zero drift the pooled rule is right by construction, and the sweep says what the model costs there.
With λ constant the pooled estimate targets the truth, its half-width is 0.5371 against the oracle’s 0.5372, and modelling a drift that is not there gives 0.5376 — three ten-thousandths wider. The model is estimating a slope that is genuinely zero and paying one degree of freedom for it.
At the other end, with the ratio spanning a factor of twenty-four, the pooled rule is 0.6129 against the oracle’s 0.4991 and the model’s 0.5031. So the model costs a tenth of a per cent where it is unnecessary and saves a fifth where it is not, which is about as favourable as an estimated-nuisance trade gets in this collection.
The reason it is favourable is the one the criterion field measured from the other direction: the model estimates two numbers from twelve blocks where the local rule estimates twelve numbers from one block each. The error it makes is nearly the same in every weight, and a weighting is invariant to a common factor.
Why the wrong weight is the forgiving mistake
The two failures in this table are not two sizes of the same failure. They are different in kind, and the reason is worth separating from the numbers, because it decides which mistake is worth working to avoid.
A weight enters the pooled estimate through a convex combination, and the variance of a convex combination is a smooth function of the weights with a minimum at the inverse-variance point. Near a minimum the loss is second order. Being wrong about λ by a factor of two does not cost twice the width; it costs the curvature times the square of the displacement, which is why one ratio for the whole trial — badly wrong at both ends of a drift from 1.785 to 35.854 — still lands inside a quarter of the oracle’s width. Equal weights, which are wrong everywhere by construction, are barely worse. The penalty for a fixed wrong weight is bounded, predictable, and computable before the trial runs, and the closed arithmetic here computes it from the block compositions alone.
An estimated weight is a different object, because it is a function of the same data it is weighting. The interval’s coverage is derived on the assumption that the weights are constants; they are not, and the correlation between a block’s weight and that block’s own contrast is what the derivation has no term for. That error is first order. It does not shrink because the estimate is nearly right — the ratio estimated inside each block is nearly right, and is aimed at exactly the quantity that varies, which is why it produces intervals narrower than the ones that know every λ. Narrower than the oracle is the tell. An interval cannot beat the oracle’s width honestly; what it has done is spend level to buy width, and 92.05% against a nominal 95% is the price.
So the rule that follows is not estimate the weight better. It is estimate the weight somewhere the estimate cannot see the contrast it will be applied to, which is what modelling the drift across blocks does: each block’s weight is fitted from the other blocks’ variances as well as its own, so the correlation that costs the level is diluted by however many blocks the model reads. At twelve blocks that is enough to recover the level in full, at 94.85%, while giving back 0.58% of width to the oracle. Twelve is not a large number, and the recovery being that complete at twelve is the practical content of the finding.
And the ordering reverses inside the same trial, which is what makes the choice unavoidable rather than academic: the block composition worth 3.591 at the start is worth 0.271 at the end. There is no single weight that is right, so there is no version of this design in which the question can be declined.
What is claimed here, and what is not
This essay takes a variance ratio that is not one number. The claims are that on a design where λ drifts by a factor of 20.1 the order of the block weights reverses, so one pooled ratio must get it wrong over half the run; that the pooled rule loses no coverage — 94.80% against a nominal 95% — and is 19.8% wider; that the rule estimating λ inside each block covers at 92.05%, nine standard errors low, while being narrower than the oracle; and that a model fitted across blocks covers at 94.85% and is 0.58% wider than knowing every ratio.
What stays out and is named as a decision: stopping. Every run here has twelve blocks, so the width is measured rather than required. A genuine fixed-width run would stop when its own weights said it had enough precision, and the rules disagree about what a block is worth — so a stopping rule that reads the weights would make this a comparison between stopping times rather than between weightings, which is a defect this field has already measured.
The boundary against the essay that found the exact weighting is that it is about a ratio that is one number and this one is about a ratio that is a function. The condition that makes the one-number case exact is named there too, and what a drifting ratio does to a slope fitted across the same blocks is a separate finding.
The checks, and the refusals that make them mean something
Two claims are gated. The modelled rule is required to hold the interval’s level while the local rule misses it by six standard errors or more, which is the finding. And the counted width ratios are required to reproduce √(calibration × efficiency) computed from the design alone, which is what separates the pooled rule is wider from the pooled rule is wider for these two reasons in these proportions.
The refusal that bites is the local estimate, and it bites harder here than where it was first written down. On a design where the ratio really does move, the rule that estimates it inside each block is the only one aimed at the right quantity and the only one that misses the level — which is the most instructive form the refusal can take, because using local information is normally the more careful choice.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Blinded, and still exact — both name blinding, coverage, degrees of freedom, efficiency, estimated variance, fixed-width interval, interval width, nuisance parameter, two-sample, variance ratio
- A block size that changes — both name blinding, blocking, coverage, degrees of freedom, fixed-width interval, nuisance parameter, weighted least squares
- What a schedule actually buys — both name blinding, blocking, coverage, degrees of freedom, efficiency, fixed-width interval, nuisance parameter
- What the blindfold costs — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, interval width, nuisance parameter
- Which weights are the inverse variances — both name blinding, coverage, degrees of freedom, fixed-width interval, two-sample, weighted least squares
- A width promised for a difference — both name blinding, coverage, degrees of freedom, fixed-width interval, two-sample
Named objects
A flat tag is an object no other essay names yet.
Allocation ratioBlindingBlockingConservative intervalCoverageDegrees of freedomEfficiencyEstimated varianceExperimental designFixed-width intervalInterval widthNuisance parameterTwo-sampleVariance ratioWeighted least squares