Blinded, and still exact
Worth reading first: The shortest interval is the one that misses · What the exactness buys.
The exact interval for a difference in the corner needs one number that the earlier construction did not: λ, the ratio of the two arm variances. Every objection to it is really an objection to that number, and there are two.
It is unknown. Nothing about a fixed-width rule supposes the variances are given; they are the things being learnt about as the trial runs.
It might be forbidden. The whole point of the construction that this one extends is that the rule may read only quantities that carry no information about the effect. A rule that reads a mean and then decides when to stop has broken the argument that makes the interval exact, and the essays that measure what happens when it does find the coverage falling by four points.
The second objection turns out to be no objection at all, and the first is worth about a hundredth of a per cent.
A ratio of spreads is a contrast
λ = σ_B²/σ_A² is estimated by ŝ_B²/ŝ_A², where each ŝ² is a within-arm sum of squares about that arm’s own mean, pooled across blocks. Every term in it is a difference between an observation and the mean of the observations beside it in the same arm and the same block. No block mean appears, and no difference between arms appears.
That is exactly the class of quantity the blinded construction permits. The argument is the same one that lets a blinded rule compute Neyman’s allocation: an allocation needs the two arm spreads separately, the spreads are within-arm contrasts, and a contrast is orthogonal to the difference the trial is about. Reading it tells the rule nothing about whether the treatment works.
So the exactness survives twice over: the interval is exact given λ, and λ is a quantity the rule was already allowed to see. The construction does not trade blinding for exactness — a trade would have been the expected shape of the answer, and there is not one.
What the estimate is
λ̂ is a ratio of two independent within-arm mean squares, which makes it λ times an F variate with the two arms’ degrees of freedom. That is not an approximation; it is what a ratio of independent sums of squares from normal observations is.
So its spread is available in closed form:
with d₁ and d₂ the two arms’ within-block degrees of freedom.
On the corner design at a required half-width of 0.3 the trial accumulates 307 degrees of freedom in one arm and 1,226 in the other, and λ̂ comes out at 25.13 against a truth of 25. Its spread is 2.277 counted and 2.259 from the F distribution. The two routes share no arithmetic: one is a simulation of the whole trial and the other is two integers and a square root.
A relative spread of nine per cent on a ratio of twenty-five is not a precise estimate of a variance ratio. It does not have to be, and the reason is the next section.
There is a subtlety in that statement worth drawing out, because estimated on hundreds of degrees of freedom is the kind of reassurance that is often given and rarely earned.
The degrees of freedom are within block and within arm. A block of forty-eight units split eight-to-forty contributes seven degrees of freedom to one arm and thirty-nine to the other, and it does so without any reference to what any other block did — so the count really does accumulate, and it accumulates faster than the block count by exactly the block size minus two.
What it does not do is help with anything else. The within-arm degrees of freedom say nothing about whether the effect is constant across blocks, nothing about whether the level drifts, and nothing about the between-block spread the interval is built from. They are informative about one parameter, and that parameter happens to be the one the new construction needs.
The counts that make it cheap
The number that settles the whole question is not λ̂’s spread. It is a comparison of two counts.
A fixed-width rule stops when the accumulated precision reaches its target, and on the way it collects N − 2b within-arm degrees of freedom. On the corner design that is 1,533. The interval it then reports has b − 1 between-block degrees of freedom: 60.
The quantity everybody worried about is estimated on twenty-five times as much information as the quantity nobody worried about. The interval’s own variance estimate — the between-block spread, on sixty degrees of freedom — is by far the noisier of the two, and it is the one the construction has always used and nobody has ever objected to.
That ordering is not an accident of this design. The within-arm count grows with N and the between- block count grows with b, and a fixed-width rule’s N grows faster than its b whenever the blocks are of any size at all. So the ratio of the two counts rises as the trial gets larger, and the estimate of λ gets relatively cheaper the longer the trial runs.
One more consequence of the counts is worth stating because it changes what a trial should report. If the within-arm degrees of freedom are the plentiful resource and the between-block ones are scarce, then the interval’s weakest component is its own variance estimate — sixty degrees of freedom on the corner design, which is a t-quantile of 2.00 rather than 1.96, and which is the entire difference between the reported width and the width a known variance would give.
That is not a defect; it is the price of not knowing the variance, and it is the price the construction was always paying. What is new is being able to say which price is which. The trial pays about two per cent in width for not knowing the between-block variance, and about a hundredth of a per cent for not knowing the variance ratio, and before this the second was the one being worried about.
The ratio of the two counts is the block size
The comparison of counts is the essay’s decisive argument and it has a closed form worth having, because it says what the ratio depends on and what it does not.
The within-arm degrees of freedom are N − 2b and the between-block ones are b − 1, so their ratio is
b(m − 2)/(b − 1)
with m the average block size. On the corner design that is 61 × 25.1/60 = 25.5, which is the twenty-five-fold the essay reports.
The expression settles a question the ratio’s size does not. It converges to m − 2 from above as the number of blocks grows, so a longer trial with the same blocks leaves the ratio essentially where it was — it does not rise with N. The plenty is a property of the block size and not of the trial’s length: blocks of forty-eight give a ratio near forty-six however long the run, and blocks of four give one near two whatever else happens.
That is the honest form of the reassurance and it is a narrower one. A design whose blocks are small does not estimate λ on twenty-five times the information; at blocks of four it estimates λ on twice, and the argument that the ratio is cheap because it is plentiful stops applying. The construction is safe on this design because these blocks are large.
Two prices, and why they differ by two hundred
The essay ends by pricing the two unknowns at about two per cent and about a hundredth of a per cent, which is a ratio of two hundred, and the degrees of freedom only account for a quarter of it.
If both quantities entered the width the same way, a ratio of 25.5 in information would give a ratio of 25.5 in price: the t multiplier’s excess over a normal is roughly 1.3/ν, so sixty degrees of freedom cost 2.0% and fifteen hundred would cost 0.085%.
The measured price of a wrong λ is an eighth of that. The extra factor is that the between-block variance enters the width linearly and λ enters it through the weights, and a weighting that is slightly wrong costs second order — the efficiency loss at the optimum of a smooth criterion is quadratic in the displacement, which is the same flatness that makes every allocation rule in this collection forgiving.
So the two hundred decomposes as 25 from the degrees of freedom and 8 from the order of the cost, and only the first half would have been guessed. That also says which way the reassurance generalises: a design with small blocks loses the first factor and keeps the second, so λ would still be the cheaper unknown — by eight rather than by two hundred.
How wrong it may be
The estimate is imprecise and the question is how much that matters, which is answerable without any estimate at all: put a deliberately wrong λ into the weights and see what happens.
λ enters only through the weights. A wrong λ leaves the estimate unbiased — any fixed weights give an unbiased weighted mean — and moves two things, both of which are closed forms of the design: the calibration factor, and the efficiency.
At half the truth the interval covers at 94.93%; at twice it, 94.27%; at three times, 94.27%; at a fifth, 95.93%. The nominal is 95% and the standard error is 0.5 points.
The curve is flat over a factor of two in either direction and starts to move outside it. The calibration factor says why: it is 1.064 at half the truth and 0.966 at twice it, so an interval built on a ratio misstated by a factor of two has its scale estimate wrong by three to six per cent — which moves a 95% coverage by well under a point.
An estimate with a relative spread of nine per cent is essentially never wrong by a factor of two. So the sensitivity curve and the estimate’s precision meet in the middle with a great deal of room, and that is what makes the feasible rule usable rather than merely definable.
The efficiency barely moves at all
The second consequence of a wrong λ is the more surprising one, and it points the same way.
The efficiency — the variance of the weighted mean against the best available — is 1.015 at a fifth of the truth, 1.001 at half, 1.000 at twice, 1.001 at five times. A ratio misstated by a factor of five costs a tenth of a per cent in variance.
The reason is that an efficiency loss from wrong weights is second-order in how wrong they are. The inverse-variance weighting is a minimum of a smooth function of the weights, and a minimum is flat: the first derivative is zero there, so a proportional error of ε in the weights costs ε² in variance. At a factor of two the weights are wrong by tens of per cent and the variance is wrong by tenths.
So misstating λ costs almost nothing in width and a little in calibration, and the calibration is the binding constraint. That reverses the natural worry, which is that a wrong weighting would be inefficient; the inefficiency is negligible and the mis-scaling is what shows up.
Why a wrong ratio is not like a wrong weight elsewhere
It is worth setting the sensitivity here against the failure the earlier field found, because they look similar and are not.
The effective-size weights in the corner are wrong by a fixed amount determined by the design, and they produce a calibration factor of 1.5808 — a fifty-eight per cent error in the interval’s own scale estimate. λ̂ misstated by a factor of two produces a calibration factor between 0.966 and 1.064, a six per cent error at worst.
The difference is not that one estimate is better than the other. It is that h_b(1) is not an estimate of anything: it is a different weighting, chosen for a condition that does not hold, and it is wrong by however far the design happens to be from that condition. There is no sense in which it converges to the right answer as the trial grows. λ̂’s error shrinks with the trial and h_b(1)'s does not.
An estimate with a known distribution and a shrinking error is a different object from a misspecification with neither, and the corner is the design that separates them. On the four designs before it they are identical, which is exactly why the distinction went unmade.
What running the feasible rule actually gives
Putting the estimate in rather than the truth, on the corner design: coverage 94.84% against the known-λ rule’s 94.84%, and a half-width of 0.3024 against 0.3023.
The two are the same rule to four significant figures. The estimate is noisy, its noise is second-order in the width and small in the calibration, and both effects are averaged over sixty blocks before anything is reported.
There is one more thing worth saying about it, and it is the reason this field exists rather than
merely restating an earlier one. The estimated-precision weights that the earlier field measured and
could not justify are this rule. Its precision weighting is one over ŝ_A²/m_A + ŝ_B²/m_B, which is
h_b(λ̂) divided by ŝ_A² — the same vector, up to a common factor a weighted mean cannot see. That is
checked as an identity rather than argued: the two columns agree to a billionth in half-width on every
design.
So the rule that field described as exact nowhere and the only one at its level in the corner was holding its level because it is the exact rule, run at an estimated ratio. What it was missing was not a correction. It was the theorem.
That is worth one more sentence, because the practical difference between the two situations is larger than it sounds. A rule that holds its level with no theorem behind it is a rule nobody can extend: it cannot be quoted for a design it has not been simulated on, it cannot be told what its coverage depends on, and its behaviour anywhere new is a matter for another simulation. A rule with a theorem behind it carries a statement of what it needs — here, that the arm variances are constant within an arm across blocks, and that λ is estimated from quantities independent of the block differences — and can be checked against those conditions rather than against a table of designs somebody happened to run.
What is claimed here, and what is not
This essay takes where the variance ratio comes from, and the claims are that it is a ratio of within-arm contrasts and so readable by a blinded rule, that λ̂ is λ times an F variate whose spread is 2.259 by closed form and 2.277 counted, that a run of the corner design collects 1,533 within-arm degrees of freedom against the interval’s 60, that coverage stays at its level over a factor of two in either direction, that the efficiency loss from a misstated ratio is second-order and under two per cent at a factor of five, and that the feasible rule and the known-ratio rule agree to four significant figures.
What stays out and is named as a decision: any variance structure other than one per arm — a variance that changes between blocks would need a ratio per block, which is the refusal the next essay measures; and non-normal observations, since the F distribution for λ̂ is a normal-theory result and the sensitivity sweep would have to be redone.
The boundary against the essay that derived the weights is that it establishes the exactness given λ and this one supplies λ. The boundary against the schedule that reads a mean is that it measures what reading a forbidden quantity costs, which is the thing the contrast argument here is avoiding.
The checks, and the refusals that make them mean something
Three claims are gated in this field’s library. The estimated-precision weights are required to be h_b(λ̂) to a billionth, which is an identity and is the whole of the connection to the earlier field. The feasible rule is required to match the known-ratio rule in both coverage and width. And λ̂’s spread is required to reproduce the F distribution’s from two degree-of-freedom counts.
The refusal for this essay is a variance ratio estimated from the spread of the block differences. That spread is what the interval is about, so a rule reading it is no longer blind — and it is simply the wrong quantity, since the between-block spread carries the variance of the difference rather than the ratio of the two arm variances.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A ratio that changes between blocks — both name blinding, coverage, degrees of freedom, efficiency, estimated variance, fixed-width interval, interval width, nuisance parameter, two-sample, variance ratio
- A width the trial has to stop for — both name blinding, coverage, estimated variance, fixed-width interval, interval width, inverse variance weighting, variance ratio
- What the blindfold costs — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, interval width, nuisance parameter
- What a schedule actually buys — both name blinding, coverage, degrees of freedom, efficiency, fixed-width interval, nuisance parameter
- What a two-arm rule may not pool — both name blinding, contrast, coverage, degrees of freedom, fixed-width interval, two-sample
- A block size that changes — both name blinding, coverage, degrees of freedom, fixed-width interval, nuisance parameter
Named objects
A flat tag is an object no other essay names yet.
BlindingCalibrationContrastCoverageDegrees of freedomEfficiencyEstimated varianceFixed-width intervalInterval widthInverse variance weightingNeyman allocationNuisance parameterPlug in estimateTwo-sampleVariance ratio