A promise about two arms

A width promised for a difference

The exact fixed-width interval was built for one mean. Two arms make the target 42.7 units of effective size and each unit costs four observations, so the same promise about a difference costs 169.4 rather than 42.7 — and the theorem survives untouched with the harmonic size in place of the block size.

Worth reading first: The shortest interval is the one that misses · When the looking happens.

The construction is one of the cleanest things in this fleet. Observations arrive in blocks; the stopping rule reads only the within-block contrasts, and the interval is built only from the block means. Those two are independent whatever the rule does, so the interval is exact — not approximately, not asymptotically, and at every block size. Letting the block size be a schedule does not touch it, because the argument never needed the sizes to be equal.

Every rule in both of those fields estimates one mean. Almost nothing anybody runs an experiment for is one mean. This field is the same promise made about a difference between two arms, and it splits cleanly: the theorem survives, and the arithmetic around it does not.

The construction, with one substitution

Block b puts m_A of its observations in one arm and m_B in the other. Its difference d_b = ȳ_A − ȳ_B has variance σ²(1/m_A + 1/m_B), which names the block’s effective size

h_b = ( 1/m_A + 1/m_B )⁻¹

the harmonic mean of the two arm counts, halved. Everything from the one-mean field goes through with h_b wherever the block size stood: the weighted difference δ̂ = Σ h_b d_b / H has variance σ²/H exactly, the between-block spread Σ h_b (d_b − δ̂)²/(b − 1) is σ² times a χ² on b − 1 independent of it, and δ^±tb1SD2/H\hat{\delta} \pm t_{b-1}\sqrt{S_D^2/H} is exact conditional on the sizes.

The reason is one sentence and it is worth having: the weights a weighted least squares decomposition needs are the inverse variances, and h_b is the inverse variance. Nothing about the substitution is a coincidence, and nothing about it is an approximation.

The construction survives a difference of two weighted meansCoverage of δ̂ ± t√(S_D²/H) on b − 1 degrees of freedom, over 900 runs at a requirement of 0.3, where δ̂ is the block differences weighted by h_b = (1/m_A + 1/m_B)⁻¹ and H is their total. The theorem the one-mean field rests on goes through with h_b in place of the block size, and the reason is that the weights a weighted least squares decomposition needs are the inverse variances — which is exactly what h_b is. The stopping rule reads only within-arm within-block contrasts, so it is a function of nothing the interval reports, whatever it does with the block sizes. Each bar is within 2.9% of the level it claims.blocks of four, split evenly93.89%blocks of eight, split evenly94.78%blocks of twenty, split evenly95.78%blocks of nine, two to one95.11%blocks of twelve, three to one94.89%a block size that halves the distance93.78%how the blocks were sized and split900 runs, requirement 0.3, nominal 95%±1.45% on each bar
Fig. 1 Coverage at four block sizes, two allocations and a schedule, with the requirement on a slider. Nothing in this figure is supposed to move.

At a requirement of 0.3, blocks of four cover 93.89%, of eight 94.78%, of twenty 95.78%; an allocation of two to one covers 95.11%, three to one 94.89%, and a block size that halves the distance to the target 93.78%. Every one is within 2.91% of its nominal level, which is four standard errors on nine hundred runs.

The stopping rule reads only within-arm within-block contrasts, so it is a function of nothing the interval reports — whatever it does with the block sizes, whatever it does with the allocation, and whether or not it changes its mind halfway.

How much headroom the coverage check has

“Within 2.91% of nominal, which is four standard errors on nine hundred runs” states a tolerance. What the six readings actually use is worth stating beside it.

A coverage near 95% on nine hundred runs carries a standard error of √(0.95 × 0.05 / 900) = 0.726 points. Against that:

  • blocks of four, 93.89% — 1.53 standard errors below nominal
  • blocks of eight, 94.78% — 0.30 below
  • blocks of twenty, 95.78% — 1.07 above
  • two to one, 95.11% — 0.15 above
  • three to one, 94.89% — 0.15 below
  • the halving schedule, 93.78% — 1.68 below

So the worst reading uses 1.7 of the four standard errors the check allows, and four of the six sit inside one. The gate is not passing narrowly; it has more than half its tolerance unused.

That matters because of what the check is for. The claim is exactness, which is a claim that cannot be established by a measurement — only refused by one. A tolerance set at four standard errors and a worst case at 1.7 is the strongest form the refusal-that-does-not-fire can take at this number of runs.

The one pattern worth watching

There is a shape in the three fixed block sizes that the tolerance does not look at, and it should be recorded rather than smoothed over.

They rise monotonically: 93.89, 94.78, 95.78 as the block goes from four to eight to twenty. Three readings in order happens by chance one time in six, so this is not evidence of anything on its own.

But the direction is the one a small-block effect would produce. A block of four with an even split has an effective size of one, so the interval’s t is on very few degrees of freedom relative to the information in it, and the schedule that halves the distance to the target — which ends in small blocks — gives the lowest reading of the six at 93.78%.

Two readings in the same direction from two different causes of smallness is a thing to check with more runs rather than a thing to report. The way to settle it is not a wider tolerance but a longer run at blocks of four alone: nine thousand runs would take the standard error to 0.23 points, at which a real one-point shortfall would read at four standard errors and a null result would close the question.

The factor of four

What does not survive is the arithmetic, and it is worth doing before any simulation.

The target is stated in effective size: to promise a half-width d, the run needs H ≥ z²σ²/d², which at d = 0.3 is 42.68. That is exactly the number of observations the one-mean version of this promise needs. Here it is a number of effective units, and effective units are not free.

Each block spends m_A + m_B observations to buy h_b of effective size, and

( m_A + m_B ) ( 1/m_A + 1/m_B ) = (1 + k)²/k ≥ 4

by Cauchy–Schwarz, with equality only at m_A = m_B.

Effective size costs four observations, and unbalance costs more. A block that puts m_A in one arm and m_B in the other is worth h = (1/m_A + 1/m_B)⁻¹ to a difference, and the total effective size is what the target is stated in. The observations it takes are (m_A + m_B)(1/m_A + 1/m_B) = (1 + k)²/k, which is 4 at balance by Cauchy–Schwarz and nothing less anywhere. So the same width promised about a difference costs four times what it costs about a mean before anything goes wrong, and an allocation of two to one costs 12.5% more than that. The line is the closed form and the points are runs that were never told it.
Fig. 2 Observations per unit of effective size, at five allocations. The line is the closed form and the points are runs that were never told it.

At balance the cost is 4.0000, counted and closed. At seven to five, 4.1143. At two to one, 4.5000 — 12.5% more. At three to one, 5.3333. At five to one, 7.2000.

So promising a width of 0.3 about a difference takes 169.4 observations where the same width about a mean takes 42.7, and an allocation of three to one takes 229.9. Both of those are counted, and both match 4H and (16/3)H to the digit.

The factor of four is not a cost of the blinding. It is a cost of asking about a difference, and it is there for any method: a difference of two means has twice the variance of one mean at the same total sample, and splitting the sample doubles it again. What the blinding costs is measured in the one-mean field and is unchanged.

Where the four comes from, and why it is exactly four

The Cauchy–Schwarz bound deserves a sentence of its own, because the number four is doing a lot of work in this field and it is worth knowing it is not an approximation of anything.

For any positive m_A and m_B,

(m_A + m_B)(1/m_A + 1/m_B) = 2 + m_A/m_B + m_B/m_A ≥ 2 + 2 = 4

with equality if and only if the two are equal, since t + 1/t ≥ 2. So the cost per effective unit is exactly four at balance and strictly more everywhere else, and the excess is m_A/m_B + m_B/m_A − 2, which is (√k − 1/√k)² in the ratio.

Two consequences are useful. The bound is flat near balance: at seven to five the cost is 4.1143, which is 2.9% over the minimum for a 40% departure from balance. So mild unbalance is nearly free, and an experimenter who has a reason to want 55/45 is paying almost nothing for it.

And the bound rises without limit: at ten to one the cost is 12.1, at a hundred to one 102.0. There is no allocation ratio beyond which the cost stops growing, because the arm with two units is the one setting the variance and nothing the other arm does can help it. That is the same arithmetic that says why a trial with a tiny arm is a trial with a tiny sample.

Neyman’s allocation, and a restriction that costs nothing

The second thing two arms bring is a second variance, and the standard result about it turns out to interact with the blinding in a way worth reporting.

With unequal arm spreads the width is set by σ_A²/m_A + σ_B²/m_B, and the allocation minimising the total observations is Neyman’s: m_A/m_B = σ_A/σ_B. Equal allocation needs 2z²(σ_A² + σ_B²)/d² and Neyman’s needs z²(σ_A + σ_B)²/d², so the saving is

1 − (σ_A + σ_B)² / ( 2(σ_A² + σ_B²) )

The interesting part is not the formula. It is that a blinded rule may use it. The two arm spreads are estimated from within-arm contrasts, which involve no block mean at all — so the quantity Neyman’s rule needs is exactly the quantity the construction already permits the rule to read. The restriction that costs the one-mean field its width costs nothing here.

The blinding costs nothing here, and what it buys is small. The allocation that minimises the observations needed for a stated width is Neyman's, m_A/m_B = σ_A/σ_B — and the two arm spreads are estimated from within-arm contrasts, which involve no block mean at all, so a rule forbidden to see a mean may compute it. The restriction that costs the one-mean field its width costs nothing at all here. What it buys is 1 − (σ_A + σ_B)²/(2(σ_A² + σ_B²)), drawn as the line: 30.6% at a fivefold difference in spread, which is a great deal less than fivefold. Coverage is unmoved at every ratio.
Fig. 3 What allocating from the within-arm contrasts saves, at four spread ratios, against the closed form. The plates are what the interval covered.

At equal spreads the saving is 1.29% against a closed form of 0.00%. At two to one, 9.64% against 10.00%. At three to one, 19.57% against 20.00%. At five to one, 30.60% against 30.77%. Coverage stays at 95.29% at the widest ratio.

And the saving is small. A fivefold difference in spread buys under a third of the observations, which is far less than the ratio suggests, because the objective is a sum of two variances and the optimum of a sum is flat near its minimum. That is the finding, and it is reported as the modest number it is rather than as a technique.

The two routes, and where each one is needed

Every number above exists twice, which is this site’s standing requirement, and the two routes divide the field unevenly.

The closed forms cover almost everything: the target H = z²σ²/d², the cost (1 + k)²/k, Neyman’s saving, the variance of the weighted difference. None of them involves a simulation and none of them is an approximation.

The counted side covers the one thing no formula can supply: whether the coverage is actually its nominal level under a stopping rule that reads the data. The exactness argument is a proof, but the proof assumes a construction that a program either implements or does not — the rule reading only contrasts, the interval reading only block means, the weights being the inverse variances — and a run that quietly violates one of those is a run whose coverage moves. So the coverage table is not confirming a theorem anybody doubts. It is confirming that the code is the theorem.

Where the two routes meet is the cost table, and that is the sharpest check in this essay: the observations per unit of effective size, counted over five hundred runs at each of five allocations, reproduce (1 + k)²/k to four decimal places at every one. The runs were never told the formula, and a mistake in what h_b is — a factor of two, an arithmetic rather than a harmonic mean — would show there first and would show nowhere else, because coverage is remarkably tolerant of getting the weights slightly wrong.

What each of the three costs is, separated

Three quantities have been called costs and they are different and it is worth keeping them apart.

Four observations per effective unit is the cost of asking about a difference. It applies to any method, blinded or not, sequential or not.

(1 + k)²/k over four is the cost of an unequal allocation. It is a design choice, it is 12.5% at two to one, and it is sometimes worth paying for reasons outside this field — a control arm somebody wants more of, a treatment that is expensive.

The width the interval reports is the cost of the exactness, and it is the one-mean field’s number unchanged: the interval reads b − 1 degrees of freedom instead of everything, so it is wider than the one that does not cover. At blocks of four it is 0.3111 and at blocks of twenty 0.3369, on runs that spend 169.4 and 178.4 observations — the larger blocks buy the stopping rule a better spread estimate and charge the interval for it.

The three do not compose into a single number and nothing here tries to make them. What they do is answer a question an experimenter can actually ask: what does an exact fixed-width interval on a difference cost against a fixed sample size of the same length? — and the answer is a factor of four before anything, plus whatever unbalance is being paid for, plus the one-mean field’s width.

Why the block sizes may still be anything

One property carries over so completely that it is easy to pass without noticing, and it is the one that makes the field usable.

The exactness holds conditional on the sizes, and the sizes may be functions of the contrasts. So a two-arm rule may do everything the one-mean rule’s schedules do: large blocks early while the spread estimate is poor, small blocks near the crossing, sizes that halve the remaining distance. The row in the coverage table for a block size that halves the distance covers 93.78% and spends 174.5 observations in 8.0 blocks, against blocks of four spending 169.4 in 42.4.

And it may now do something the one-mean rule could not, because there is a second dial: it may change the allocation as well as the size. Neyman’s rule is exactly that — an allocation chosen from the within-arm contrasts, updated as the two spread estimates improve — and it is admissible for the same reason a schedule is.

Two dials that may both be read from contrasts is more freedom than the one-mean field has, and the next essay is where that freedom stops being free: the weights that make the decomposition exact depend on the two dials together, and there is a corner where no fixed choice of them is right.

A schedule is exact for the same reason a fixed size is. Coverage of the interval built from the block means, over 1500 runs at a requirement of 0.25. The block means are independent of every within-block contrast whatever the sizes are, so a rule that stops on the contrasts and reports the weighted mean with Σm_b(ȳ_b − x̄)²/(b − 1) behind it is exact conditional on the sizes — and the sizes may then be chosen from the contrasts, adaptively, without touching the argument. Each bar is within 2.3% of the level it claims. Nothing here is an approximation that improves with the sample size.
Fig. 4 The one-mean construction this extends: a block size chosen from the contrasts leaves the interval exact, which is the theorem the harmonic substitution carries over.
Three sets of weights, five designs, and no estimator that is exact everywhere. Coverage of the same interval under three weightings. h_b is the inverse variance when the arms share a variance or the allocation is constant; equal weights are right when every block has the same two counts; the estimated precision weights are right in the limit and exact nowhere, because the decomposition needs the weights to be the constants they are only estimating. In the corner — two variances, changing sizes, changing allocation — the two exact estimators are the ones that miss, at 98.45% and 95.65%, and the one with no theorem behind it is at 95.05%. That is the whole statement: there is an exact estimator under either condition, and none under both.
Fig. 5 And the boundary of the substitution: three weightings against five designs, where the corner is the one with two variances, changing block sizes and a changing allocation all at once.

What a fixed sample would have done

The comparison an experimenter actually faces is not against the one-mean version. It is against deciding a sample size in advance from a guess at σ, which is what almost every trial does.

A fixed design needs 4z²σ²/d² observations for the same width, which at d = 0.3 and σ = 1 is 170.7. The blinded rule spends 169.4 at blocks of four and 178.4 at blocks of twenty. So at a correct guess the rule costs about what the fixed design costs, and it buys the thing the fixed design cannot have: it does not need the guess.

That is the whole trade and it is worth stating without decoration. A fixed design with a wrong guess produces an interval of the wrong width — too wide if σ was overestimated, too narrow to be useful if it was underestimated, and in neither case is there anything to be done afterwards. The blinded rule produces an interval of the promised width at any σ, and pays for it by not knowing in advance how long the trial will be.

The uncertainty in the length is not small. At blocks of four the run takes 42.4 blocks on average and its length varies from run to run; a trial that has to book a ward or a machine for a fixed period cannot use a rule whose end is a random variable, and that is a practical constraint this field has no answer to. What it does have is the observation that the expected cost of not needing a guess is, here, under one per cent.

What is claimed here, and what is not

This essay takes the blinded fixed-width construction on a difference, and the claims are the harmonic substitution leaving the interval exact at every block size and allocation, the factor of four and its closed form, and Neyman’s allocation being computable by a rule forbidden to see a mean.

What stays out and is named as a decision: which weights are the inverse variances when the two arms do not share a variance, which is the next essay; the degrees of freedom, which come out one per block short of what the one-mean identity would predict and are the third; and everything the rule may not read, which is the fourth.

The boundary against the one-mean field is the second arm. The independence of block means and within-block contrasts, the exactness that follows, the schedules, and the cost of the exactness in width are all established there and used here without being re-derived.

The checks, and the refusals that make them mean something

Two claims are gated in this field’s library. The interval is required to cover at its nominal level at six rules and two requirements, which is the theorem and which is the whole reason the substitution is worth making. And the observations per unit of effective size are required to match (1 + k)²/k at five allocations and to be exactly four at balance, on runs that were told neither — which fails if the harmonic size were being defined into agreement rather than measured.

The refusals for this field belong to its last essay, and the reason is that nothing in this one is a choice. The construction is forced: given that the rule may read only contrasts and the interval only block means, the weights are the inverse variances and the cost is Cauchy–Schwarz. What can go wrong goes wrong later, when somebody decides which contrasts count.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationBlindingBlockConfidence intervalContrastCoverageDegrees of freedomEffective sample sizeFixed-width intervalHarmonic meanNeyman allocationSequentialStopping ruleTwo-sampleVariance