A promise about two arms

What a two-arm rule may not pool

A spread computed "within the block" without the arm label carries a share of the effect, so the trial runs 173 observations at a null and 282 at an effect of 1.5. The stopping rule is reading the thing it exists to measure, and the phrase that produced it is one word long.

Worth reading first: Not half and half · When the looking happens.

The whole construction rests on one restriction: the stopping rule reads quantities that are independent of what the interval will report. In one arm that restriction is easy to state and easy to violate — the rule may read the within-block contrasts and may not read the block means, and a schedule that shrinks its blocks when the running mean is small loses six points of coverage.

Two arms change the forbidden quantity, and the change is subtle enough to be worth three separate measurements. It is no longer the mean that is off limits. It is the running estimate of the difference, and one of the three ways to read it by accident does not look like reading anything.

A word that means two things

A blinded rule needs an estimate of σ². The natural phrase is the spread within a block, and that phrase names two different quantities.

Within arm, within block. Compute each arm’s mean inside the block, take deviations from it, pool. This estimates σ² and involves no comparison between arms at all.

Within block. Compute the block’s mean over all its observations, take deviations from that, pool. This estimates σ² plus a share of δ², because the two arms have different means and the deviations about the common mean contain the gap between them.

The second is what a program computes if nobody tells it about the arm column, and it is one word away from the first in every description of the procedure.

The consequence runs the wrong way in a way that is worth being precise about. The rule stops when the accumulated precision reaches z²/d², and its estimate of the variance is inflated by δ², so the trial gets longer the larger the effect is. At a true null the naive rule spends 173 observations; at an effect of 1.5 it spends 282. The honest rule spends 172 and 172.

That is a dependence between the stopping time and the quantity being reported, which is the one thing the construction forbids. Its practical shape is unusual: the trial does not lose coverage so much as lose all its efficiency exactly where it should have been cheapest — a large effect is the case a trial most wants to finish early on, and this rule runs 64% longer on it.

The estimate is not what is wrong

The interesting part is that the naive rule’s interval is not obviously broken.

The interval is still built from block differences and their between-block spread, both of which are computed correctly. What has changed is only when the run stops. So the failure is not that the reported quantity is wrong; it is that the sample size is a function of δ, which means the procedure — the whole map from data to interval — has a property nobody stated and nobody would want.

This is the same shape as a p-value under an unreported sampling plan: each individual number is computed correctly, and the object they compose into is not the object it is described as. The refusal here fires on the sample size rather than on the coverage, which is why it has to be a separate check — a coverage table would pass it.

Free until the sums stop seeing what the differences see. Coverage with and without the block sums pooled into the interval's variance estimate. With one effect and one level they are free. With an effect that varies between blocks they are still free, because a block's sum picks that variation up exactly as its difference does. With a level that varies they make the interval 37% wider and conservative. And where the effect falls as the level rises — a ceiling, and not an exotic thing to suppose — the sums carry none of the between-block variation while the differences carry all of it, the pooled estimate is short, and the interval that uses it covers 88.75% on a width 20% narrower than the honest one.
Fig. 1 The second refusal, with the requirement on a slider: pooling the block sums into the interval under four things a trial’s blocks might do, one of which costs six points.
The blinding costs nothing here, and what it buys is small. The allocation that minimises the observations needed for a stated width is Neyman's, m_A/m_B = σ_A/σ_B — and the two arm spreads are estimated from within-arm contrasts, which involve no block mean at all, so a rule forbidden to see a mean may compute it. The restriction that costs the one-mean field its width costs nothing at all here. What it buys is 1 − (σ_A + σ_B)²/(2(σ_A² + σ_B²)), drawn as the line: 30.6% at a fivefold difference in spread, which is a great deal less than fivefold. Coverage is unmoved at every ratio.
Fig. 2 What the blinding costs on a difference, which is nothing: the two arm spreads are estimated from within-arm contrasts and involve no block mean, so a rule forbidden to see a mean may compute Neyman’s allocation. What it buys is small — 30.6% at a fivefold difference in spread.

The sums, where the effect moves against the level

The second thing a two-arm rule may not do is the one the previous essay measures, and it belongs in this list because it is a permission that holds under a condition and is usually stated without one.

The block sums are a free second estimate of the same variance, worth 2.16% of width, and they stay free under a constant effect, under a varying effect, and under a varying level — the last conservatively, at 99.20% coverage on an interval 37.35% wider.

They stop being free when the effect falls where the level rises. The sums then carry none of the between-block variation and the differences carry all of it, the pooled estimate is short, and the interval covers 88.75% against the honest 94.95% on a width 20.4% narrower.

The reason to list it as a refusal rather than as a caveat is the direction. Every other way of getting the sums wrong makes the interval wider; this one makes it narrower, and narrower is the direction nobody checks. A trial that reports a suspiciously tight interval and a coverage argument that does not mention period effects has produced exactly this.

Stopping when the interval is narrow enough

The third is the direct analogue of the one-mean field’s most expensive mistake, and it is worth repeating here because the forbidden quantity has changed.

A rule that stops when the interval it is about to report is narrow enough is reading the between-block spread — which is the interval’s own variance estimate, and therefore the reported width itself. It covers 91.22% against the honest rule’s 95.23%, on 146 observations against 173.

It looks like a saving. It spends 15% fewer observations and delivers an interval of the promised width, and there is nothing in its output to indicate what happened. What it has done is stop early on the runs where the between-block spread happened to come out small, which are exactly the runs whose variance estimate is too low, so the interval it reports is systematically the narrow one.

The difference between this and the honest rule is one line and there is nothing in the result to distinguish them, which is why it is a check rather than a warning. The honest rule stops on the within-arm contrasts, which say nothing about the between-block spread; the dishonest one stops on the between-block spread directly.

Why the first one is the dangerous one

Of the three, the pooled spread is the one worth guarding against hardest, and not because it costs the most.

The other two are decisions. Somebody has to decide to pool the block sums, and somebody has to decide to stop when the interval looks narrow enough. Both appear in a protocol as a sentence, and a reviewer who knows to look can find them.

The pooled spread is a default. It is what happens when a variance is computed from a block without a by arm in the call, and the resulting number is a perfectly ordinary within-block variance that looks right, prints right, and is right for the question how spread out are the observations in this block. Nothing about it announces that it is the wrong quantity for a stopping rule.

The size of it is arithmetic and it explains the measurement exactly. Pooling a block of m per arm about its own overall mean gives an expected mean square of σ² + δ²m/(2(2m − 1)), which at four per arm and an effect of 1.5 is 1.643σ². The target is proportional to the variance estimate, so the run should be 64% longer — and it is 64% longer, at 282 observations against 173. Two routes, and the closed form is the one that says the failure has nothing to do with the stopping rule’s mechanics and everything to do with which sum of squares was taken.

And the direction hides it further. The trial runs longer, which is the direction nobody complains about: a trial that overruns looks like a trial with a conservative spread estimate, which is a respectable thing to be. The version of this mistake that made trials too short would have been found years ago.

What each one costs, in the currency the trial is paid in

The three failures are usually stated in coverage, and coverage is only one of the two things this construction promises. It is worth restating them in both.

The pooled spread costs nothing in coverage and 64% in observations at an effect of 1.5. A fixed-width procedure’s two promises are a width and a coverage; this one keeps both and pays in the third quantity, the length, which the procedure never promised and which is what anybody is actually budgeting.

The pooled sums under a ceiling cost 6.2 points of coverage and buy 20.4% of width. This is the only one of the three that breaks a stated promise, and it breaks it in exchange for the other stated promise, which is what makes it hard to see: the interval is narrower, which is the thing the trial was run to get.

Stopping on the reported interval costs 4.0 points of coverage and 15% of the observations. It breaks a promise and pays for it in the budget, which is the most legible of the three and, on the evidence of the one-mean field, still the one people write.

Three failures, three different currencies, and none of them detectable from the interval that gets reported. That is what makes them checks rather than advice.

What the three have in common

All three failures are one thing said three ways: a quantity that is a function of the reported difference has been allowed into the stopping rule or the variance estimate.

The pooled spread lets δ into the target. The sums, under a ceiling, let the between-block variation into the variance estimate asymmetrically. The narrow-enough rule lets the reported width into the stopping time.

And all three are cheap to avoid. The pooled spread needs an arm column. The sums need a sentence about whether the effect is expected to depend on the baseline. The narrow-enough rule needs the stopping condition written on the pooled within-arm estimate rather than on the between-block one.

What none of them needs is more data, a correction, or a different test. The whole of this field’s protection is a rule about which numbers may be looked at, and every failure in it is a number being looked at that should not have been.

Three failures, and how each one would be found

It is worth asking, of each refusal, what would have caught it if the check did not exist — because on this site that question is the one that produced most of the checks.

The pooled spread. Nothing in the analysis would catch it. The interval is correct, the coverage is correct, and the only symptom is a trial that took longer than expected — which every trial does, for a dozen ordinary reasons. It would be found by somebody comparing the realised sample size against the one the observed variance implies, which nobody computes, or by somebody rerunning the rule on simulated data at two effect sizes, which is what the check does.

The pooled sums. A coverage simulation at a null with no period effects passes. A coverage simulation with period effects passes conservatively. Only a simulation with the effect regressed on the level fails, and nobody writes that simulation unless they already suspect the mechanism. This is the one that would have shipped.

Stopping on the reported interval. A coverage simulation catches this immediately, at four points below nominal, which is why the one-mean field found it. It is on the list because being catchable is not the same as being caught: the rule is easier to write than the honest one, its output is indistinguishable, and it looks like a saving.

The pattern is the fleet’s usual one. The failure that no ordinary check sees is the one where every individual number is correct, and both of the first two are of that kind. What separates them is that the second is at least visible in a coverage table given the right world to simulate, and the first is not visible in a coverage table at all.

What a two-arm rule may read

The permission list is short and it is worth stating positively, since three refusals in a row read as though everything is forbidden.

Within-arm within-block contrasts, pooled or by arm. These are what the stopping rule is built on, they estimate σ_A² and σ_B² separately, and they are independent of every block mean.

The block sizes and the allocation, which the rule sets and may set from the contrasts — including Neyman’s allocation, which needs the two arm spreads separately and is therefore computable by a rule forbidden to see a mean.

The block sums, under a constant allocation and where the effect is not expected to move against the level.

And the number of blocks so far, which is a function of everything above and of nothing else.

That is a rule that may size its blocks adaptively, rebalance its arms adaptively, stop when it has the precision it wanted, and report an exact interval. The restriction is real and the room inside it is larger than the one-mean field’s, because a second arm adds a second spread to read.

The construction survives a difference of two weighted means. Coverage of δ̂ ± t√(S_D²/H) on b − 1 degrees of freedom, over 900 runs at a requirement of 0.3, where δ̂ is the block differences weighted by h_b = (1/m_A + 1/m_B)⁻¹ and H is their total. The theorem the one-mean field rests on goes through with h_b in place of the block size, and the reason is that the weights a weighted least squares decomposition needs are the inverse variances — which is exactly what h_b is. The stopping rule reads only within-arm within-block contrasts, so it is a function of nothing the interval reports, whatever it does with the block sizes. Each bar is within 2.9% of the level it claims.
Fig. 3 And what the honest rule does at six settings: covers, at every block size and allocation, whatever it does with them.
Three sets of weights, five designs, and no estimator that is exact everywhere. Coverage of the same interval under three weightings. h_b is the inverse variance when the arms share a variance or the allocation is constant; equal weights are right when every block has the same two counts; the estimated precision weights are right in the limit and exact nowhere, because the decomposition needs the weights to be the constants they are only estimating. In the corner — two variances, changing sizes, changing allocation — the two exact estimators are the ones that miss, at 98.45% and 95.65%, and the one with no theorem behind it is at 95.05%. That is the whole statement: there is an exact estimator under either condition, and none under both.
Fig. 4 Coverage under three weightings across five designs. In the corner — two variances, changing sizes, changing allocation — the two exact estimators are the ones that miss, at 98.45% and 95.65%, and the estimated precision weights, which are exact nowhere, hold their level.

The one that is not a refusal

There is a fourth way a two-arm blinded interval goes wrong and it is deliberately not on the list, because it is not something anybody chooses.

The weights are the inverse variances when the two arms share a variance or when the allocation ratio is the same in every block. Where neither holds — two variances, a block size that changes, and an allocation that changes with it — the effective-size weighting covers 98.45% at a nominal 95%, and the interval is systematically too wide.

That is a failure of the same kind in every respect except one: it is a property of the design rather than a step somebody took. Nobody decided to weight wrongly; they decided to alternate the allocation, and the weighting followed. So it belongs in a table of designs rather than in a list of refusals, and the check for it is arithmetic on the plan rather than a rule about what to look at.

The distinction is worth keeping because the repairs are different in kind. A refusal is repaired by not doing the thing. A design failure is repaired by changing the design — here, by holding the allocation ratio constant, which costs nothing at all — or, failing that, by using a weighting with no exactness argument and saying so.

What is claimed here, and what is not

This essay takes the three things a blinded two-arm rule may not read, and the claims are the sample size growing 64% with the effect under a pooled spread, the six points lost by pooling the sums where the effect moves against the level, and the four points lost by stopping on the interval about to be reported.

What stays out and is named as a decision: any repair for the ceiling configuration, which needs the between-block variation modelled rather than pooled; the power of any of these rules, since every measurement here is at a stated effect rather than against an alternative; and an adaptive allocation that reads the outcome rather than the contrasts, which is the adaptive field’s problem and has an error-rate answer there rather than here.

The boundary against the one-mean field is which quantity is forbidden. That a stopping rule may not read what the interval reports, and what it costs when it does, are established there; what is new is that in two arms the forbidden quantity is a difference, and that one of the three ways to read it is a missing column rather than a decision.

The checks, and the refusals that make them mean something

This field’s library holds three refusals and they are the whole of this essay.

A spread pooled without the arm label: the check runs the rule at a null and at an effect of 1.5 and requires the naive version to get longer while the honest one does not — 173 → 282 against 172 → 172 — and refuses it on the sample size, because its coverage would pass.

The block sums pooled where the effect falls as the level rises: 88.75% against 94.95% on an interval 20.4% narrower, refused while the same pooling under three other configurations is allowed, so that the check is about the condition rather than about the pooling.

And a rule that stops when the interval it is about to report is narrow enough: 91.22% against 95.23% on 146 observations against 173. It is refused for spending fewer observations and covering less, which is the pair that makes it look like a saving.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • A block size that changes — both name blinding, confidence interval, coverage, degrees of freedom, fixed-width interval, independence, sample size, stopping rule
  • The rule that cannot see the mean — both name blinding, coverage, degrees of freedom, fixed-width interval, independence, sample size, stopping rule
  • Blinded, and still exact — both name blinding, contrast, coverage, degrees of freedom, fixed-width interval, two-sample
  • Two degrees of freedom, one total — both name confidence interval, coverage, degrees of freedom, fixed-width interval, sample size, stopping rule
  • What a schedule actually buys — both name blinding, coverage, degrees of freedom, fixed-width interval, sample size, stopping rule
  • What the blindfold costs — both name blinding, coverage, degrees of freedom, fixed-width interval, sample size, stopping rule

Named objects

A flat tag is an object no other essay names yet.

AllocationBlindingConfidence intervalContrastCoverageDegrees of freedomFixed-width intervalIndependencePeriod effectPooled varianceSample sizeSequentialStopping ruleTreatment effectTwo-sample