Stopping on the arms
Worth reading first: The shortest interval is the one that misses · When the looking happens.
A fixed-width trial that stops when its own interval is short enough covers at 91.5% against a nominal 95%, and it does so even when every nuisance parameter in the design is known exactly. The stopping time and the reported spread are correlated at 0.938: the rule is stopping on the quantity it is about to quote.
The repair follows from asking what the rule needs to know. It does not need to know the answer — it needs to know how precise the answer is going to be, and on this design those are two different quantities that a normal sample keeps apart exactly.
The width is predictable without looking at the answer
By construction, each block’s difference has variance σ²_A/h_b, where h_b is the block’s effective size — the quantity the exact weighting is built from. So the half-width a set of weights will produce is
which for weights that are the effective sizes collapses to t·√(σ²_A/W).
Everything in that expression except σ²_A is a design constant or an estimate of the variance ratio. And σ²_A can be estimated from the within-arm sums of squares, which under normality are independent of the arm means — so the whole prediction is independent of every block difference the interval is about.
That is the property a blinded rule turns on, applied one level up: there it lets a rule choose its weights without unblinding, and here it lets a rule choose when to stop without unblinding.
And a rule that stops on it keeps its coverage
Stopping on the predicted width instead of on the reported one, the five weightings cover at 95.10%, 94.50%, 95.60%, 91.80% and 94.40% — against 91.50%, 91.45%, 90.20%, 89.30% and 90.05% for the rule that reads its own interval, and against 95.20%, 94.65%, 95.40%, 92.30% and 95.60% at a fixed twelve blocks.
Every weighting is back where it was before there was a stopping rule at all. The correlation that caused the damage is gone with it: between the weighted spread and the number of blocks it is −0.072, which is nothing.
The one row that does not recover is the local estimate, at 91.80%, and it should not — it undercovers at fixed length too, for a reason that has nothing to do with stopping. A repair that fixed that as well would be a repair that was doing something else, and a reader should be suspicious of one that claimed to.
How much of the gap the repair closes
Three rows of five readings invite one arithmetic, and it is the arithmetic the field’s claim rests on: the blinded rule against the fixed-length design, weighting by weighting.
Against a fixed twelve blocks, the rule that reads its own interval is short by 3.70, 3.20, 5.20, 3.00 and 5.55 points — a mean of 4.13.
The blinded rule is short by 0.10, 0.15, −0.20, 0.50 and 1.20 — a mean of 0.35, with one weighting actually covering more than at a fixed length.
So the repair closes 3.78 of the 4.13 points, which is 92% of the gap, and what is left is smaller than the standard error of any single reading at the run lengths this line of fields uses. On a thousand runs a coverage near 95% carries about 0.69 points; on two thousand, 0.49. A residual of 0.35 points is not a residual anybody can measure.
One weighting has a different problem
The fourth column is worth separating out, because averaging it in hides a defect the repair is not aimed at and does not fix.
At a fixed twelve blocks it covers 92.30% — already 2.7 points below nominal, with no stopping rule involved at all. Under the naive rule it covers 89.30%, so of its 5.7-point shortfall, 2.7 points are the weighting and 3.0 are the stopping.
The blinded rule takes it to 91.80%, which is 0.5 below its own fixed-length figure and still 3.2 below nominal. The repair did its job on that weighting exactly as well as on the others; the weighting was short to begin with.
That is worth saying because the five-column layout invites reading the columns as five instances of one phenomenon. Four of them are. The fourth is two phenomena, and only one of them is this essay’s.
What 0.938 is, as a share
The correlation between the stopping time and the reported spread converts into the quantity that actually names the defect.
Squared, 0.880: nearly nine tenths of the variation in when the trial stops is variation in the number it is about to quote. That is not a rule contaminated by the answer at the margin; it is a rule whose timing is almost entirely a function of the answer’s own noise.
Which is why the repair is as large as it is. Removing a dependence that accounts for 88% of a rule’s behaviour ought to move the coverage a great deal, and it moves it 3.78 points of a 4.13-point gap.
What it costs in blocks
The blinded rule runs longer: 18.1 blocks against 15.6 for the modelled weighting, and the same ordering for every other weighting in the table.
That is not a coincidence and it is not a cost that could be engineered away. The report-reading rule is quicker because it stops on lucky runs — the ones whose block differences happened to be tight — and a rule that cannot see them has to wait for the precision to arrive rather than for the noise to be small. Two and a half blocks is the price of the three and a half points, and paying it is what makes the interval mean what it says.
Put in the units a sponsor uses, that is a 16% longer trial. It is worth comparing against the other ways of buying three and a half points of coverage on this design: switching from the local ratio estimate to the modelled one is worth two and a half points and costs nothing, and running four more blocks at fixed length is worth nothing at all. The stopping rule is the only place on this design where a large amount of coverage is available, and it is available at a price that can be stated in advance.
The second promise, which it does not keep
Here the honest reporting is that the repair trades one promise for another.
A fixed-width procedure promises two things: that the interval covers at its nominal rate, and that it is no wider than the width asked for. The report-reading rule keeps the second almost exactly — 1.9% of runs come out wider than 0.34 — and misses the first by three and a half points. The blinded rule keeps the first and comes out wider than promised on 42.2% of runs.
The reason is structural rather than a defect of the implementation. The prediction is unbiased for the width the trial will report, so the realised width lands on either side of it about equally often. Making it land below the target reliably would mean stopping on the realised width — which is the rule that costs the coverage.
The two promises are in conflict, and the conflict is exactly the conditioning that the coverage needs to avoid. A procedure cannot both guarantee the realised width and be independent of the quantity that determines it.
Which promise a trial should keep
The conflict has a resolution and it is a question about what the promise is for rather than about statistics.
A fixed width is asked for because somebody downstream needs the answer to be decisive: an interval wider than 0.34 does not separate the hypotheses the trial was run to separate. On that reading the width is the binding constraint and the coverage is a technical property, and the report-reading rule looks defensible.
It is not, and the reason is what an interval is. An interval that covers at 91.5% is not a 95% interval that is sometimes slightly wide; it is a different interval, narrower than a 95% interval by about the amount that makes it a 91.5% one. The decisiveness it delivers is borrowed from the confidence level, and a reader who takes the width at face value and the level at face value is being told two things that are not simultaneously true.
So the order is: keep the coverage, then buy back as much of the width promise as the recruitment budget allows. The next section prices that.
Buying the width back
The conflict is not total, because the blinded rule can aim below the promise.
Aiming at 0.30 rather than 0.34 takes the share of runs that exceed 0.34 from 43.2% to 19.0%, and the trial from 18.1 blocks to 25.9. Aiming at 0.27 takes it to 5.5% on 34.4 blocks, which is nearly twice the trial for a promise kept nineteen times in twenty.
Coverage is 94.27%, 94.93% and 95.33% along that sweep, which is the useful part: the margin does not buy coverage, it buys the width promise. The coverage was already there and it stays there, so the two properties can be tuned separately — one by what the rule reads, the other by what it aims at.
That separation is the practical result of the field. A trial that wants both can have both, at a stated price in recruitment, and can compute the price in advance: the margin needed is a function of how variable the predicted width is, which is a design quantity.
Two routes to the predicted width
The prediction is a closed expression and it is checked against what the trial actually reports, which is the site’s usual pair and earns its place here for a specific reason: the expression is derived under the assumption that the weights are the effective sizes, and one of the five weightings ignores that.
For the four rules whose weights are the estimated effective sizes, Σw²/h collapses to W and the prediction is t·√(σ̂²_A/W). For the rule that weights every block alike it does not collapse, because a rule that treats blocks as equal still has blocks of different precision — so the general form is computed rather than the collapsed one, and the equal-weight rule stops at 21.2 blocks rather than the 18.1 the collapsed expression would have told it to.
The realised half-widths confirm the prediction is unbiased: against a promise of 0.34 the blinded rule reports 0.337, 0.334, 0.337, 0.325 and 0.333 for the five weightings. A prediction that was systematically low would show up as a realised width above the target on nearly every run rather than on about half of them, which is the next section’s number and is the check that this one is right.
What the rule actually computes, block by block
The construction is short enough to write out, and writing it out is what shows how little it needs.
After each block the trial has, for every block so far: the two arm sizes, the within-arm sums of squares, and the block difference. The stopping rule uses the first two and never the third.
From the within-arm sums of squares it forms the variance ratio estimate its weighting calls for — pooled across blocks, modelled across them, or one per block — and from the ratio and the sizes it forms each block’s effective size h_b, which is (1/m_A + λ/m_B)⁻¹. It sums those to get W, estimates σ²_A by pooling the first arm’s sums of squares, and reads the predicted half-width off the expression above.
Every one of those is a within-arm quantity or a design constant. The block differences enter the reported interval and nothing else, which is the whole of the argument and is checkable by reading the code rather than by measuring anything: the stopping decision is a function whose arguments do not include the differences.
That is a stronger statement than a low correlation, and the correlation is the check on it rather than the claim. A measured −0.072 says the implementation does what the argument says; the argument says it must be zero.
Why the independence is exact rather than approximate
The argument rests on a fact about normal samples: within a group, the sample mean and the sample sum of squares are independent. It is exact rather than asymptotic, it is the same fact that makes a t statistic a t statistic, and it is why this repair is available at all.
Two things follow, and they mark the boundary of the result.
It is a fact about normality, not about large samples. On skewed data the within-arm sums of squares and the arm means are correlated, the stopping time would carry information about the differences again, and the shortfall would return in some amount nothing here measures. A trial that expects skew should treat the repair as approximate.
And it is a fact about the arms separately, which is what makes the quantity blinded. σ̂²_A is computed within arm A alone, and the estimated variance ratio the weights need is a contrast between within-arm sums of squares, so neither reads the difference between the arms. A rule built from them can be run by somebody who has never seen the treatment codes — which is what a blinded rule is, and is a governance property as well as a statistical one.
What it does not repair
Three things survive the change, and a reader should not take a restored coverage number as more than it is.
The weighting still matters. The local estimate covers at 91.80% under the blinded rule, and it covers at 92.30% at fixed length: the stopping rule is repaired and the weighting is still wrong. The two defects are independent and both have to be fixed.
The trial length is still random. A blinded stopping rule stops when the accumulated precision is enough, and how long that takes depends on the sample: 18.1 blocks on average with a spread across runs that a sponsor still has to plan against. What the repair removes is the dependence between the length and the answer, not the randomness of the length.
And the cap still binds sometimes. Under the tightest margin measured here the trial reaches the thirty-six-block cap on a small share of runs, and a run that stops because it ran out of blocks is a run whose interval is whatever it happens to be. Every fixed-width procedure has that boundary, it is the reason the promise is conditional on the cap, and no stopping rule removes it.
What is claimed here, and what is not
This essay takes what a fixed-width stopping rule is allowed to read. The claims are that the half-width a weighting will produce is predictable from σ²_A, the weights and the block compositions alone; that a rule stopping on that prediction covers at 95.10%, 94.50%, 95.60%, 91.80% and 94.40% for the five weightings against 89.30% to 91.50% for a rule stopping on the reported interval; that the correlation between the reported spread and the stopping time falls from 0.938 to −0.072; that the repair costs two and a half blocks; and that it exceeds the promised width on 42.2% of runs, which aiming at 0.27 instead of 0.34 reduces to 5.5% at 34.4 blocks.
What stays out and is named as a decision: normal errors. The independence the whole construction rests on is a property of the normal distribution, and this field measures nothing else. A trial with skewed outcomes has a different problem and possibly the same answer, and the honest statement is that the argument stops where normality does.
Also out: a stopping rule that reads the effect on purpose. Everything here is about a rule that reads the effect by accident, through the spread. A rule that stops when the effect is large enough is a different construction with a much larger literature and a much larger defect, and this collection treats it elsewhere.
The boundary against the essay that measured the damage is that it asks what stopping on the report costs and this one asks what stopping on something else recovers.
The checks, and the refusal
Two claims are gated. The blinded rule is required to run at least as long as the report-reading one for every weighting, which is the assertion that the coverage was bought rather than found. And aiming below the promise is required to reduce the share of over-wide runs by at least a factor of four while increasing the trial length by at least a half — a pair, because either half alone would be consistent with the margin doing nothing.
The refusal is the interval this field started from. A fixed-width interval whose stopping rule read its own half-width is rejected, and the replacement is not a wider interval or a longer trial but a different quantity to stop on: one the interval is not about.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The bias that lands in the slope — both name blinding, blocking, coverage, estimated variance, fixed-width interval, interval width, variance ratio, weighted least squares
- The condition that cannot be dropped — both name blinding, coverage, estimated variance, fixed-width interval, interval width, inverse variance weighting, variance ratio, weighted least squares
- A block size that changes — both name blinding, blocking, coverage, fixed-width interval, independence, stopping rule, weighted least squares
- A schedule that reads the mean — both name blinding, blocking, coverage, fixed-width interval, independence, optional stopping, stopping rule
- What the blindfold costs — both name blinding, coverage, fixed-width interval, interval width, sequential analysis, stopping rule
- What a schedule actually buys — both name blinding, blocking, coverage, fixed-width interval, stopping rule
Named objects
A flat tag is an object no other essay names yet.
BlindingBlockingCoverageEstimated varianceFixed-width intervalIndependenceInterval widthInverse variance weightingOptional stoppingRandom sample sizeSequential analysisStopping ruleVariance ratioWeighted least squares