A width the trial has to stop for
Worth reading first: When the looking happens · The shortest interval is the one that misses.
The exact weighting for a difference needs one number per block — the ratio of the two arms’ variances — and when that ratio drifts across a trial, modelling it across blocks covers at 94.85% where estimating it inside each block covers at 92.05%.
Every run in that comparison has twelve blocks. The width is therefore measured rather than required, and the reason it was left that way is stated there: the rules disagree about what a block is worth, so a rule that stops when its own weights say it has enough precision would turn a comparison between intervals into a comparison between stopping times.
But a fixed width is the whole point of the construction. A trial that enrols until its interval is short enough is the ordinary thing to build, and this essay builds it.
The promise, and what happens to it
The trial enrols blocks one at a time, alternating between an even allocation and a lopsided one, on a design whose variance ratio drifts by a factor of twenty from the first block to the last. After the fourth block it computes its interval, and it stops when the half-width falls to 0.34 or when thirty-six blocks are enrolled, whichever comes first.
At twelve blocks fixed in advance the five weightings cover at 95.20%, 94.65%, 95.40%, 92.30% and 95.60% — the earlier result, reproduced on a sequential design.
Stopping when the reported interval is short enough, the same five cover at 91.50%, 91.45%, 90.20%, 89.30% and 90.05%. Every one of them loses between three and five points, against a standard error of 0.49.
Including the rule that knows everything
The row that decides what this is about is the first one. The oracle weighting is handed every block’s true variance ratio — it estimates nothing about the design, it is the infeasible best, and it is the benchmark the whole field is measured against.
It covers at 95.20% at fixed length and 91.50% when the trial stops on its own interval.
So the shortfall is not the weights. It is not the drift, it is not the estimation of the ratio, and it is not the modelling: a rule with nothing left to estimate about the design loses as much as the rules that estimate everything. What is left is the stopping.
And it is not the length either
A trial that stops on its own report runs for 15.6 blocks on average, not twelve, so an obvious explanation is that a longer trial has different properties. It does not survive a control.
Run the same design to sixteen blocks fixed in advance — the average the stopping rule reaches — and the five weightings cover at 94.85%, 94.20%, 94.90%, 90.95% and 95.00%. Within noise of the twelve-block values and three points above the stopped ones.
Same design, same average length, four points of coverage between them. The only difference is whether the number of blocks was decided in advance or by the data.
What a fixed-width procedure is for
It is worth being clear about why anybody stops a trial this way, because the alternative is not obviously worse and the reason is practical rather than statistical.
A trial with a fixed number of blocks delivers whatever precision it happens to deliver. On a design whose variance ratio drifts by a factor of twenty, that precision varies enormously between runs: at twelve blocks the realised half-width here ranges over a factor of more than two across runs of the same design. A sponsor who needs an interval narrow enough to make a decision cannot plan against that — the trial either overshoots and wastes recruitment, or undershoots and answers nothing.
A fixed-width rule replaces that with a guarantee about the answer rather than about the effort: enrol until the interval is short enough, and the trial ends when it has what it came for. The cost is a random sample size, which is a nuisance for planning and is usually thought of as the whole of the trade.
The measurement above says the trade has a second half. The random sample size is not random in a harmless way: it is a function of the data the interval reports, and that is what the coverage pays for.
What the promise was, and what was delivered
Coverage is the natural way to count the failure and it is not the quantity the trial promised. The trial promised a half-width of 0.34, and converting each coverage back into a width says what it actually handed over.
An interval built at ±1.96 standard errors covers c when its width is short by a factor r, with . Inverting the five stopping-rule coverages:
- 91.50% → 13.7% too narrow
- 91.45% → 13.9%
- 90.20% → 18.5%
- 89.30% → 21.6%
- 90.05% → 19.0%
Applied to the target, the trial stops believing it has a half-width of 0.34 and has one of 0.387 to 0.413.
That is the failure stated in the units the construction exists for. Not “the coverage is three to five points low” — the interval is a fifth wider than the number the trial stopped on, and a fixed-width design is a design whose entire output is that number.
And the fixed-length designs, for comparison
The same conversion applied to the twelve-block figures says how much of this is the stopping rule and how much was already there.
95.20% corresponds to an interval 0.9% too wide — over-covering slightly, which is what an exact construction with a t-based width does. 94.65% and 95.40% and 95.60% are within a point of nominal either way.
The fourth weighting is the exception at 92.30%, which is 10.8% too narrow before any stopping rule is applied, delivering 0.377 where it claims 0.34.
So four of the five weightings deliver very nearly what they promise at a fixed length and are out by fourteen to nineteen per cent once the trial stops on its own interval. The fifth is out by eleven per cent to begin with and by twenty-two per cent afterwards — the two effects nearly adding, which is what one would expect of two independent sources of narrowness and is worth checking rather than assuming.
The mechanism, in one correlation
The interval’s half-width is t·√(S²/W), where W is the accumulated weight and S² is the weighted spread of the block differences. A rule that stops when that falls below a target is, by construction, stopping when S² is small — and S² is computed from the very differences the interval is about.
Measured over two thousand runs, the correlation between the weighted spread and the number of blocks is 0.938. The stopping time is very nearly a function of the quantity being reported.
A run that stops early is a run whose differences happened to be tight, so the interval it reports is conditioned on tightness, so it is too short for the truth it is estimating. That is the same argument this collection makes about an interval after a stopping rule that read the effect — arriving here through the variance rather than through the mean, and in a place where the analyst is not looking at the effect at all.
What the trial does keep
The rule keeps the promise it was written for, and that is worth stating precisely, because it is why the defect is not obvious from the output.
Of two thousand runs, 1.9% report an interval wider than the 0.34 they promised. The realised half-width averages 0.323 against a promise of 0.34. Every one of those runs did exactly what it said: enrol until the interval is short enough, then report it.
So the trial delivers a short interval, on time, that covers at 91%. Nothing in its own output says otherwise — the width is right, the weights are right, the model of the drift fits, and the coverage is the one quantity a single trial cannot see.
That is the shape this collection keeps meeting and it is worth naming again here. An advertised property is a statement about a procedure run many times, and every check a single run can make is a check on something else: the estimate, the residuals, the fit, the width. Coverage is counted over runs or it is not counted at all.
The order of the damage is not the order of the weights
One more reading, because it says something about which repairs are worth making.
At fixed length the spread across weightings is 92.30% to 95.60% — three and a third points, and the whole of the previous round’s argument lives inside it. Under stopping the spread is 89.30% to 91.50%, two and a fifth points, and every weighting is below every fixed-length value.
The stopping rule costs more than the choice of weighting does. A trial that models the drift carefully and then stops on its own interval has spent its effort on the smaller of the two problems: the modelled weighting under stopping covers at 91.45%, which is worse than the worst weighting at fixed length.
The ordering between weightings survives — the local estimate is still last, the oracle still first — so the earlier comparison is not overturned. It is relegated: it is a comparison between rules that all share a larger defect.
There is a version of that sentence which is too strong and should be resisted. The weighting still decides how long the trial runs and how much of the drift it can exploit — the local rule stops after 15.3 blocks and the pooled one after 16.6, and both deliver about the same width — so a trial choosing a weighting is choosing its own cost as well as its coverage. What the measurement says is narrower and more useful: fix the stopping rule first. A three-point defect that every weighting shares is worth more attention than a two-point spread between them, and the two repairs are independent of each other.
How the shortfall moves with the promise
A promise is a dial, and the shortfall is not the same at every setting of it.
Asking for a half-width of 0.55 stops the trial after 7.1 blocks on average and covers at 93.92%. Asking for 0.45 takes 9.4 blocks and covers at 92.83%. Asking for 0.34 — the setting everything above uses — takes 16.1 blocks and covers at 91.83%. Asking for 0.28 takes 25.2 blocks and covers at 92.08%.
So the damage deepens as the promise tightens and then stops deepening. Two things are moving against each other. A tighter promise means more chances to stop, and every additional look is another opportunity for a lucky run of tight differences to end the trial early — which is the ordinary optional-stopping arithmetic and makes the shortfall worse. But a tighter promise also means a longer trial, and in a longer trial each block’s contribution to the weighted spread is smaller, so a single tight block cannot carry the decision.
The second effect wins eventually, and where it wins is a fact about this design’s cap rather than a general statement. What is general is the first: a fixed-width rule looks at its own interval after every block, and the number of looks is decided by the promise.
That is worth putting beside the way this collection normally counts looks. A group-sequential trial declares how many interim analyses it will make and spends an error rate across them. A fixed-width trial makes an interim analysis after every block and declares nothing, because it is looking at a width rather than at a p-value — and the looks cost coverage anyway.
Why nobody sees it
The defect has three properties that keep it out of sight, and they are worth naming together because each one alone would be survivable.
The output is right. The interval is the width it promised, the weights are the right weights, the model of the drift fits, and the estimate is unbiased. Nothing in the report is wrong except the confidence level, which is the one number a single trial cannot check.
The instinct points the wrong way. A trial that stops early feels like a trial that got lucky and saved money. What it actually is, is a trial selected for having a small spread, reporting an interval scaled by that spread.
And the fix that suggests itself does not work. Widening the interval — using a larger critical value — restores the coverage on average and destroys the width promise, because the runs that need widening are precisely the ones that stopped early with a short interval. The correct repair is not to change what is reported but to change what is read, which is the subject of the essay beside this one.
What is claimed here, and what is not
This essay takes what a fixed-width promise costs when the trial stops on the interval it is about to report. The claims are that five weightings covering at 92.30% to 95.60% at fixed length cover at 89.30% to 91.50% when the trial stops on its own interval; that the oracle weighting, told every block’s true ratio, loses 3.7 points of coverage like the rest; that the same design run to the average stopped length covers at nominal, so the shortfall is the stopping rather than the length; and that the weighted spread and the stopping time are correlated at 0.938.
What stays out and is named as a decision: one promise, one design, one cap. The width asked for is 0.34, which the design reaches in about sixteen blocks of a possible thirty-six; a tighter promise runs into the cap and a looser one stops at the minimum, and both change the arithmetic. The drift is a factor of twenty, which is the design the weighting field was built on, and a flatter drift would make the weightings agree with each other while leaving the stopping shortfall untouched.
What is also out is the repair, which is long enough to be its own essay: the width can be predicted from quantities the interval is not about, and a rule that stops on the prediction keeps its coverage.
What it would take to keep the promise honestly
Three repairs suggest themselves and only one of them is a repair.
Widen the interval. Use a critical value large enough that the stopped interval covers at 95%. It works on average and it abandons the promise: the runs that need widening are the ones that stopped early with a short interval, so widening them makes the realised width exceed the target on exactly the runs the target was for.
Stop later. Add a margin to the promise — stop at 0.30 and report the 0.34 that was asked for. That buys coverage, and it buys it by running longer for a reason the trial cannot state: the correct margin depends on how much the stopping rule reads, which is the quantity nobody has measured.
Stop on something else. The width the trial will report is predictable from quantities the interval is not about — the within-arm sums of squares, which are independent of the arm differences under normality, and the block compositions, which are design constants. A rule that stops on the prediction has a stopping time that carries no information about the answer, and the conditioning that costs the coverage never happens.
The third is the one that works, it needs nothing the trial does not already have, and it is measured next door. What it costs is two blocks and a promise about the realised width, which is a genuine trade rather than a free repair — and the essay that measures it says so.
The checks, and the refusal
Two claims are gated. Every weighting is required to lose coverage under the report-reading rule, including the oracle, which is the assertion that this is about stopping and not about estimation. And the correlation between the weighted spread and the stopping time is required to be above 0.8 for the rule that reads its own report and below 0.2 for the rule that does not — a pair, because one of them alone is a number and the two together are a mechanism.
The refusal is the construction itself. A fixed-width interval reported after a stopping rule that read its own half-width is rejected: it is the interval a run was selected for having, and it covers at 91.5% against a nominal 95% even when every nuisance parameter in the design is known exactly.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The bias that lands in the slope — both name allocation ratio, blinding, blocking, coverage, estimated variance, fixed-width interval, interval width, variance ratio, weighted least squares
- The condition that cannot be dropped — both name blinding, coverage, estimated variance, fixed-width interval, interval width, inverse variance weighting, variance ratio, weighted least squares
- A width rule on skewed outcomes — both name blinding, coverage, estimated variance, fixed-width interval, inverse variance weighting, stopping rule, variance ratio
- Blinded, and still exact — both name blinding, coverage, estimated variance, fixed-width interval, interval width, inverse variance weighting, variance ratio
- A block size that changes — both name blinding, blocking, coverage, fixed-width interval, stopping rule, weighted least squares
- A schedule that reads the mean — both name blinding, blocking, coverage, fixed-width interval, optional stopping, stopping rule
Named objects
A flat tag is an object no other essay names yet.
Allocation ratioBlindingBlockingCoverageEstimated varianceFixed-width intervalInterval widthInverse variance weightingOptional stoppingRandom sample sizeSequential analysisStopping ruleVariance ratioWeighted least squares