When a fixed width is reached

The trials that stopped early

A fixed-width trial that stops when its own interval is short enough covers 91.45% — an average of 78.2% among the 22.3% of runs that stop within eight blocks and 96% to 99% among those that run longer. Widening every interval by 17.1% brings the average to 95% and leaves the early stops at 85.6%, while 92.8% of runs now report an interval wider than the width they promised. Even doubling every interval leaves the early stops short.

Worth reading first: The shortest interval is the one that misses · When the looking happens.

A trial that stops when its interval is short enough covers 91.45% where it promised 95%, and a rule that stops on the arms instead brings the same trial back to 94.50%. Both numbers are averages over two thousand runs of one design: a drifting variance ratio, blocks enrolled one at a time, a promised half-width of 0.34, and a cap of thirty-six blocks. An average over runs is a statement about the procedure.

A trial is not the procedure. It is one run, and it knows something about itself that the average does not: how many blocks it enrolled before it stopped. That is the one fact every fixed-width trial reports, because it is the trial’s cost, and it turns out to decide how far the interval beside it can be trusted. What follows are the same two thousand runs, kept one by one rather than averaged, so that what a trial covers can be read at the length it actually stopped at.

What a fixed-width interval covers, by the number of blocks the trial ran before it stopped. Two thousand runs of each rule, the modelled weighting, a promise of 0.34. Reading its report: 4–8 blocks, 22.3% of runs, 78.2%; 9–12 blocks, 16.6% of runs, 90.4%; 13–16 blocks, 18.4% of runs, 96.2%; 17–20 blocks, 17.4% of runs, 96.0%; 21–28 blocks, 17.9% of runs, 96.4%; 29–36 blocks, 7.4% of runs, 99.3% — 91.45% overall. Reading the arms: 4–8 blocks, 0.0%, none; 9–12 blocks, 0.9%, 94.4%; 13–16 blocks, 30.4%, 95.6%; 17–20 blocks, 50.0%, 93.9%; 21–28 blocks, 18.0%, 94.4%; 29–36 blocks, 0.7%, 92.3% — 94.50% overall.
Fig. 1 The coverage of the final interval among the runs that stopped in each band of lengths, for the rule that stops on its own report and the rule that stops on the arms, with the share of runs in each band at the right.

Coverage, by when the trial stopped

Reading its own report, the modelled weighting stops within eight blocks on 22.3% of runs, and those runs cover 78.2%. Stopping at nine to twelve blocks, on 16.6% of runs, they cover 90.4%. From thirteen blocks on, the coverage is above the nominal level in every band: 96.2% at thirteen to sixteen, 96.0% at seventeen to twenty, 96.4% at twenty-one to twenty-eight, and 99.3% for the 7.4% of runs that reach twenty-nine blocks or more.

So the 91.45% is not a trial slightly short of its promise. It is a mixture of trials badly short of it and trials that over-deliver, weighted by how often each happens. Of the 171 runs out of two thousand whose interval misses, 97 stopped within eight blocks: 57% of the misses come from 22% of the runs.

Reading the arms, the picture is flat. No run stops within eight blocks. The 0.9% that stop at nine to twelve cover 94.4%, and the three bands that hold nearly every run cover 95.6%, 93.9% and 94.4%. A rule whose stopping time carries no information about the answer leaves nothing in the stopping time to condition on, so the average it reports is also, near enough, the coverage at every length.

Why the early stops are the ones that miss

Two thousand runs stopping on their own report: the error each reported, against when it stopped. Each point is one run; below the line at one its interval contains the truth. 171 of 2000 miss (8.55%). 445 runs stop within eight blocks and 21.8% of those miss. Errors above 2.6 times the half-width are drawn at 2.6.
Fig. 2 Two thousand runs stopping on their own report: each run’s error as a multiple of the half-width it reported, against the number of blocks it enrolled. Below the line at one, the interval contains the truth.

The misses crowd the left of the picture, and the reason is the rule itself. It stops as soon as tS2/Wt\sqrt{S^2/W} falls to 0.34, where WW is the weight accumulated so far and S2S^2 the weighted spread of the block differences. Early in a trial WW is small, so only a small S2S^2 can satisfy the rule — the spread of the differences has to come out well below its expectation — and a run whose spread came out small reports a half-width that understates how far its estimate can be from the truth. The correlation between the reported spread and the stopping time is 0.938; this is the same fact read one stopping time at a time.

Later in a trial WW has grown, and an ordinary S2S^2 satisfies the rule. The runs that reach twenty blocks are the runs whose spreads were not small enough to stop them sooner, which is the other side of the same selection, and it is why they cover above 95%: selected for spreads that did not come out small, their intervals are if anything too wide for the error they carry.

The two halves of that selection are not symmetric, which is the whole of the damage. The runs that stop early are fewer, but their shortfall is large — seventeen points — and the runs that stop late over-deliver by one to four points each. An average weighted by how often each happens is short by three and a half.

How far past its half-width an early stop’s error goes

The coverage figures count how many runs miss; the runs themselves say by how much. Among the runs that stopped within eight blocks, the typical error — the median — is 0.554 of the half-width the run reported, and one run in ten has an error beyond 1.503 half-widths. Among the runs that stopped at thirteen to twenty-eight blocks, the median is between 0.287 and 0.332 half-widths, and the ninetieth percentile between 0.695 and 0.830: well inside the interval.

Some of the difference is ordinary arithmetic about degrees of freedom, and that part is already in the interval. A run that stops at four blocks uses the t quantile on three degrees of freedom, 3.182, where one that stops at twenty uses 2.093, so an honest interval from four blocks is half as wide again for the same spread. The rule accounts for that. What it cannot account for is that the spread it multiplies was selected: a t quantile is right for a spread drawn at random and too small for a spread that stopped the trial because it came out small. After the degrees of freedom have been paid, the early runs’ errors still sit nearly twice as far out in their own intervals as the late runs’ do.

One widening for the whole trial

Every interval widened by 17.1%, the one constant that restores 95% overall, by when the trial stopped. Reading the report, unwidened: 4–8 blocks 78.2%, 9–12 blocks 90.4%, 13–16 blocks 96.2%, 17–20 blocks 96.0%, 21–28 blocks 96.4%, 29–36 blocks 99.3%. Widened by a factor of 1.171: 4–8 blocks 85.6%, 9–12 blocks 94.6%, 13–16 blocks 98.4%, 17–20 blocks 98.6%, 21–28 blocks 98.0%, 29–36 blocks 100.0% — 95.00% overall. The share of runs reporting an interval wider than 0.34 rises from 1.9% to 92.8%.
Fig. 3 The coverage in each band of stopping lengths as reported and with every interval widened by the one constant factor that brings the runs to 95% overall.

The obvious repair for a 91.45% is a larger critical value. Among the two thousand runs, multiplying every reported half-width by 1.171 — widening every interval by 17.1% — brings the coverage to 95.00%. It is the one constant that does.

At each stopping length it does something different. The runs that stopped within eight blocks now cover 85.6%, still nearly ten points short; nine to twelve blocks, 94.6%; and every band from thirteen blocks on covers between 98.0% and 100.0%. The constant has been set by the average, so it is too small for the runs that needed it and larger than the rest ever needed.

What the widening does to the width

A fixed-width trial promised a half-width of 0.34, and as reported only 1.9% of runs came out wider. Widened by 17.1%, 92.8% do. The rule stopped at the first block where the half-width reached 0.34, so almost every run’s reported half-width sits just below it, and multiplying by 1.171 carries almost every run past it.

That is the conflict stopping on the arms found between the two promises a fixed-width procedure makes, arriving from the other side. Stopping on the arms keeps the coverage and breaks the width on 42.2% of runs. Stopping on the report and widening keeps the average coverage, breaks the width on 92.8% of runs, and still leaves the early stops short — so it is worse than the blinded rule on both promises at once.

No factor fits every stopping time

What a constant widening buys and what it breaks, as the factor grows. Reading the report. At no widening: 91.5% coverage overall, 78.2% among the runs that stopped within eight blocks, 1.9% of runs wider than promised. At 1.171, 95% overall and 85.6% early, with 92.8% wider than promised. The early stops do not reach 95% by a factor of two.
Fig. 4 Across factors from one to two: the coverage of all runs, the coverage of the runs that stopped within eight blocks, and the share of runs wider than promised.

The factor can be pushed further, and the picture does not change shape. At 1.2 the runs cover 95.3% overall and the early stops 85.6%, with 94.5% of runs wider than promised. At 1.5, 97.4% overall and 89.7% early. Even at 2 — every interval doubled — the runs that stopped within eight blocks cover 94.2%, while all runs together cover 98.7% and 99.2% of them are wider than they promised. No constant factor brings the early stops to 95% without first turning every other run’s interval into something much wider than a 95% interval.

The reason is that an early stop’s problem is not a width that is a fixed fraction too small. It is a half-width computed from a spread that was selected for being small, on few degrees of freedom, and a selected spread understates by an amount that varies from run to run. Multiplying by a constant rescales the typical run; the runs that miss are not typical, and the ones that miss by most are the ones whose spreads were luckiest.

A critical value for each stopping time

If one factor cannot fit every stopping time, a factor for each can. Among the runs that stopped within eight blocks, multiplying the half-width by 2.064 makes 95% of them cover; among those that stopped at nine to twelve blocks, 1.220; from thirteen blocks on, no widening is needed at all. Chosen on the same runs it is then counted on, that schedule covers 95% in every band by construction, which is the most it could ever be credited with.

What it does to the width promise is the reverse of what a fixed-width trial is for. Every run that stopped within eight blocks now reports an interval about twice as wide as the one it stopped on, and 96.4% of them exceed the promised 0.34; of the runs that stopped at nine to twelve blocks, 100.0% do. The runs from thirteen blocks on keep the promise because they needed no widening, apart from the runs that reached twenty-nine blocks or more, 25.5% of which hit the cap at thirty-six without their interval ever reaching the width. So the promise is kept by exactly the runs that took long enough not to need it, and broken by every run that stopped early — the runs the rule was built to reward for reaching its width quickly.

A first look that waits

Both stopping rules with the first look moved from the fourth block to the eighth, twelfth and sixteenth. first look at 4, reading the report: 91.45% on 15.65 blocks, 1.9% of runs wider than promised; first look at 4, reading the arms: 94.50% on 18.13 blocks, 42.2% of runs wider than promised; first look at 8, reading the report: 93.05% on 16.67 blocks, 1.8% of runs wider than promised; first look at 8, reading the arms: 94.50% on 18.13 blocks, 42.1% of runs wider than promised; first look at 12, reading the report: 94.15% on 17.89 blocks, 1.7% of runs wider than promised; first look at 12, reading the arms: 94.45% on 18.13 blocks, 42.0% of runs wider than promised; first look at 16, reading the report: 94.15% on 19.55 blocks, 2.1% of runs wider than promised; first look at 16, reading the arms: 94.65% on 18.40 blocks, 41.3% of runs wider than promised.
Fig. 5 Both stopping rules with the first look allowed at the fourth, eighth, twelfth and sixteenth block: coverage with two standard errors, and at the right the mean number of blocks and the share of runs wider than promised.

The rule’s worst runs are the ones that stop at its first few looks, and the simplest repair is not to look that early. Every run so far could stop from its fourth block. Requiring eight blocks before the first look, the rule that reads its report covers 93.05% and uses 16.67 blocks on average; requiring twelve, 94.15% on 17.89 blocks; requiring sixteen, 94.15% again, on 19.55.

Set beside the blinded rule, the twelve-block wait is a genuine alternative rather than a half-measure. Stopping on the arms covers 94.50% on 18.13 blocks and reports an interval wider than the 0.34 it promised on 42.2% of runs. Stopping on the report after a twelve-block wait covers 94.15% on 17.89 blocks — a quarter of a block fewer — and breaks the width on 1.7% of runs. The coverage gap between them, 0.35 points, is inside the standard error of either count, and the width gap is forty points. A wait does not change what the rule reads; it removes the looks at which reading it did the most damage.

It does not remove the pattern. With a twelve-block wait, the runs that stop at the first look they are allowed cover 88.0%, and those that stop at thirteen to sixteen blocks cover 94.9%. The selection moves to wherever the first look is, and waiting shrinks it because a spread built from twelve blocks has more degrees of freedom to be lucky against. The blinded rule’s coverage does not move with the wait at all — 94.45% with a twelve-block wait and 94.65% with sixteen — because it never had early looks to lose.

The rule that never stops early

Two thousand runs stopping on the arms: the error each reported, against when it stopped. Each point is one run; below the line at one its interval contains the truth. 110 of 2000 miss (5.50%). No run stops within eight blocks. Errors above 2.6 times the half-width are drawn at 2.6.
Fig. 6 The same picture for two thousand runs stopping on the arms: each run’s error as a multiple of its reported half-width, against the number of blocks it enrolled.

Stopping on the arms, the same picture has no crowd at the left, because it has no left: no run stops within eight blocks. The blinded rule stops when a width predicted from the first arm’s within-arm sums of squares reaches 0.34, and its stopping time is independent of the block differences, so a run that stops early is not a run whose differences happened to be tight. Its 110 misses out of two thousand are spread along the stopping times rather than gathered at one end.

That is also why the blinded rule’s average is a number a single trial can use. Among the runs that stop at thirteen to sixteen blocks, at seventeen to twenty and at twenty-one to twenty-eight, the coverage is within a point and a half of the average in each band, so a reader who knows when such a trial stopped learns almost nothing more about its interval. A reader who knows when a report-reading trial stopped learns a great deal, and most of it is bad news for the trials that stopped soonest.

The stopping time is information

A reader of one fixed-width trial knows when it stopped. For a trial that stopped on its own report, the relevant coverage is the coverage at that stopping time, and it runs from 78.2% for a trial that stopped within eight blocks of a planned thirty-six to 99.3% for one that ran past twenty-eight. The design’s average of 91.45% describes neither trial.

The same shape turned up in two other settings in this series. Every ordering of a group-sequential trial’s outcomes gave an interval that covered exactly 95% of trials and none that covered 95% at the look a trial stopped at, and the estimate a stopped trial reports was nearly unbiased across trials and badly biased at every early look. A group-sequential boundary and a fixed-width rule are different procedures with the same property: the stopping time is a function of the data, and a guarantee averaged over it is a guarantee about trials other than the one in hand.

What a fixed-width report should carry

What the stopping rule read. A rule that read the interval it reports and a rule that read the arms produce intervals that look identical and mean different things, and only the protocol distinguishes them.

The number of blocks at the stop, against the minimum and the cap. For a rule that read its own report, that number is the most informative thing in the report about how far to trust the interval — more informative than the interval’s own width.

The first block at which the rule could stop. A report-reading rule that could stop at four blocks and one that could not stop before twelve are different procedures — 91.45% and 94.15% on the same design — and a report that gives only the number of blocks enrolled cannot tell them apart.

The coverage at that length, if the rule read the report. It is a property of the design and computable before the trial by running the design, exactly as it was computed here, and a trial that stopped early can report it beside its interval instead of borrowing the design’s average.

What is counted here

Counted. Everything: two thousand runs of each stopping rule on the modelled weighting, the same runs — seed for seed — that the earlier figures of this series average, so the 91.45% and the 94.50% here are those figures’ numbers split rather than new ones. The widening factors are read off those runs, and the coverage at each factor is counted on the same runs it was chosen from, which flatters the widening slightly; it does not change which runs it fails.

Particular to this design. One drifting variance ratio, a promise of 0.34, a minimum of four blocks and a cap of thirty-six, normal outcomes, and the modelled weighting. The length bands are chosen in advance and are coarse; within the first band the runs that stop at four blocks are presumably worse off than the ones that stop at eight, and nothing here resolves that.

Still open: a critical value computed rather than counted, and a rule on skewed outcomes

The critical values for each stopping time were read off the runs they were then counted on, and the first band was too coarse to separate a stop at four blocks from a stop at eight. Both are limits of counting rather than of the idea. For the weighting that is told every true ratio, the spread at each look is a sum of squared normal contrasts and grows by one independent term a block, so the probability of stopping at a given look with a given error over half-width is an integral over the path of that spread — the same kind of recursion a group-sequential boundary is computed by. That would give a critical value at every look, exactly, and say how much of each look’s shortfall belongs to the selection and how much to the few degrees of freedom. Whether the modelled weighting, whose weights are themselves estimated from the spreads, admits anything like it is open.

The rule that never stops early keeps its promise because its stopping quantity, a within-arm sum of squares, is independent of the arm means — and that independence is a property of normal samples. On skewed outcomes a sample’s variance moves with its mean, and stopping on the arms named what should then happen: the stopping time would carry information about the differences again, and the shortfall would return by an amount nothing had measured. A width rule on skewed outcomes measures it: the overall coverage barely moves, and with the skew in one arm the runs that stop within twelve blocks cover about 90%.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BlindingConditional distributionCoverageCritical valueFixed-width intervalInterval widthMonte CarloOptional stoppingRandom sample sizeSequential analysisStopping rule