The trials that stopped early
Worth reading first: The shortest interval is the one that misses · When the looking happens.
A trial that stops when its interval is short enough covers 91.45% where it promised 95%, and a rule that stops on the arms instead brings the same trial back to 94.50%. Both numbers are averages over two thousand runs of one design: a drifting variance ratio, blocks enrolled one at a time, a promised half-width of 0.34, and a cap of thirty-six blocks. An average over runs is a statement about the procedure.
A trial is not the procedure. It is one run, and it knows something about itself that the average does not: how many blocks it enrolled before it stopped. That is the one fact every fixed-width trial reports, because it is the trial’s cost, and it turns out to decide how far the interval beside it can be trusted. What follows are the same two thousand runs, kept one by one rather than averaged, so that what a trial covers can be read at the length it actually stopped at.
Coverage, by when the trial stopped
Reading its own report, the modelled weighting stops within eight blocks on 22.3% of runs, and those runs cover 78.2%. Stopping at nine to twelve blocks, on 16.6% of runs, they cover 90.4%. From thirteen blocks on, the coverage is above the nominal level in every band: 96.2% at thirteen to sixteen, 96.0% at seventeen to twenty, 96.4% at twenty-one to twenty-eight, and 99.3% for the 7.4% of runs that reach twenty-nine blocks or more.
So the 91.45% is not a trial slightly short of its promise. It is a mixture of trials badly short of it and trials that over-deliver, weighted by how often each happens. Of the 171 runs out of two thousand whose interval misses, 97 stopped within eight blocks: 57% of the misses come from 22% of the runs.
Reading the arms, the picture is flat. No run stops within eight blocks. The 0.9% that stop at nine to twelve cover 94.4%, and the three bands that hold nearly every run cover 95.6%, 93.9% and 94.4%. A rule whose stopping time carries no information about the answer leaves nothing in the stopping time to condition on, so the average it reports is also, near enough, the coverage at every length.
Why the early stops are the ones that miss
The misses crowd the left of the picture, and the reason is the rule itself. It stops as soon as falls to 0.34, where is the weight accumulated so far and the weighted spread of the block differences. Early in a trial is small, so only a small can satisfy the rule — the spread of the differences has to come out well below its expectation — and a run whose spread came out small reports a half-width that understates how far its estimate can be from the truth. The correlation between the reported spread and the stopping time is 0.938; this is the same fact read one stopping time at a time.
Later in a trial has grown, and an ordinary satisfies the rule. The runs that reach twenty blocks are the runs whose spreads were not small enough to stop them sooner, which is the other side of the same selection, and it is why they cover above 95%: selected for spreads that did not come out small, their intervals are if anything too wide for the error they carry.
The two halves of that selection are not symmetric, which is the whole of the damage. The runs that stop early are fewer, but their shortfall is large — seventeen points — and the runs that stop late over-deliver by one to four points each. An average weighted by how often each happens is short by three and a half.
How far past its half-width an early stop’s error goes
The coverage figures count how many runs miss; the runs themselves say by how much. Among the runs that stopped within eight blocks, the typical error — the median — is 0.554 of the half-width the run reported, and one run in ten has an error beyond 1.503 half-widths. Among the runs that stopped at thirteen to twenty-eight blocks, the median is between 0.287 and 0.332 half-widths, and the ninetieth percentile between 0.695 and 0.830: well inside the interval.
Some of the difference is ordinary arithmetic about degrees of freedom, and that part is already in the interval. A run that stops at four blocks uses the t quantile on three degrees of freedom, 3.182, where one that stops at twenty uses 2.093, so an honest interval from four blocks is half as wide again for the same spread. The rule accounts for that. What it cannot account for is that the spread it multiplies was selected: a t quantile is right for a spread drawn at random and too small for a spread that stopped the trial because it came out small. After the degrees of freedom have been paid, the early runs’ errors still sit nearly twice as far out in their own intervals as the late runs’ do.
One widening for the whole trial
The obvious repair for a 91.45% is a larger critical value. Among the two thousand runs, multiplying every reported half-width by 1.171 — widening every interval by 17.1% — brings the coverage to 95.00%. It is the one constant that does.
At each stopping length it does something different. The runs that stopped within eight blocks now cover 85.6%, still nearly ten points short; nine to twelve blocks, 94.6%; and every band from thirteen blocks on covers between 98.0% and 100.0%. The constant has been set by the average, so it is too small for the runs that needed it and larger than the rest ever needed.
What the widening does to the width
A fixed-width trial promised a half-width of 0.34, and as reported only 1.9% of runs came out wider. Widened by 17.1%, 92.8% do. The rule stopped at the first block where the half-width reached 0.34, so almost every run’s reported half-width sits just below it, and multiplying by 1.171 carries almost every run past it.
That is the conflict stopping on the arms found between the two promises a fixed-width procedure makes, arriving from the other side. Stopping on the arms keeps the coverage and breaks the width on 42.2% of runs. Stopping on the report and widening keeps the average coverage, breaks the width on 92.8% of runs, and still leaves the early stops short — so it is worse than the blinded rule on both promises at once.
No factor fits every stopping time
The factor can be pushed further, and the picture does not change shape. At 1.2 the runs cover 95.3% overall and the early stops 85.6%, with 94.5% of runs wider than promised. At 1.5, 97.4% overall and 89.7% early. Even at 2 — every interval doubled — the runs that stopped within eight blocks cover 94.2%, while all runs together cover 98.7% and 99.2% of them are wider than they promised. No constant factor brings the early stops to 95% without first turning every other run’s interval into something much wider than a 95% interval.
The reason is that an early stop’s problem is not a width that is a fixed fraction too small. It is a half-width computed from a spread that was selected for being small, on few degrees of freedom, and a selected spread understates by an amount that varies from run to run. Multiplying by a constant rescales the typical run; the runs that miss are not typical, and the ones that miss by most are the ones whose spreads were luckiest.
A critical value for each stopping time
If one factor cannot fit every stopping time, a factor for each can. Among the runs that stopped within eight blocks, multiplying the half-width by 2.064 makes 95% of them cover; among those that stopped at nine to twelve blocks, 1.220; from thirteen blocks on, no widening is needed at all. Chosen on the same runs it is then counted on, that schedule covers 95% in every band by construction, which is the most it could ever be credited with.
What it does to the width promise is the reverse of what a fixed-width trial is for. Every run that stopped within eight blocks now reports an interval about twice as wide as the one it stopped on, and 96.4% of them exceed the promised 0.34; of the runs that stopped at nine to twelve blocks, 100.0% do. The runs from thirteen blocks on keep the promise because they needed no widening, apart from the runs that reached twenty-nine blocks or more, 25.5% of which hit the cap at thirty-six without their interval ever reaching the width. So the promise is kept by exactly the runs that took long enough not to need it, and broken by every run that stopped early — the runs the rule was built to reward for reaching its width quickly.
A first look that waits
The rule’s worst runs are the ones that stop at its first few looks, and the simplest repair is not to look that early. Every run so far could stop from its fourth block. Requiring eight blocks before the first look, the rule that reads its report covers 93.05% and uses 16.67 blocks on average; requiring twelve, 94.15% on 17.89 blocks; requiring sixteen, 94.15% again, on 19.55.
Set beside the blinded rule, the twelve-block wait is a genuine alternative rather than a half-measure. Stopping on the arms covers 94.50% on 18.13 blocks and reports an interval wider than the 0.34 it promised on 42.2% of runs. Stopping on the report after a twelve-block wait covers 94.15% on 17.89 blocks — a quarter of a block fewer — and breaks the width on 1.7% of runs. The coverage gap between them, 0.35 points, is inside the standard error of either count, and the width gap is forty points. A wait does not change what the rule reads; it removes the looks at which reading it did the most damage.
It does not remove the pattern. With a twelve-block wait, the runs that stop at the first look they are allowed cover 88.0%, and those that stop at thirteen to sixteen blocks cover 94.9%. The selection moves to wherever the first look is, and waiting shrinks it because a spread built from twelve blocks has more degrees of freedom to be lucky against. The blinded rule’s coverage does not move with the wait at all — 94.45% with a twelve-block wait and 94.65% with sixteen — because it never had early looks to lose.
The rule that never stops early
Stopping on the arms, the same picture has no crowd at the left, because it has no left: no run stops within eight blocks. The blinded rule stops when a width predicted from the first arm’s within-arm sums of squares reaches 0.34, and its stopping time is independent of the block differences, so a run that stops early is not a run whose differences happened to be tight. Its 110 misses out of two thousand are spread along the stopping times rather than gathered at one end.
That is also why the blinded rule’s average is a number a single trial can use. Among the runs that stop at thirteen to sixteen blocks, at seventeen to twenty and at twenty-one to twenty-eight, the coverage is within a point and a half of the average in each band, so a reader who knows when such a trial stopped learns almost nothing more about its interval. A reader who knows when a report-reading trial stopped learns a great deal, and most of it is bad news for the trials that stopped soonest.
The stopping time is information
A reader of one fixed-width trial knows when it stopped. For a trial that stopped on its own report, the relevant coverage is the coverage at that stopping time, and it runs from 78.2% for a trial that stopped within eight blocks of a planned thirty-six to 99.3% for one that ran past twenty-eight. The design’s average of 91.45% describes neither trial.
The same shape turned up in two other settings in this series. Every ordering of a group-sequential trial’s outcomes gave an interval that covered exactly 95% of trials and none that covered 95% at the look a trial stopped at, and the estimate a stopped trial reports was nearly unbiased across trials and badly biased at every early look. A group-sequential boundary and a fixed-width rule are different procedures with the same property: the stopping time is a function of the data, and a guarantee averaged over it is a guarantee about trials other than the one in hand.
What a fixed-width report should carry
What the stopping rule read. A rule that read the interval it reports and a rule that read the arms produce intervals that look identical and mean different things, and only the protocol distinguishes them.
The number of blocks at the stop, against the minimum and the cap. For a rule that read its own report, that number is the most informative thing in the report about how far to trust the interval — more informative than the interval’s own width.
The first block at which the rule could stop. A report-reading rule that could stop at four blocks and one that could not stop before twelve are different procedures — 91.45% and 94.15% on the same design — and a report that gives only the number of blocks enrolled cannot tell them apart.
The coverage at that length, if the rule read the report. It is a property of the design and computable before the trial by running the design, exactly as it was computed here, and a trial that stopped early can report it beside its interval instead of borrowing the design’s average.
What is counted here
Counted. Everything: two thousand runs of each stopping rule on the modelled weighting, the same runs — seed for seed — that the earlier figures of this series average, so the 91.45% and the 94.50% here are those figures’ numbers split rather than new ones. The widening factors are read off those runs, and the coverage at each factor is counted on the same runs it was chosen from, which flatters the widening slightly; it does not change which runs it fails.
Particular to this design. One drifting variance ratio, a promise of 0.34, a minimum of four blocks and a cap of thirty-six, normal outcomes, and the modelled weighting. The length bands are chosen in advance and are coarse; within the first band the runs that stop at four blocks are presumably worse off than the ones that stop at eight, and nothing here resolves that.
Still open: a critical value computed rather than counted, and a rule on skewed outcomes
The critical values for each stopping time were read off the runs they were then counted on, and the first band was too coarse to separate a stop at four blocks from a stop at eight. Both are limits of counting rather than of the idea. For the weighting that is told every true ratio, the spread at each look is a sum of squared normal contrasts and grows by one independent term a block, so the probability of stopping at a given look with a given error over half-width is an integral over the path of that spread — the same kind of recursion a group-sequential boundary is computed by. That would give a critical value at every look, exactly, and say how much of each look’s shortfall belongs to the selection and how much to the few degrees of freedom. Whether the modelled weighting, whose weights are themselves estimated from the spreads, admits anything like it is open.
The rule that never stops early keeps its promise because its stopping quantity, a within-arm sum of squares, is independent of the arm means — and that independence is a property of normal samples. On skewed outcomes a sample’s variance moves with its mean, and stopping on the arms named what should then happen: the stopping time would carry information about the differences again, and the shortfall would return by an amount nothing had measured. A width rule on skewed outcomes measures it: the overall coverage barely moves, and with the skew in one arm the runs that stop within twelve blocks cover about 90%.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The rule that cannot see the mean — both name blinding, conditional distribution, coverage, fixed-width interval, optional stopping, sequential analysis, stopping rule
- A schedule that reads the mean — both name blinding, coverage, fixed-width interval, monte carlo, optional stopping, stopping rule
- What the blindfold costs — both name blinding, coverage, fixed-width interval, interval width, sequential analysis, stopping rule
- A block size that changes — both name blinding, coverage, fixed-width interval, monte carlo, stopping rule
- Stopping when it is precise enough — both name coverage, fixed-width interval, monte carlo, random sample size, stopping rule
- The interval after a stop it chose — both name coverage, fixed-width interval, monte carlo, random sample size, stopping rule
Named objects
A flat tag is an object no other essay names yet.
BlindingConditional distributionCoverageCritical valueFixed-width intervalInterval widthMonte CarloOptional stoppingRandom sample sizeSequential analysisStopping rule