Concept

Random sample size — where it appears

A number of observations decided by the data rather than in advance, which costs nothing unless it depends on the same numbers the interval is built from. Where it does depend on them, the interval reported at the stopping time covers less than the same rule's sample size would with a fresh sample.

Named by 5 essays across 2 fields — each of them below, with the objects they name alongside it.

The stopping rule costs more than the weighting does. Coverage over 2000 runs of a trial whose variance ratio drifts by a factor of twenty, at three ways of deciding when to stop. Twelve blocks fixed in advance is the top line and reproduces what a trial of fixed length delivers. Stopping when the reported interval is short enough is the bottom line, and it costs between 3.0% and 5.5% of coverage — including for the rule that is told every block's true ratio, which is what says the shortfall belongs to the stopping and not to the weights. Stopping on a width predicted from the within-arm sums of squares is the middle line, and it is back at the fixed-length values. The standard error on each point is 0.49%.

A width the trial has to stop for

The weighting that covers at 94.9% on twelve blocks covers at 91.5% when the trial stops as soon as its interval is short enough — and so does the rule that is told every block's true variance ratio. The shortfall is the stopping, not the weights.

stop · Stopping
Two promises, and no rule here keeps both. A fixed-width procedure promises two things: that the interval covers at its nominal rate, and that it is no wider than the width asked for. Over 1500 runs of the modelled weighting, a rule that stops when the interval it will report is short enough keeps the width — only 2.0% of runs come out wider than 0.34 — and covers at 91.13% against a nominal 95%. A rule that stops on a width predicted from the within-arm sums of squares covers at 94.80% and comes out wider than promised on 42.3% of runs. The two promises are in conflict because keeping the second one exactly requires conditioning on the very quantity that has to be independent of the stopping time for the first.

Stopping on the arms

The width a trial will report is predictable from quantities the interval is not about. A rule that stops on the prediction covers at 94.5% where one that stops on the interval covers at 91.5, and it costs two blocks and half of the width promise.

stop · Width
One rule keeps its promise and the other keeps its budget. Both stopping rules at five requirements, 1,500 experiments each, with a first stage of 5. The upper curve is the two-stage rule: 97.1%, 96.2%, 95.9%, 96.1%, 96.0% — at or above 95% at every point, which is a theorem rather than a tendency, because its interval is built from a spread estimated before the stopping point was chosen. It pays 2.06×, 2.01×, 2.00×, 1.99×, 1.99× the observations that knowing σ would need. The lower curve is the rule that re-estimates after every observation: 94.5%, 89.3%, 91.0%, 91.5%, 94.3%, on 0.98×, 0.87×, 0.88×, 0.93×, 0.96×. The second rule is the one anybody would run and the first is the one whose claim is true.

Stopping when it is precise enough

An experiment that runs until its estimate is precise enough is the natural design and the one with a theorem against it. Its two-stage cousin keeps its promise exactly, for every unknown spread, and pays twice the observations for it.

guarantee · Stopping
What a fixed-width interval covers, by the number of blocks the trial ran before it stopped. Two thousand runs of each rule, the modelled weighting, a promise of 0.34. Reading its report: 4–8 blocks, 22.3% of runs, 78.2%; 9–12 blocks, 16.6% of runs, 90.4%; 13–16 blocks, 18.4% of runs, 96.2%; 17–20 blocks, 17.4% of runs, 96.0%; 21–28 blocks, 17.9% of runs, 96.4%; 29–36 blocks, 7.4% of runs, 99.3% — 91.45% overall. Reading the arms: 4–8 blocks, 0.0%, none; 9–12 blocks, 0.9%, 94.4%; 13–16 blocks, 30.4%, 95.6%; 17–20 blocks, 50.0%, 93.9%; 21–28 blocks, 18.0%, 94.4%; 29–36 blocks, 0.7%, 92.3% — 94.50% overall.

The trials that stopped early

A fixed-width trial that stops when its own interval is short enough covers 91.45% — an average of 78.2% among the 22.3% of runs that stop within eight blocks and 96% to 99% among those that run longer. Widening every interval by 17.1% brings the average to 95% and leaves the early stops at 85.6%, while 92.8% of runs now report an interval wider than the width they promised. Even doubling every interval leaves the early stops short.

stop · Width
Four intervals at one stopping time. 3,500 experiments under the sequential rule with a first stage of 5 and a required half-width of 0.4, which is a demand that knowing σ would meet with 24.0 observations and which the rule meets with 20.8. The rule's own interval covers 90.3%. Replacing the fixed width by a t interval on the same data gives 92.0%. Keeping the rule's own random sample size and drawing a fresh sample of that size gives 89.8% at the fixed width and 95.5% for a t interval — so the sample size being random costs nothing, and the sample size being chosen by the data the interval is built from costs the rest. The spread estimated at the stopping moment is 17.1% below the truth, which is the same fact one level down.

The interval after a stop it chose

A rule that stops when the estimated precision is good enough stops on the samples whose estimate was small. Its interval covers 90% and claims 95%, and a fresh sample of the same random size covers 95.4%.

guarantee · Stopping

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageFixed-width intervalStopping ruleBlindingInterval widthMonte CarloOptional stoppingSequential analysisBlockingEstimated varianceExperimental designInverse variance weighting

All concepts