Concept

Sequential design — where it appears

An experiment that revises what it does next after every observation or batch, rather than fixing the whole plan before anything is measured. It makes the design a function of the data, which is the shape an adaptive allocation already has and which needs the same conditioning to analyse.

Named by 7 essays across 3 fields — each of them below, with the objects they name alongside it.

One experiment, with the blocks getting smaller as the target comes into range. A single run at a requirement of 0.25, with the block sizes 5, 5, 11, 25, 11, 8, 3, 2 and a total of 70 observations in 8 blocks. The rule stops when the observations in hand reach z²σ̂²/d², with σ̂² pooled from the within-block contrasts — an estimate that moves as the run goes on, so the target moves too. Early blocks are large because the target is far away and cannot be overshot; late ones are small because a block is the granularity of the answer. The interval afterwards is built from the 8 block means and from nothing the rule looked at, and it has 7 degrees of freedom against the rule's 62.

A block size that changes

The blinded rule's exactness never needed the blocks to be the same size. Letting the size be chosen from the contrasts as the run goes on leaves the coverage exactly where it was — and runs straight into an identity that says what a schedule can and cannot buy.

pace · Stopping
No block size is best at both things the procedure claims. Two claims and one dial. The honest interval's half-width falls as the blocks get smaller, because the interval's degrees of freedom are the number of blocks: 0.2602 at blocks of two against 0.2933 at blocks of sixteen. The fixed-width claim — that the mean is within 0.25 of the truth — gets more reliable as they get larger, because the sample size is less variable: 92.40% against 95.00%. Both are computed from the same runs, and the second is reproduced to within a tenth of a point by E[2Φ(d√N/σ) − 1], which needs the sample-size distribution and nothing else. The schedules sit at the bottom left: as narrow as the smallest fixed block and as few observations, with the rule's spread estimate on half as many degrees of freedom again.

Two degrees of freedom, one total

The block size is a dial, and the two things a fixed-width procedure claims move in opposite directions along it. Divide the width by the square root of the sample size and one of them turns out to depend on the number of blocks and on nothing else.

pace · Width
One rule keeps its promise and the other keeps its budget. Both stopping rules at five requirements, 1,500 experiments each, with a first stage of 5. The upper curve is the two-stage rule: 97.1%, 96.2%, 95.9%, 96.1%, 96.0% — at or above 95% at every point, which is a theorem rather than a tendency, because its interval is built from a spread estimated before the stopping point was chosen. It pays 2.06×, 2.01×, 2.00×, 1.99×, 1.99× the observations that knowing σ would need. The lower curve is the rule that re-estimates after every observation: 94.5%, 89.3%, 91.0%, 91.5%, 94.3%, on 0.98×, 0.87×, 0.88×, 0.93×, 0.96×. The second rule is the one anybody would run and the first is the one whose claim is true.

Stopping when it is precise enough

An experiment that runs until its estimate is precise enough is the natural design and the one with a theorem against it. Its two-stage cousin keeps its promise exactly, for every unknown spread, and pays twice the observations for it.

guarantee · Stopping
The overshoot is the last block size and nothing else. A run stops at a multiple of its own block sizes and cannot land between them, so it ends past its own target by about half a block. Fixed sizes overshoot by 1.5, 2.2, 3.1, 4.7, 8.5 observations as the size goes 2, 3, 5, 8, 16. Every schedule here ends in blocks of two and every one of them lands where blocks of two land — 1.62, 1.32, 1.37 against 1.48 — while having spent most of the run inside blocks four and eight times larger. That is the one thing on this page a schedule genuinely takes from both ends.

What a schedule actually buys

Big blocks early and small blocks late is the right instinct and it does not take both ends of the trade, because there are not two ends to take. What it does take is the overshoot — about four per cent of the observations — and a steadier stopping point.

pace · Nuisance
What a schedule is allowed to read, and what happens when it reads more. The construction allows the block sizes to be anything at all as long as they are functions of the within-block contrasts, which are independent of every block mean. A schedule that shrinks the block whenever the between-block spread is running above what the contrasts say is a direct attempt to hold down the quantity the interval will be built from, and it succeeds: the estimate lands at 0.8373σ² against the honest 0.9831, and the coverage goes with it. Reading the running mean instead pushes the other way and over-covers — which is not a repair, it is the same violation with the sign reversed, and the level is no longer a property of the procedure at all.

A schedule that reads the mean

The block sizes may be anything at all provided they are functions of the contrasts. Two natural schedules break that, in opposite directions — and the most natural mistake of the three is not a schedule at all but a stopping rule, at 86.87% coverage and fewer observations.

pace · Stopping
Four intervals at one stopping time. 3,500 experiments under the sequential rule with a first stage of 5 and a required half-width of 0.4, which is a demand that knowing σ would meet with 24.0 observations and which the rule meets with 20.8. The rule's own interval covers 90.3%. Replacing the fixed width by a t interval on the same data gives 92.0%. Keeping the rule's own random sample size and drawing a fresh sample of that size gives 89.8% at the fixed width and 95.5% for a t interval — so the sample size being random costs nothing, and the sample size being chosen by the data the interval is built from costs the rest. The spread estimated at the stopping moment is 17.1% below the truth, which is the same fact one level down.

The interval after a stop it chose

A rule that stops when the estimated precision is good enough stops on the samples whose estimate was small. Its interval covers 90% and claims 95%, and a fresh sample of the same random size covers 95.4%.

guarantee · Stopping
The certificate when 4 runs are already spent. a 2² factorial has already been run and 2 further runs are to be placed. The stationarity condition is no longer max d = p; it is max d = (p − λ·tr(M⁻¹M_fixed))/(1 − λ) with λ = 0.6667 the share of runs already spent, which is 7.0985 here. The search reaches 7.098508790 against it, and the largest value anywhere on a 41×41 grid is 7.098508790.

Augmenting a design that has already run

The equivalence theorem still certifies when some runs are already spent, and one number in it changes: the bound is no longer p but (p − λ·tr(M⁻¹M_fixed))/(1 − λ). It equals p again exactly when the runs already made can still be absorbed into the design that would have been chosen — so the certificate says whether the experiment is still recoverable.

optimality · Equivalence

Named alongside it

The objects these essays reach for when they reach for this one.

CoverageFixed-width intervalMonte CarloStopping ruleSample sizeBlockingConfidence intervalDegrees of freedomNuisance parameterBlindingExperimental designPrecision

All concepts