A design that assumes less

The design that stops guessing

Every repair so far protects a guess. The alternative is to run part of the experiment, estimate the parameter from it, and design the rest at the estimate — which recovers most of what a threefold wrong guess costs, and has a best moment to stop guessing that is earlier than anyone expects.

Worth reading first: The design that needs the answer · The design for the worst case.

Both repairs so far take the guess as given and try to survive it. Averaging over a prior spreads the runs across the values the guess might be wrong about; the maximin design places them so that no value in a stated range is badly served. Neither of them ever finds out what the parameter is.

The experiment does. Runs produce responses, responses estimate the parameter, and the estimate can be used to place the runs that have not happened yet. The design becomes a function of the data — which is the shape the adaptive field already has an error-rate answer for, and the answer here is different, because the quantity being read is the one being estimated rather than an outcome somebody is being treated with.

One experiment finding out where to look

One experiment finding out where to lookA single run of the fully sequential design: 40 runs, the first 8 placed at the guess K = 1, then the model refitted and the design revised after every 2. The marks are the settings the runs were made at. The horizontal lines are where a design built at the truth K = 3 would have put them — 1.875 and 10.00 — and the rule walks onto them without being told: its estimate of K after the first eight runs was 2.694, and by the end 2.765 against a truth of 3. The whole experiment is 96.5% as efficient as the design that knew the answer, where running all 40 at the guess would have been 81.1%.02.5057.501010203040run, in the order it was madethe setting it was made atthe guessoptimal at the truth: 1.88optimal at the truth: 10.00one experiment, seed 91050, guess K = 1, truth K = 3, noise 0.0596.5% efficient, K̂ = 2.765
Fig. 1 Forty runs in the order they were made. The first eight are placed at the guess K = 1. After that the model is refitted and the design revised every two runs, and the settings walk onto the two the truth would have chosen — without being told the truth.

The mechanism is as plain as the picture. The locally D-optimal design for this model is two settings, KT/(2K + T) and the end of the range, with half the runs at each; the first is a function of K and the second is not. So the visible movement in that picture is one setting migrating, and where it stops is a statement about what the experiment has learned. The estimate after the first eight runs is 2.69 against a truth of 3, and by the end 2.765.

The whole experiment comes out 96.5% as efficient as the design that knew the answer. Running all forty runs at the guess would have been 81.1%.

One experiment finding out where to look. A single run of the fully sequential design: 40 runs, the first 8 placed at the guess K = 0.25, then the model refitted and the design revised after every 2. The marks are the settings the runs were made at. The horizontal lines are where a design built at the truth K = 3 would have put them — 1.875 and 10.00 — and the rule walks onto them without being told: its estimate of K after the first eight runs was 2.557, and by the end 2.765 against a truth of 3. The whole experiment is 90.6% as efficient as the design that knew the answer, where running all 40 at the guess would have been 34.6%.
Fig. 2 The same forty runs from a twelvefold wrong guess. The first eight are placed far too early, the first refit moves the design most of the way, and what is lost is the eight runs — not the experiment.

What it recovers, against what it costs

Updating between runs, against never updating at all. A threefold wrong guess — the design is built for K = 1 and the truth is 3 — and three ways of spending the runs. The flat line is the local design at the guess, 81.1% efficient at any size, because running the wrong design more times does not make it a better design. The two-stage design spends 40% of its runs there and the rest at the estimate, and lands near 93% throughout. The fully sequential design updates after every 2 runs and improves with size, from 87.8% at 12 runs to 99.1% at 160 — crossing the two-stage design where its early estimates stop being noise.
Fig. 3 Three ways of spending the runs, from twelve to a hundred and sixty of them, at a threefold wrong guess. The flat line is the local design at the guess. The middle curve spends 40% of its runs there and the rest at the estimate. The rising curve revises after every two runs.

The flat line is the reason the field exists: 81.1% at every size, because running the wrong design more times does not make it a better design. Efficiency is a ratio to what was available, and what was available scales with the runs exactly as the wrong design does.

The two-stage design lands between 90.8% and 93.0% across the whole sweep and does not improve much with size — which is a property of its 40% first stage rather than of two-stage designs in general, and the next section is about that. The fully sequential design goes from 87.8% at twelve runs to 99.1% at a hundred and sixty.

They cross. On a small experiment the two-stage design is better; on a large one the sequential one is. The crossing is at about twenty runs here, and the reason is that an early estimate is mostly noise: a design revised after four runs is a design placed at a number with a large standard error, and acting on it wastes the runs that follow. The sequential design’s advantage is that it acts on every estimate, and its disadvantage is the same sentence.

Why the estimate improves as well as the design

Efficiency is a statement about the design and it is not the only thing that moves. The same experiments produce an estimate of K, and it is better too.

At forty runs and a threefold wrong guess the root mean squared error of K̂ is 0.292 from the fixed design at the guess, 0.264 after two stages, and 0.256 after full sequential updating — against 0.259 for the design that knew the answer. The sequential design’s estimate is as good as the oracle’s, which is what an efficiency of 96.5% means when it is cashed into the quantity anybody reports.

That is not a second result. D-efficiency is a statement about the determinant of the information matrix, the information matrix is what the standard errors come from, and an estimate from a well-placed experiment is more precise than an estimate from a badly placed one. It is worth measuring separately because it is the currency the experiment is actually spent in: nobody reports a design’s efficiency in a paper, and everybody reports a standard error.

How long to go on guessing

If the first stage is where the parameter is learned and the second is where the runs are spent well, the obvious question is how to split them — and the obvious answer, half and half, is wrong by a long way.

How long to go on guessing. 40 runs in total, a guess of K = 0.25 against a truth of 3, and the first stage varied from 5% of the runs to 70%. The optimum is interior and it is early: 89.0% at 10%, which is 4 runs — barely more than the four a two-parameter model needs to be fitted at all. Spending less leaves the second stage designed at an estimate with nothing in it; spending more leaves most of the experiment at a setting already known to be wrong, and by 70% the whole procedure has given back 27.4 points of what it bought.
Fig. 4 Forty runs, a twelvefold wrong guess, and the first stage swept from a twentieth of the experiment to seven tenths of it. The optimum is interior and early: 89.0% at a tenth, which is four runs.

The sweep runs 87.1%, 89.0%, 87.4%, 83.6%, 73.9%, 61.6% at first stages of 5%, 10%, 20%, 30%, 50% and 70%. The best is at a tenth of the experiment — four runs, which is the smallest number that can fit a two-parameter model at two settings and estimate anything at all.

Both ends of that curve are informative. Spending less than four runs leaves the second stage designed at an estimate with essentially no information in it, and the efficiency falls back towards what the guess alone would have given. Spending seven tenths of the experiment at the guess gives back most of what the procedure was for: at 61.6% it is closer to the local design’s 34.6% than to the 89% available.

How long to go on guessing. 20 runs in total, a guess of K = 0.25 against a truth of 3, and the first stage varied from 5% of the runs to 70%. The optimum is interior and it is early: 85.2% at 5%, which is 1 runs — barely more than the four a two-parameter model needs to be fitted at all. Spending less leaves the second stage designed at an estimate with nothing in it; spending more leaves most of the experiment at a setting already known to be wrong, and by 70% the whole procedure has given back 24.2 points of what it bought.
Fig. 5 And on twenty runs, where a tenth is two runs and the model is exactly identified by them. The curve is flatter at the left — there is nothing below the minimum to fall off — and the same collapse at the right.

The practical rule this suggests is not split the experiment but stop guessing as soon as possible: place the smallest number of runs that make the model fittable, refit, and design everything else at what comes back. That is close to what the fully sequential design does, which is why the two are so near each other at moderate sizes.

The two-stage design’s fixed cost

The middle curve in the size picture barely moves — 90.8%, 92.7%, 92.9%, 93.0%, 93.0% across sizes from twelve to a hundred and sixty. That flatness is not a limitation of two-stage designs; it is a consequence of holding the first stage at 40% of the runs whatever the size.

A fixed fraction spent at the guess is a fixed proportional loss, so it cannot wash out with scale. A fixed number of runs spent at the guess is a loss that shrinks: four runs out of forty is a tenth, and four out of a hundred and sixty is a fortieth. The sweep in the previous section is what says four is enough, and a two-stage design that spends four runs and then commits reaches 97.1% at forty runs — above the fully sequential design at that size.

So the honest summary of the two constructions is narrower than sequential beats two-stage: almost all of the gain is in the first revision, and revising repeatedly earns a few more points on large experiments and costs a few on small ones. What matters is not how often the design is revised but how early.

How wrong the guess can be

A design that is right once, and one that is never wrong by much. Three designs for the Michaelis–Menten model, scored at every true value of K across a 16-fold range. The peaked curve is the two-point local design built at K = 1: 100% there and 66.7% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 67.9% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 3 settings, never above 80.8% and never below 78.8%. Its worst case is 12.0 points better, and what it gives up is the 21.2 points at the one value the local design was built for.
Fig. 6 For contrast, the two designs that never find out: the local design at the guess and the maximin design over a range. Their efficiency is a fixed curve, decided before the experiment starts.

The sequential design’s efficiency is not a curve over K in the same sense, because it depends on the truth only through how far the first stage has to travel. Measured at a truth of K = 3 with the guess varied: at a guess of 2 the local design is already 97.4% and the sequential one 99.4%. At a guess of 1 they are 81.1% and 96.5%. At 0.5, 56.7% and 92.8%. At 0.25, 34.6% and 90.4%.

That last row is the field’s headline. A design that reads its own data turns a twelvefold wrong guess from a two-thirds loss into a ten percent one. The maximin design over a range that included the truth would have delivered something in the seventies; the sequential design delivers ninety, and it needs no range at all — only the willingness to fit the model in the middle of the experiment.

What it does need is that the experiment can be interrupted. Every design in this field is a set of settings and a number of runs at each; a sequential design is a set of settings, a number of runs, a pause, an analysis, and a decision. In an experiment where the runs are made in one batch, none of this is available and the previous essay’s answer is the only one there is.

Updating between runs, against never updating at all. A threefold wrong guess — the design is built for K = 0.25 and the truth is 3 — and three ways of spending the runs. The flat line is the local design at the guess, 34.6% efficient at any size, because running the wrong design more times does not make it a better design. The two-stage design spends 40% of its runs there and the rest at the estimate, and lands near 79% throughout. The fully sequential design updates after every 2 runs and improves with size, from 63.2% at 12 runs to 97.7% at 160 — crossing the two-stage design where its early estimates stop being noise.
Fig. 7 The same three designs from the twelvefold wrong guess. The gap between the flat line and everything else is what the guess cost; the gap between the two rising curves is what the update frequency is worth.
One experiment finding out where to look. A single run of the fully sequential design: 80 runs, the first 8 placed at the guess K = 0.25, then the model refitted and the design revised after every 2. The marks are the settings the runs were made at. The horizontal lines are where a design built at the truth K = 3 would have put them — 1.875 and 10.00 — and the rule walks onto them without being told: its estimate of K after the first eight runs was 2.557, and by the end 2.765 against a truth of 3. The whole experiment is 95.3% as efficient as the design that knew the answer, where running all 80 at the guess would have been 34.6%.
Fig. 8 Eighty runs from the same twelvefold wrong guess. The early settings are wrong, the middle ones are approximately right, and the late ones sit on the oracle’s two lines — an experiment whose first tenth is a pilot it did not have to plan for.

What is being estimated, and what is being designed

There is a symmetry worth stating because it explains why this works at all and why the adaptive field’s version did not.

The design here reads the data through one number — the current estimate of K — and that number is a sufficient summary of everything the design needs. It does not read the responses individually, it does not chase a favourable-looking arm, and it has no incentive to. Its objective is precision about the same parameter the analysis will report, so the design’s interest and the analysis’s interest are the same interest.

In the adaptive field, an allocation rule that reads outcomes has a different objective from the analysis — it wants to treat units well, the analysis wants to estimate an effect — and the two pull against each other, which is where the broken error rates come from. Here they do not pull at all.

That is a claim about the structure and it needs a measurement, which is the next essay: whether the interval computed after a sequential design covers what it says.

What this does to the number of runs, if anything

A design that is 96.5% efficient rather than 81.1% is worth a certain number of runs, and the conversion is worth doing because it is the form the decision usually takes.

D-efficiency for a two-parameter model scales the whole information matrix, so an experiment of n runs at efficiency e carries the information of about n·e runs of the best available design. Forty runs at 81.1% is the information of thirty-two ideal runs; forty at 96.5% is the information of thirty-nine. Updating the design bought seven runs out of forty, at a cost of one interruption and one refit.

At a twelvefold wrong guess it buys twenty-two: forty runs at 34.6% carry fourteen ideal runs’ worth, and at 90.4% they carry thirty-six. That is the whole experiment over again, for the price of stopping to look at it once.

What the update recovers, as a share of what the guess cost

The four rows of the guess sweep are usually read as two columns of efficiencies. Read instead as losses — the distance each design falls short of the one that knew the answer — they say something the efficiencies do not.

The local design loses 2.6, 18.9, 43.3 and 65.4 points as the guess goes from one-and-a-half-fold wrong to twelvefold wrong. The sequential design loses 0.6, 3.5, 7.2 and 9.6. Divide the two and the share of the loss that the update recovers runs 76.9%, 81.5%, 83.4%, 85.3% — rising, in step with how wrong the guess is.

That direction is the opposite of what a hedging argument would predict, and it is the strongest thing this construction has going for it. A repair that recovers a larger fraction of a larger error is one whose worst case is bounded rather than merely improved. Averaging over a prior and protecting a range both do the reverse: they buy their protection in advance, so the further outside the anticipated range the truth turns out to be, the smaller the fraction of the damage they prevent.

The same numbers say what one more doubling of the guess error costs each design. Between successive rows the guess is wrong by twice as much again, and the local design gives up about twenty points each time — 16.3, 24.4 and 22.1 — while the sequential design gives up about three: 2.9, 3.7 and 2.4. Updating divides the marginal cost of being wrong by about seven, and it does so without anyone having to state a range, a prior or a worst case.

The floor is what the last row is approaching. Ninety per cent is roughly what remains once the first few runs have been spent at a badly wrong setting and everything after them is placed well, and the sweep is flattening towards it rather than continuing down. A guess wrong by a hundredfold would lose those same early runs and no more, which is why the first stage’s length is the quantity worth arguing about and the guess is not.

The sweep is not symmetric about its optimum

The first-stage sweep — 87.1%, 89.0%, 87.4%, 83.6%, 73.9%, 61.6% — has an interior maximum, and the useful part is that the two sides of it are not the same shape.

Measured from the peak at a tenth of the experiment, moving earlier to a twentieth costs 1.9 points. Moving later costs 1.6 points at a fifth, then 5.4 at three tenths, 15.1 at a half and 27.4 at seven tenths. One direction is bounded by how little a pilot can be and still fit a two-parameter model; the other runs all the way down to the local design’s 34.6%, since a first stage of the whole experiment is the local design by definition.

Over the last stretch that is about a point of efficiency for every percent of the experiment spent guessing, and the rate is not falling. So the two errors are not comparable, and a rule for choosing under uncertainty follows from the asymmetry alone: when unsure how long the pilot should be, make it shorter. The cost of stopping too early is one or two points and is capped by identifiability; the cost of stopping too late is unbounded within the experiment.

That is the same shape of reasoning the design that hedges applies to a guessed parameter — pick the side of the uncertainty whose penalty is the flatter one — arriving here at a quantity the experimenter controls exactly rather than at one that has to be guessed.

What is claimed, and what is not

The claim is sequential design for a non-linear model: the two-stage and fully sequential constructions, what they recover from a wrong guess, where they cross, and the interior optimum in the first stage’s length. The criteria field named sequential design as unclaimed when it was written.

What stays out: sequential designs whose stopping rule is also adaptive, which combines this with the stopping-rule field’s subject and where the error-rate question returns in force; batch sizes chosen adaptively rather than fixed; and designs that update a prior rather than a point estimate, which is the Bayesian version and would need the full posterior rather than a fitted value.

The boundary against the internal-pilot machinery in the adaptive field is that a pilot re-estimates a nuisance parameter — a variance — in order to fix the sample size, and nothing about where the runs go changes. Here the parameter being re-estimated is the one of interest and what changes is the design.

What an experimenter has to be able to do

Three requirements, and they are not always available. Naming them is the difference between a technique and a recommendation.

The runs have to be separable in time. A batch of forty tubes processed together cannot be revised in the middle. Where that is the situation, the previous essay’s design is the answer and this one is not available at any price.

The model has to be fittable on the first few runs. Four runs at two settings identify two parameters exactly, which is the minimum and is also fragile: one aberrant response moves the second stage a long way. A design that revises at four runs is buying most of the available gain and carrying most of the available risk, and revising at eight is a defensible compromise the sweep above prices at under a point.

Somebody has to do the analysis mid-experiment. That is an organisational cost rather than a statistical one, and it is the reason most experiments that could do this do not.

Against those, what is bought is the largest single improvement measured anywhere in this field: from 34.6% to 90.4% at a badly wrong guess, which no amount of hedging or worst-case protection comes close to.

The checks, and the refusal

Three claims are gated in this field’s library. A two-stage design must beat the local design at the guess by more than ten points at a threefold error, and the fully sequential one must be at least as good at forty runs. The estimate that comes out must have lower root mean squared error than the fixed design’s, which is the same information seen from the other side. And the first stage must have an interior optimum — better than both the shortest and the longest offered — because a monotone sweep would mean the trade-off asserted here does not exist.

The refusal for this half of the field is in the next essay, where the interval after a data-chosen design is measured against the interval after a fixed one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Adaptive designD-optimalityExperimental designLeast squaresLocally optimal designMichaelis–MentenMonte CarloThe non-linear modelOptimal designPilot studyTwo-stage design