The criterion, and what it assumes

The design that needs the answer

Every design this site has computed is optimal whatever the experiment turns out to say, because X′X does not contain the parameters. For a non-linear model it does, so the best place to take a measurement is a function of the number the measurement exists to find — and guessing it three times too low costs two and a half times more than guessing it three times too high.

Worth reading first: A design is a number.

Every design in the optimality field is optimal before any data exists, and the reason is one line of algebra that is so ordinary it is never pointed at. The information matrix is

M = X′X / N

and X is built from the settings — the factor levels, their products, their squares. The parameters are not in it. So a design can be chosen, defended, catalogued and published without anybody knowing what the experiment is going to say, and every number in that field is a fact about an arrangement rather than about a response.

That is a property of linear models, and it is a strong one. Break it and the whole field changes character.

Where to look depends on the answer. The information a single run at time t carries about the rate of an exponential decay, (∂η/∂θ)² = t²·exp(−2θt), at three values of θ. Each curve has one maximum and it is at t = 1/θ exactly — marked, and found by a search over 8,001 settings that was never told the formula. Nothing in a linear model behaves this way: there the information matrix is X′X and the parameters are not in it, so a design can be chosen once and used whatever the answer turns out to be. Here the design is optimal at a guess, and the three curves are three different experiments for one model.
Fig. 1 The information a single run carries about the decay rate of an exponential, at three different true rates. Each curve has one maximum and each maximum is somewhere different. There is no design here that is good whatever the answer is.

One parameter, one run, one closed form

Take the simplest non-linear model there is: an exponential decay, η(t) = exp(−θt), with one unknown rate. A run at time t contributes information equal to the squared derivative of the response with respect to the parameter:

(∂η/∂θ)² = t²·exp(−2θt)

That expression has θ in it. Differentiate and set to zero and the maximum is at

t = 1/θ

exactly, and at nothing else. The D-optimal design for one parameter is a single point — there is nothing else to spend runs on — and that point is the reciprocal of the number the experiment exists to measure.

The two limits say why. At t near zero nothing has decayed yet, so the response is one whatever θ is and the run says nothing. At t large everything has decayed, the response is zero whatever θ is, and the run says nothing again. The information is squeezed between two kinds of ignorance and peaks at the one time when the response is changing fastest relative to what is left of it.

The closed form is checked against a search that was never told it: a grid of eight thousand settings, at four different rates, each maximum required to land on 1/θ to within one grid step. That is the site’s usual arrangement — a formula and a search that share no arithmetic — and it matters more than usual here, because everything below is a statement about how wrong the formula’s input can be.

What a wrong guess costs, in closed form

A design has to be built before the data exists, so the experimenter supplies a guess θ₀ and runs at 1/θ₀. What that design is worth when the truth is θ is a ratio of two evaluations of one function, and it simplifies to something with no design left in it at all:

eff(ρ) = ρ²·exp(2(1−ρ)), where ρ = θ/θ₀

One at ρ = 1, falling away on both sides. And not at the same rate.

  • A threefold under-estimate of the rate leaves 16.5% of the information.
  • A threefold over-estimate leaves 42.2%.
  • Fourfold: 4.0% against 28.0%.
What a wrong guess costs, and which direction is the cheap one. The design for the exponential model is a single run at t = 1/θ, so the efficiency of a design built at a guess θ₀ and used where the truth is θ has a closed form with no design in it: ρ²·exp(2(1−ρ)) at ρ = θ/θ₀. It is 1 at ρ = 1 and falls away on both sides, and not at the same rate — 16.5% at a threefold underestimate of the rate against 42.2% at a threefold overestimate. Underestimating the rate means measuring too late, which is the direction that feels cautious, and it is the expensive one: information about a decay is destroyed by the decay itself, so a run placed past the optimum is measuring something that has mostly happened.
Fig. 2 The efficiency curve, with reciprocal errors marked. The left half is guessing the rate too high and the right half too low, and the two halves are not mirror images. A factor of four in the safe direction costs 72 points; the same factor the other way costs 96.

The direction is the surprising one and it is worth stating as advice, because it inverts what caution suggests. Guessing the rate too low means assuming the process is slower than it is, which means measuring later — and measuring late is the expensive mistake, because information about a decay is destroyed by the decay. A run placed past the optimum is measuring a quantity that has mostly already happened, and the exponential in the denominator punishes it exponentially. Measuring early is wasteful and survivable: nothing has decayed, the run is uninformative, and it is uninformative in a way that only costs the square of a small number.

So the cautious-sounding instruction — leave plenty of time, the process might be slower than it looks — is the one that destroys the experiment.

The design moves with the guess, and the picture says how much

The single point at 1/θ is the whole design, so the whole design is one number and a reader can watch it move.

Where to look depends on the answerThe information a single run at time t carries about the rate of an exponential decay, (∂η/∂θ)² = t²·exp(−2θt), at three values of θ. Each curve has one maximum and it is at t = 1/θ exactly — marked, and found by a search over 8,001 settings that was never told the formula. Nothing in a linear model behaves this way: there the information matrix is X′X and the parameters are not in it, so a design can be chosen once and used whatever the answer turns out to be. Here the design is optimal at a guess, and the three curves are three different experiments for one model.00.0500.1000.1500246when the run is takeninformation one run carries about θθ = 1, t = 1.00θ = 2, t = 0.50θ = 4, t = 0.25t²·exp(−2θt) at θ = 1, 2, 4, over 241 settingseach maximum at 1/θ, searched rather than substituted
Fig. 3 The same information function at three faster rates. The peaks are at 1, ½ and ¼, and the peak heights fall as the square of the reciprocal rate — a fast process supplies less information per run than a slow one, whatever time it is measured at, because there is less of it left to observe. Drag the rate the design is built for and watch the single optimal setting move.

That the peak height falls with θ is worth a sentence, because it is a separate fact from where the peak is and it changes what a design is worth rather than what it is. The maximum of t²exp(−2θt) is 1/(θ²e²), so a process twice as fast supplies a quarter of the information per run at its own best time. An experimenter facing a fast decay needs four times as many runs as one facing a decay half the speed, and no arrangement of settings recovers any of it.

Two parameters, two points, and the same problem

One parameter is a clean illustration and not an experiment anybody runs. The two-parameter case is, and it has the same shape with one extra piece of luck.

The Michaelis–Menten model, η(t) = V·t/(K + t), rises from zero to an asymptote V with a half-way point at K. Its gradient is (t/(K+t), −V·t/(K+t)²), and the second component is where V and K both appear.

Run the same multiplicative algorithm over a grid of times on (0, T] and it returns a design with two support points at half the runs each, at

t = KT/(2K + T) and t = T

which is a closed form again, and again a function of the unknown. At K = 1 on a range to 10 the lower point is 0.8333; at K = 0.5 it is 0.4545; at K = 2 it is 1.4286; at K = 5 it is 2.5000. The search finds all four to within a grid step, having been told nothing but the criterion.

The piece of luck is that V does not appear in either setting. It scales the second gradient component and therefore the determinant, and drops out of the argmax entirely. So the guess that has to be made is a single number, which is the smallest version of this problem that is not the one-parameter one.

The upper point sitting at the end of the range is worth noticing rather than passing: the design wants one run where the asymptote is best seen, and the best available place for that is as far right as the experiment is allowed to go. That makes T a design decision with the same standing as the settings — extending the range changes the answer, and the closed form says by how much.

Where to look depends on the answer. The information a single run at time t carries about the rate of an exponential decay, (∂η/∂θ)² = t²·exp(−2θt), at three values of θ. Each curve has one maximum and it is at t = 1/θ exactly — marked, and found by a search over 8,001 settings that was never told the formula. Nothing in a linear model behaves this way: there the information matrix is X′X and the parameters are not in it, so a design can be chosen once and used whatever the answer turns out to be. Here the design is optimal at a guess, and the three curves are three different experiments for one model.
Fig. 4 The information function at three slower rates, on a longer range. The optima are at 4, 2 and 1, and the peak heights are correspondingly larger — a slow process is worth more per run than a fast one, at its own best time.

The theorem still works, which is not obvious

One thing survives the move from linear to non-linear models and it is worth checking rather than assuming, because most of the machinery does not.

The general equivalence theorem — the search is finished exactly when the largest directional derivative equals the number of parameters — was derived for a criterion that is a concave function of the information matrix. Local optimality holds the parameters fixed while it optimises, so at any one guess the problem is the linear problem with a different X, and the theorem applies unchanged. The Michaelis–Menten search stops when max δ reaches 2, to twelve decimal places, exactly as the quadratic model’s search stops at 6.

What does not survive is everything between guesses. Two locally optimal designs at two values of K are each certified optimal, and there is no sense in which one of them is better than the other — they are answers to different questions, and comparing them requires nominating a truth, which is the step the refusal below is about.

The variance touches p and never crosses it. d(x) = f(x)′M⁻¹f(x) along the diagonal of a square region, for the D-optimal measure. The line at 6 is the number of parameters in the model. Kiefer and Wolfowitz's theorem says a design is D-optimal exactly when the largest d anywhere in the region is p — not approximately, equals — so the optimal curve is tangent to that line at its support points and below it everywhere else. Here the largest value anywhere on a 41×41 grid is 6.000000000.
Fig. 5 The theorem drawn in the linear case, from the optimality field: the prediction variance under the optimal measure touches the number of parameters and never exceeds it. The non-linear case produces the identical picture at any fixed guess — and a different one at every guess, which is the whole difficulty stated as a figure that would have to be redrawn for each value of K.

The efficiency of a guess, measured rather than derived

For the two-parameter model there is no closed form for the efficiency of a wrong guess, so it is computed: build the design at K = 1, score it at every true K against the local optimum there.

  • at K = 0.5, 90.2%
  • at K = 2, 91.3%
  • at K = 0.25, 66.7%
  • at K = 4, 73.2%
  • at K = 8, 56.4%

A design built for the right order of magnitude is fine. A design built for a guess that is out by a factor of four has lost between a quarter and a third of its information, and by a factor of eight nearly half.

A design that is right once, against one that is never wrong by much. Two designs for the same two-parameter model, scored at every true value of K across a range of 16-fold. The local design is the two-point optimum for a guess of K = 1: it reaches 100% there and 66.7% at the worst point of the range. The hedged design maximises the average of log|M| over a prior spanning a factor of 4 either side, uses 3 settings rather than two, and is never below 75.4%. What it costs is 10.4 points at the one value the local design was built for — which is the whole trade, and it is only available to somebody willing to say how wrong the guess might be.
Fig. 6 The peaked curve is that design, scored across a range of true K. It reaches 100% at exactly one value and falls away on both sides — which is what “locally optimal” means, stated as a picture. The flatter curve is what the next essay does about it.

The refusal, and it is a habit rather than a mistake

There is a way of scoring these designs that produces an answer of exactly 1 everywhere, and it is what a linear-model habit produces without anybody deciding anything.

Take the design optimal at the guess. Score it — against the optimum at the guess. In a linear model that is the only optimum there is, so the phrase “the design’s efficiency” is unambiguous and the answer is the design’s efficiency. Here there is one optimum per value of K, and scoring against the guess’s optimum reports 1.0000 at K = 0.25, 1.0000 at K = 1 and 1.0000 at K = 4 — a design that is perfect everywhere, which is the one answer that cannot be true.

The refusal in this field’s library makes exactly that call and requires the honest score to disagree. It is not a mistake anybody would defend once it is pointed at; it is a mistake that is invisible because the wrong quantity and the right quantity have the same name.

The same shape appears in the allocation field, and the two are worth reading together. There, a rule for splitting units between two arms is a function of the two spreads nobody has, and plugging in estimates makes the experiment 12.0% worse than not bothering at small pilot sizes. Here a rule for choosing settings is a function of a rate nobody has. Both fields end at the same place from different directions: an optimality result whose input is the answer is a conditional statement, and the condition is the thing being measured.

The one structural difference is worth naming, because it decides which repair is available. In the allocation field the unknown can be estimated from the experiment itself — a pilot gives an estimate of the spreads, at a price the field measures. Here the unknown is what the experiment is for, and there is no pilot that supplies it without being an experiment. What is available instead is a prior: a statement about how wrong the guess might be rather than a better guess. That is the whole of the next essay.

A wider prior buys another setting, and the arithmetic says when. The number of distinct settings in the design that maximises the average of log|M| over a prior on the unknown K, against how wide that prior is. It is a staircase because the answer is an integer: a prior reaching a factor of three either side of the guess is still answered by the two settings a local design uses, the third arrives at a spread of 3.36 and the fourth at a spread of 8.86. Neither threshold was put in — both are found by bisecting on the design the algorithm returns. This is what "hedge the guess" means concretely: runs have to be spent at settings that would be right if the guess were wrong, and a two-point design has nowhere to put them.
Fig. 7 What the repair looks like before it is explained: the number of settings a design has to use once the guess is replaced by a range. Two while the range is narrow, and then an integer that steps.

The asymmetry is a tail phenomenon

The closed form makes the asymmetry computable at every ratio rather than at the two the essay quotes, and doing so says something the two do not: near the guess there is no asymmetry at all.

Divide the two directions. eff(1/r)/eff® = r⁻⁴·exp(2(r − 1/r)), which is exactly 1 at r = 1 and grows from there. At a factor of 1.5 it is 1.05; at 2, 1.26; at 3, 2.56; at 4, 7.06. So the penalty for a mis-guess is essentially symmetric in the ratio until the ratio is about two, and then the two directions separate very fast — the asymmetry roughly squares with each further factor.

The reason is in the logarithm. log eff = 2 log ρ + 2(1 − ρ), whose second-order expansion about ρ = 1 is −(ρ − 1)² — an even function, so the leading behaviour cannot distinguish the two directions. The asymmetry lives entirely in the third and higher terms, which is why it is invisible inside a factor of 1.5 and dominant past a factor of three.

That gives the practical threshold a sharper form than the essay’s “about a factor of 1.5”. Solving ρ²exp(2(1 − ρ)) = 0.95 puts the 95%-efficient band at ρ from 0.79 to 1.24 — a factor of 1.27 either way, near enough symmetric. Inside that band the direction of the guess does not matter and neither does very much else. Outside a factor of three it is the only thing that matters.

Two models, and a seventy-fold difference in the asymmetry

The two-parameter model’s efficiency curve is measured rather than derived, and setting its numbers beside the exponential’s says how little the closed form transports.

At a fourfold error the Michaelis–Menten design built at K = 1 is 66.7% efficient when the truth is a quarter of the guess and 73.2% when it is four times it — a ratio of 1.10. The exponential model at the same fourfold error gives 4.0% against 28.0%, a ratio of 7.06.

Same kind of failure, two two-parameter-scale problems, and the asymmetry differs by a factor of seventy. The direction agrees — both punish the guess that puts the measurement too late relative to the process — and nothing else about the two curves is comparable: one falls to 4% at a fourfold error and the other to 67%.

So the exponential’s closed form is a worked example rather than a template. What survives from it is the structure — an efficiency that is a function of the ratio alone, unity at the guess, falling asymmetrically — and the numbers do not, in either their size or their spread. A designer who has read the closed form knows what shape of curve to compute and has not been told any of its values, which is the honest statement of what a one-parameter result buys for a two-parameter problem and is why the second curve is counted rather than quoted.

Where an experimenter’s guess comes from

The efficiency curves above are useless without some sense of how wrong a guess typically is, and that is not a statistical question — it is a question about a subject.

Three sources are ordinary. A previous experiment on the same system, which gives a guess with a standard error attached and is the best case. A related system — the same reaction with a different substrate, the same drug in a different species — which gives an order of magnitude and no standard error. And a range from theory, which often gives bounds rather than a value.

The efficiencies say what each is worth. A previous experiment that pinned the rate to within twenty per cent leaves the design above 96% efficient and there is nothing to discuss. An order of magnitude — a guess that might be out by a factor of three either way — leaves it somewhere between 16.5% and 42.2%, which is where the design decision starts to matter more than anything else in the experiment. Bounds from theory are the case the next essay is for, because bounds are a prior.

What a wrong guess costs, and which direction is the cheap one. The design for the exponential model is a single run at t = 1/θ, so the efficiency of a design built at a guess θ₀ and used where the truth is θ has a closed form with no design in it: ρ²·exp(2(1−ρ)) at ρ = θ/θ₀. It is 1 at ρ = 1 and falls away on both sides, and not at the same rate — 16.5% at a threefold underestimate of the rate against 42.2% at a threefold overestimate. Underestimating the rate means measuring too late, which is the direction that feels cautious, and it is the expensive one: information about a decay is destroyed by the decay itself, so a run placed past the optimum is measuring something that has mostly happened.
Fig. 8 The same efficiency curve on a wider range, which is the picture to read a guess against. The practical reading is a threshold rather than a number: inside about a factor of 1.5 the design is fine, and outside about a factor of three it is not a design.

It is worth noticing what this does to the usual order of operations. An experiment is normally planned by choosing a sample size for a stated effect, which is the design field’s own calculation and is a statement about power. Here there is a decision before that one, it is a choice of settings rather than of a count, and being wrong about it cannot be fixed by running more units: a design at 16.5% efficiency needs six times the runs to recover, and no power calculation carried out at the wrong settings will say so.

What a wrong guess costs, and which direction is the cheap one. The design for the exponential model is a single run at t = 1/θ, so the efficiency of a design built at a guess θ₀ and used where the truth is θ has a closed form with no design in it: ρ²·exp(2(1−ρ)) at ρ = θ/θ₀. It is 1 at ρ = 1 and falls away on both sides, and not at the same rate — 16.5% at a threefold underestimate of the rate against 42.2% at a threefold overestimate. Underestimating the rate means measuring too late, which is the direction that feels cautious, and it is the expensive one: information about a decay is destroyed by the decay itself, so a run placed past the optimum is measuring something that has mostly happened.
Fig. 9 The efficiency curve once more, on a range that reaches a fivefold error either way. Everything to the right of about ρ = 2 is a design that has lost more than half its information, and the curve gets there faster on that side than on the other.

What is being claimed here, and what is not

This essay claims local D-optimality for two non-linear models: the information function, the closed-form optima, the closed-form efficiency for the one-parameter case, and the counted efficiency curve for the two-parameter one.

What stays out: sequential design, where the guess is updated as the data arrives and the next run is placed at the current estimate — that is a genuinely different construction with its own error-rate questions, and it is the obvious neighbour of everything here; non-linear models with three or more parameters, where the number of support points grows and the closed forms stop; and any claim about the estimation of these models, which is a numerical problem this site does not enter.

The boundary against the response-surface field is worth stating because it looks like an overlap. A quadratic is a non-linear function of the settings and a linear function of the parameters, which is why its information matrix is free of them and why that field could catalogue designs. The models here are non-linear in the parameters, which is the distinction the whole subject turns on and is lost by the word “non-linear” being used for both.

The checks

Two claims are gated in this field’s library.

The optima are where the closed forms say, at four rates for the exponential model and four values of K for Michaelis–Menten, with the two-point support and the half-and-half weights required as well — because a grid search does not have to return two points carrying 0.5 each, and that it does is the theory being confirmed rather than assumed.

And the efficiency of a wrong guess matches ρ²exp(2(1−ρ)), checked against the design search at four values of ρ, with the asymmetry asserted as an inequality: a threefold underestimate has to cost more than twice what a threefold overestimate costs. Asserting the direction rather than only the values is what stops a future change that flattened the curve from passing.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Bayesian optimal designD-optimalityEquivalence theoremExperimental designInformation matrixLocally optimal designMichaelis–MentenThe non-linear modelOptimal designPlug in estimate