A design that assumes less

The design for the worst case

A design for a non-linear model is optimal at a guess about the answer. Averaging over a prior repairs that on average; protecting the worst value in a range is a different problem, with a different answer, and it needs a third setting to reach it.

Worth reading first: The design that needs the answer · A design is a number.

The criteria field ends with a design that is optimal only if a guess is right. For a model whose information depends on the parameter being estimated — a decay rate, a half-saturation constant, any quantity that decides where the response is changing fastest — the best place to put the runs is a function of the answer. The efficiency of a design built at a guess of θ₀ when the truth is θ has a closed form in the exponential case, ρ²exp(2(1−ρ)) with ρ = θ/θ₀, and it is not symmetric: 16.5% at a threefold underestimate against 42.2% at a threefold overestimate.

That field measures one repair. Average the criterion over a prior on the parameter, and the design that comes back is better than the local one nearly everywhere and needs more settings as the prior widens. It is a good repair and it requires something an experimenter often does not have: a prior.

There is a weaker thing an experimenter usually does have, which is a range. The half-saturation constant is somewhere between a quarter and four. Nothing is claimed about where in that interval it is likely to be. A design chosen to make its worst efficiency over that range as large as possible is asking a different question from a design chosen to make its average as large as possible, and the two questions have different answers.

Three designs on one picture

The model throughout is the one the criteria field built for exactly this purpose: a response Vt/(K + t), two parameters, and a locally D-optimal design consisting of two settings — KT/(2K + T) and the end of the range — with half the runs at each. Everything below is measured against that design’s own optimum at each true K, so an efficiency of 1 means as good as it was possible to be if the answer had been known.

A design that is right once, and one that is never wrong by muchThree designs for the Michaelis–Menten model, scored at every true value of K across a 16-fold range. The peaked curve is the two-point local design built at K = 1: 100% there and 66.7% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 67.9% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 3 settings, never above 80.8% and never below 78.8%. Its worst case is 12.0 points better, and what it gives up is the 21.2 points at the one value the local design was built for.00.2500.5000.7501-0.602-0.30100.3010.602the true value of K, on a log scaleD-efficiency of the design that was actually runthe guess, K = 1worst 78.8%flat: maximin · peaked: built at the guess · between them: averaged over a priorthree designs scored at 13 true values across a factor of 16worst case 66.7%, 67.9%, 78.8%
Fig. 1 Three designs scored at thirteen true values across a sixteenfold range. The peak is the design built at K = 1. Just under it, and nearly indistinguishable, is the design that averages over a uniform prior on the same range. The flat line is the maximin design.

The local design reaches 100% where the guess is right and 66.7% at the worst point of the range. The design that averages the criterion over a uniform prior on the whole range reaches 67.9% at its worst — barely a point better — because an average is dominated by the middle of the range, where the local design is already good, and the ends contribute a thirteenth each.

The maximin design is at 78.8% at its worst and 80.8% at its best. It is never good, and it is never bad, and the flatness is not a coincidence: it is what optimising a minimum produces, and the next essay is about why.

What the flatness costs, stated where it is paid

A design whose efficiency curve is flat at 79% gives away twenty-one points at the value the local design was built for. That is the whole of the trade and it should be stated in those terms rather than as a preference for robustness.

If the guess is right, the local design is better by twenty-one points. If the guess is wrong by a factor of four, the maximin design is better by twelve. Which of those matters depends on how much the guess is worth, and that is a judgement about the situation rather than about design theory. What the arithmetic supplies is the exchange rate.

A design that is right once, and one that is never wrong by much. Three designs for the Michaelis–Menten model, scored at every true value of K across a 16-fold range. The peaked curve is the two-point local design built at K = 2: 100% there and 44.5% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 67.9% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 3 settings, never above 80.8% and never below 78.8%. Its worst case is 34.2 points better, and what it gives up is the 19.5 points at the one value the local design was built for.
Fig. 2 The same range with the local design built at K = 2 rather than 1. Its peak moves, its worst case moves with it, and the maximin design does not move at all — it is a function of the range and not of anybody’s belief about where in the range the truth sits.

That last property is the one that makes maximin usable when averaging is not. The maximin design does not ask where the truth is likely to be. It asks only what values are possible, which is a statement most experimenters will sign.

What a design is, before any of this is optimised

One thing has to be pinned down before efficiencies can be compared at all, because it is where the non-linear case differs from every design on this site that came before it.

A design here is a set of settings and the share of the runs made at each — a measure on the range of settings, exactly as in the optimality field. What is new is that the information the design carries is a function of the parameter as well as of the design, so there is no such thing as the information matrix of a design. There is a family of them, one per parameter value, and an efficiency is only meaningful once it is said which one is being scored and what it is being scored against.

The convention used throughout — and it is a convention with consequences, which the next essay is about — is standardised efficiency: the design’s information at K, divided by the information the best possible design at K would have carried, raised to the power one over the number of parameters so that it reads as a per-parameter quantity. That ratio is 1 when the design is the local optimum for K and less than 1 otherwise, so the curves in these pictures all live in the same units and can be compared across the horizontal axis.

A design that is right once, and one that is never wrong by much. Three designs for the Michaelis–Menten model, scored at every true value of K across a 32-fold range. The peaked curve is the two-point local design built at K = 1: 100% there and 56.4% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 63.4% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 3 settings, never above 80.0% and never below 76.3%. Its worst case is 19.9 points better, and what it gives up is the 23.6 points at the one value the local design was built for.
Fig. 3 The same three designs over an eightfold range, where the local design’s peak has become narrow relative to the range it is being asked about: 42.1% at the worst point, against 64.2% for the averaged design and 72.2% for the maximin one.

The range decides how much protection is available

What each design can promise, as the range it must cover widens. Three designs for the same two-parameter model, each scored at thirteen true values spanning the range it was built for, and the worst of those thirteen plotted. The local design is the two-point optimum at K = 1: it is 100% efficient there and 96.6% over a factor of 1.5, 90.2% over a factor of 2, 66.7% over a factor of 4, 42.1% over a factor of 8, 23.9% over a factor of 16. The averaged design holds up far better than the local one past a factor of four, and the maximin design — which maximises exactly this quantity — is above both everywhere, by 10.9 points at a factor of 4. All three use two settings on the narrowest range; the robust pair buy a third at a factor of 4.
Fig. 4 The worst case each of the three designs can be promised, as the range it has to cover widens from a factor of one and a half to a factor of sixteen either side of the guess.

The table this draws is the field’s summary and it is worth reading as one:

At a factor of 1.5, all three designs are above 96% and there is nothing to discuss. At 2, they are 90.2%, 90.1% and 90.8%. At 4, the local design is at 66.7%, the averaged one at 67.9%, and the maximin one at 78.8%. At 8: 42.1%, 64.2%, 72.2%. At 16: 23.9%, 59.8%, 64.6%.

Two things happen across that sweep and only one of them is obvious. No design can hold its worst case as the range widens — the maximin curve falls like everything else, because a design covering a sixteenfold range has to spend runs at settings that are wrong for most of it. And the averaged design separates from the local one long before the maximin design separates from either: at a factor of eight, averaging has already bought twenty-two points of worst case and maximin only eight more.

So the case for maximin over averaging is narrower than the case for either over a local design. That is worth saying because the two robust answers are usually presented as rivals, and the measurement says they mostly agree, and that where they disagree the disagreement is worth a few points rather than a factor.

Two settings cannot do it

The local design for this model uses two settings at every guess. The averaged design uses two until the prior is wide enough to force a third. The maximin design uses three from a factor of four onwards, and the third setting is not a detail.

Where the runs go, when the answer is not known in advance. The two designs as they would be run. The local design at K = 1 puts half its runs at 0.833 and half at 10.00, which is KT/(2K + T) and the end of the range, both from a closed form. The maximin design over K from 0.25 to 4 uses 3 settings: 23.4% at t = 0.298, 32.2% at t = 2.029, 44.4% at t = 9.992. The extra setting is the whole of the protection — a two-point design cannot cover the range however its two points are placed, because the point that is right for a small K is wrong for a large one and there is nowhere in between that is right for both.
Fig. 5 Where the runs go. The local design at K = 1 puts half at 0.833 and half at the end of the range. The maximin design puts 23.4% at 0.298, 32.2% at 2.029 and 44.4% at 9.992 — three settings, unequal weights, and neither of the local design’s two among them.

Run the same search restricted to two settings, allowing them anywhere and in any proportion, and the best worst case available is 69.5%. Allow a third and it is 78.8%. Nine points, bought by a setting rather than by a better placement of the two.

The reason is structural rather than numerical. A small K puts the informative setting early and a large K puts it late; there is no single early-or-late setting that is right for both, because the optimum moves monotonically with K and the criterion is not flat between. A design that has to protect both ends has to have runs at both ends, and that takes a point the two-setting design does not have to spare.

Where the runs go, when the answer is not known in advance. The two designs as they would be run. The local design at K = 1 puts half its runs at 0.833 and half at 10.00, which is KT/(2K + T) and the end of the range, both from a closed form. The maximin design over K from 0.25 to 16 uses 3 settings: 23.6% at t = 0.396, 32.8% at t = 3.002, 43.6% at t = 10.000. The extra setting is the whole of the protection — a two-point design cannot cover the range however its two points are placed, because the point that is right for a small K is wrong for a large one and there is nowhere in between that is right for both.
Fig. 6 And across a wider range, where the same three settings spread out to cover it. The count does not keep growing with the range in the way the averaged design’s support count does — three is enough to tie the worst case at both ends, and a fourth setting has nothing left to do.

The arithmetic of a wrong guess, which is what all of this is against

It is worth having the size of the problem in mind before deciding what to spend on it, and the criteria field supplies it in closed form. For the exponential model an efficiency at ratio ρ = θ/θ₀ is ρ²exp(2(1−ρ)), which is 16.5% at ρ = 3 and 42.2% at ρ = ⅓. For the model used here the same asymmetry appears without a closed form: at a guess of K = 1 the local design is 66.7% efficient when the truth is a quarter of the guess and 73.2% when it is four times it.

Both statements say the same thing and neither says it in the same direction, which is the part worth carrying: a wrong guess is not paid for symmetrically, the expensive direction is a property of the model rather than a general rule, and it has to be worked out rather than assumed. For the exponential model the expensive error is guessing the rate too low; for this one, at this range, it is guessing the constant too high. A designer who cannot say which side their guess is likely to fall on has no way to exploit either fact, and that is precisely the designer maximin is for.

A wider prior buys another setting, and the arithmetic says when. The number of distinct settings in the design that maximises the average of log|M| over a prior on the unknown K, against how wide that prior is. It is a staircase because the answer is an integer: a prior reaching a factor of three either side of the guess is still answered by the two settings a local design uses, the third arrives at a spread of 3.36 and the fourth at a spread of 8.86. Neither threshold was put in — both are found by bisecting on the design the algorithm returns. This is what "hedge the guess" means concretely: runs have to be spent at settings that would be right if the guess were wrong, and a two-point design has nowhere to put them.
Fig. 7 The averaged design’s answer to the same problem, from the criteria field: how many settings the design needs as the prior widens. The support count is an integer that steps at particular widths, and the maximin design’s three is the same phenomenon reached from the other side.

The averaged design, for comparison, on its own terms

It would be unfair to leave the averaged design measured only by a criterion it was not built for. Its own criterion is the mean, and on the mean it wins: 87.3% against the maximin design’s 79.7% over the sixteenfold range drawn first. That is not a small margin and it is the right comparison for anybody who can honestly write down a prior.

The distinction is the same one the interval fields make about coverage: a procedure can be right on average over a distribution of situations, or right in every situation, and those are different requirements that happen to coincide when the situation is narrow enough. A prior is what turns a range into a distribution, and if the prior is real then averaging is the better answer. Maximin is what to do when the prior would be invented.

Where maximin’s advantage actually lives

The width sweep is read above as two separate observations — no design holds its worst case, and the averaged design separates from the local one before maximin separates from either — and differencing the three columns makes the second one sharper than it looks.

Maximin’s margin over the averaged design runs 0.7, 10.9, 8.0 and 4.8 points at widths of two, four, eight and sixteen. It is not monotone: it peaks at a fourfold range and falls away on both sides. Averaging’s margin over the local design, by contrast, runs −0.1, 1.2, 22.1 and 35.9 — monotone and still climbing at the widest range measured.

So the two robust answers are not two points on one scale. Averaging is worth more the wider the range gets, and maximin is worth most in the middle. At a factor of two there is nothing to protect and all three designs agree; at a factor of sixteen the averaged design has itself been pushed onto enough settings to be nearly flat, and the specifically minimax construction is buying five points rather than eleven.

That locates the case for this field’s construction precisely, and it is narrower than the field’s own opening suggests. A range of about a factor of four either side of the guess is where a maximin design earns its extra setting and its twenty-one points at the guess; outside that band an experimenter who is willing to write down a uniform prior gets most of the same protection from machinery the criteria field already had.

The exchange rate at the far end is worth stating too, since both means are available there. Over the sixteenfold range the averaged design’s mean efficiency is 87.3% against maximin’s 79.7%, and its worst case is 59.8% against 64.6%. Maximin buys 4.8 points of worst case for 7.6 points of average — an exchange of about one and a half to one, which is a defensible trade for a promise and a poor one for an expectation.

Two models, two asymmetries, opposite in sign and tenfold in size

The essay says a wrong guess is not paid for symmetrically and that the expensive direction has to be worked out rather than assumed. Both models it quotes are available, so the claim can be given a size.

For the exponential model the closed form gives 16.5% when the truth is three times the guess and 42.2% when it is a third of it — a ratio of 2.56 between the two directions. For the saturating model used here, a local design at K = 1 is 66.7% efficient at a quarter of the guess and 73.2% at four times it, a ratio of 1.10.

Two two-parameter non-linear models, and the asymmetry of a wrong guess differs between them by more than an order of magnitude: 156% against 10%. And it points the other way. The exponential model punishes guessing the rate too low; this one punishes guessing the constant too high.

That is what makes the asymmetry unusable as a general rule and useful as a per-model calculation. A designer who knows which side their guess is likely to fall on can exploit it — worth a factor of two and a half in one model and almost nothing in the other — and a designer who does not know cannot even tell which side to lean towards without doing the arithmetic for their own model first.

Where the extra runs come from

A design has a fixed number of runs, so a third setting is paid for by the two that were there. The maximin design’s weights — 23.4%, 32.2%, 44.4% — are not equal, and the ordering is the one the arithmetic of the model forces: the end of the range carries the most because it is informative at every K, and the early setting carries the least because it is informative only if K is small.

That is a general feature of protection rather than a fact about this model. The runs a design spends on being wrong-proof are spent where they would be wasted if the guess were right, so they come out of the settings the local design was confident about. An experimenter who cannot afford to give up twenty-one points at the guess is an experimenter who should not buy the protection, and the picture above is what they are declining rather than an argument that they should.

A design that is right once, and one that is never wrong by much. Three designs for the Michaelis–Menten model, scored at every true value of K across a 64-fold range. The peaked curve is the two-point local design built at K = 1: 100% there and 44.9% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 62.3% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 4 settings, never above 80.1% and never below 74.5%. Its worst case is 29.6 points better, and what it gives up is the 24.2 points at the one value the local design was built for.
Fig. 8 The sixteenfold range, where the local design has fallen to 23.9% at its worst and the maximin design holds 64.6%. This is the width at which the two robust designs part company visibly: the averaged one is at 59.8%, five points below, and needs a prior to be written down before it can be computed at all.

What is claimed, and what is not

This field takes maximin-efficient design: the design that maximises the worst standardised efficiency over a stated range, the settings it needs, what it gives up at the guess, and how the comparison against a local and an averaged design moves as the range widens. The criteria field named maximin design as the E-optimality of the non-linear half and did not build it.

What stays out and is named: maximin designs for subsets of the parameters, which combine this essay’s problem with Ds-optimality’s; maximin over a range of models rather than of parameter values, which is a different robustness and needs a different standardisation; and the whole sequential alternative, which is the second half of this field and answers the same complaint by not guessing at all.

The boundary against the criteria field is that it owns the family Φₚ and the local and averaged designs, and this owns the minimax problem over parameter values. They meet at the observation that both are optimisations of a non-differentiable objective, which is the next essay.

What this does not promise

Two limits are worth stating so that a flat efficiency curve is not read as more than it is.

The range is an assumption and it is the only one that matters. A maximin design over a quarter-to-four range says nothing at all about a truth of eight, and its efficiency there is not protected by anything: at K = 8 the design drawn above falls to well under its own worst case. The protection is exactly as good as the statement of what is possible, which is a smaller assumption than a prior and is still an assumption.

Standardised efficiency is a ratio and not an amount. A design that is 79% efficient everywhere is 79% of the best available at each K, and the best available at a large K may itself be poor. What this field optimises is the relative loss from not knowing the answer, which is the only part a designer controls; the absolute precision is decided by the model, the noise and the number of runs.

The same words, the wrong quantity. Both curves are designs that maximise a worst case. The flat one maximises the worst efficiency — how good the design is at K, relative to the best design there is at K. The rising one maximises the worst determinant of the information matrix, which sounds like the same thing and is not: information at one parameter value is not comparable with information at another, the determinant is smallest at K = 4 whatever any design does, and the search therefore ends up placing every run where a design for K = 4 would place it. Scored on efficiency it runs from 29.0% to 100.0%, against a flat 78.7% and above. A robust design has to be told what it is being robust about.
Fig. 9 The next essay’s refusal, for the reader who wants to know now what the standardisation is doing: the same search run on the worst determinant rather than the worst efficiency, and what it returns.

The checks, and the refusal

Four claims are gated in this field’s library. The maximin design’s worst case must exceed the local design’s by more than ten points at the range drawn, and must be at or above the averaged design’s at every width measured. The local design must still be the best thing to run at the guess, which is what makes the trade a trade. The worst case must fall as the range widens, so that no design is claimed to protect an arbitrary range for free. And a two-setting maximin search must fall at least five points short of a free one, which is the third setting’s value asserted rather than described.

The refusal is a design chosen by maximising the worst determinant instead of the worst efficiency — the same words, the wrong quantity — and it is the next essay’s subject.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • An efficiency that is a ratio — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, locally optimal design, michaelis–menten, the non-linear model, optimal design
  • Augmenting a design that has already run — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design
  • The criterion with no derivative — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design
  • The family behind the letters — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design
  • The guess with two numbers in it — both name d-optimality, design measure, equivalence theorem, information matrix, the non-linear model, optimal design
  • The theorem that says when to stop — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design

Named objects

A flat tag is an object no other essay names yet.

Bayesian optimal designD-optimalityDesign measureEquivalence theoremExperimental designInformation matrixLocally optimal designMaximin designMichaelis–MentenThe non-linear modelOptimal designPrior