What the design is asked to guarantee

Protecting one parameter over a range

A design for a non-linear model is optimal at a guess. A design for one of its parameters over a range of guesses is a worst case of a ratio of two determinants, and it is not a special case of either problem it is made of.

Worth reading first: The design that needs the answer · A design is a number.

Two repairs to the same complaint exist and this field has to compose them.

The complaint is that a design for a non-linear model is optimal only at a guess about the answer. The robust field’s repair is to protect the worst parameter value in a range instead of one value, which turns the design problem into a maximin problem whose optimum sits on a tie. The previous essay’s repair is to design for the parameter somebody wanted instead of the whole vector, which turns the criterion into a ratio of two determinants.

Composing them gives a worst case, over a range, of an efficiency that is itself a ratio. That is not a special case of either, and the four designs it produces disagree about where the runs go by enough that a design described as “robust” has said nothing until it says robust about what.

The four designs

Fix a range: K anywhere between 0.25 and 4, a factor of sixteen. Four designs are available.

Two are local, built at the middle of the range as though the guess were right: the D-optimal design for the pair, and the closed-form Ds-optimal design for K. Two are maximin, built to maximise the worst efficiency over the whole range: one for the pair, from the robust field, and one for K, which is what this essay is about.

Four designs, scored on the one parameter that was wanted. Every design scored by its Ds-efficiency for K at 13 true values across a 16-fold range. The peaked curve is the subset design built at the guess K = 1: 100% there and 42.1% at the worst point of the range. The flat curve is the maximin-Ds design, never above 64.8% and never below 61.2%. Between them is the maximin design for the pair — a robust design, protecting something else, and worth 41.9% at worst here. The lowest curve is the D-optimal design at the guess, which is what an experimenter who wanted K and looked up a design for the model would actually run: 29.1% at the worst point, against 61.2% available.
Fig. 1 All four scored on the one criterion the experimenter asked about, at thirteen true values across the range. The peaked curves are the two local designs; the flat one is the maximin design for K.

Scored on the subset criterion, their worst cases over the range are 29.1%, 42.1%, 41.9% and 61.2%. The ordering is worth reading slowly.

The design an experimenter who wants K would actually run — look up the D-optimal design for the model, use the best guess — is the worst of the four, at 29.1%. It loses on both counts at once: it is built for the wrong question and at a single value.

Designing for the right question at a single value gets to 42.1%. Designing for the wrong question across the whole range gets to 41.9%. Those two repairs are worth the same amount, and an experimenter who applied either one alone would land in the same place by two completely different routes.

Applying both reaches 61.2%.

Robust about what

The reverse scoring is what makes the composition a third problem rather than a refinement.

Robust about what?. Four designs over a range of K from 0.25 to 4, each shown by its worst case on both criteria: the bar is the worst efficiency for the half-saturation constant alone, the mark beside it the worst efficiency for the pair. The maximin design for the subset holds 61.2% on the subset and 58.7% on the pair; the maximin design for the pair holds 78.7% on the pair and 41.9% on the subset. Each is about twenty points better on its own criterion and about twenty points worse on the other, so a design described as robust over this range has said nothing until it says which of the two it means.
Fig. 2 The same four designs, each shown by its worst case on both criteria: the bar is the guarantee for K alone, the rule beside it the guarantee for the pair.

The maximin design for the pair holds 78.7% on the pair and 41.9% on the subset. The maximin design for the subset holds 61.2% on the subset and 58.7% on the pair. Each is about twenty points better on its own criterion and about twenty points worse on the other.

So “this design is maximin-efficient over a sixteenfold range of K” is not a statement until the criterion is named. Both designs satisfy that sentence; they guarantee different things, they put their runs in different places, and neither one’s guarantee transfers.

That is the same lesson the criterion field records for the letters — a design that wins on D is nearly worst on E — arriving in a place where it is easier to miss, because “robust” sounds like a property of a design and “D-optimal” obviously does not.

Two repairs worth the same, and why that is not a coincidence

The coincidence in the middle of that table is worth an explanation rather than a shrug. Designing for the right question at one value is worth 42.1%; designing for the wrong question over the whole range is worth 41.9%. Two entirely different repairs, one number.

They are not the same repair in disguise. The first fixes the criterion and leaves the design local, so it is excellent at K = 1 — 100%, by construction — and collapses at the ends of the range. The second fixes the range and leaves the criterion wrong, so it is mediocre everywhere: 51.9% at K = 1, which is worse there than the design built at the guess for the wrong question.

What the two share is that each has one thing right and one thing wrong, and over a sixteenfold range those two errors happen to cost about the same. Widen the range and they separate — at a factor of eight the local design for the subset is at 16.3% and the maximin design for the pair at 33.4% — because a wrong criterion costs a roughly constant amount and a wrong value costs more the wider the range gets.

So the equality is a fact about this range rather than a law, and it is the kind of fact worth writing down because it is what makes the two repairs look interchangeable to anybody who measures at one width.

The optimum is a tie, and a second route says the same

The maximin design for the subset uses three settings, at t = 0.2589, 1.6565 and 9.9949, weighted 0.4205, 0.3412 and 0.2384. Its efficiency across the range is flat to within 3.59 points, and its minimum, 61.19%, is attained at four true values including both ends: K = 0.2500, 0.8909, 1.0000 and 4.0000.

The worst case is a tie, and the prior says where. The maximin-Ds design's efficiency for the half-saturation constant at every true value across a 16-fold range. It is flat to within 3.6 points and its minimum, 61.19%, is attained at 4 values including both ends — a minimum attained at one point could be improved by moving towards it, so the tie is what says the search has finished rather than stalled. The bars along the bottom are the least favourable prior found by a completely different route: a search over weightings, with Nelder–Mead inside it rather than the multiplicative algorithm the pair's version uses, because the Ds sensitivity's fixed point is not the one that algorithm converges to. It reaches 60.15% and puts its weight at K = 0.25, 1.00, 4.00, which is where the ties are.
Fig. 3 The maximin design’s efficiency curve with the values that attain its minimum marked, and the least favourable prior’s weights as bars along the bottom.

The equalisation is what says the search has finished rather than stalled: a minimum attained at a single value could be improved by moving the design towards it, so an optimum has to be a tie. The robust field records the same signature for the pair criterion and the criterion field records it one field further back as E-optimality’s repeated eigenvalue. Three different criteria, one structural fact — an optimum over a range is where the worst cases meet.

The second route is the least favourable prior. A maximin problem is an averaging problem under the weighting that makes the average as bad as possible, so a search over priors, with a Bayes design inside it, has to find the same answer. It reaches 60.15% against the direct search’s 61.19%, and it puts its weight at K = 0.25, 1.00 and 4.00 — the values where the ties are — with efficiencies there of 61.45%, 60.15% and 61.38%, a spread of 1.3 points.

The two routes share no code below the efficiency function itself, and the optimiser under the second is different by necessity rather than by choice: the multiplicative algorithm that finds the pair criterion’s Bayes design does not converge for this one. Run at its usual step it stalls with a sensitivity of 1.43 against a theorem that says 1, so the prior search uses a derivative-free method underneath. That is recorded here because it is the kind of detail that reads as an implementation note and is actually a fact about the criterion: the Ds sensitivity’s fixed point is not the one the D algorithm climbs to.

What the design actually looks like

The three settings are worth reading against the two the local designs use, because the pattern is not “the same design, spread out”.

The local design for K at the middle of the range puts 70.7% of its runs at t = 0.604 and 29.3% at the ceiling. The maximin design puts 42.1% at t = 0.259, 34.1% at t = 1.657 and 23.8% at the ceiling. The lowest setting has moved down — below where the local design at the smallest K in the range would put it — and a middle setting has appeared where no local design puts anything.

That middle setting is the one that does the work at the top end of the range. At K = 4 the response does not bend until much later, so a design with runs only at 0.26 and at the ceiling learns almost nothing about where the bend is; the run at 1.66 is what keeps the efficiency at 61% there rather than at the 5.2% the local design manages at a factor of sixteen.

And the weights are no longer a closed form. The 1/√2 the previous essay derives is a property of a two-point design for a single value of K; over a range there are three settings and three weights and none of them is a constant.

What the guarantee costs as the range widens

A design cannot protect a range it is not told about, and every quantity here is a function of how wide the range is.

What a guarantee about one parameter costs, as the range widens. Four designs, each scored by the worst Ds-efficiency it can be promised over a range of true values spanning the stated factor either side of 1. The upper curve is the design that maximises exactly this quantity: 80.8% at a factor of 2, 61.2% at a factor of 4, 51.6% at a factor of 8, 41.8% at a factor of 16. Below it is the maximin design for the pair, which is a robust design and is robust about something else — it holds 26.3% where the subset design holds 41.8% at a factor of 16. The two designs built at the guess fall away fastest, and the one built at the guess for the pair is the worst of the four everywhere past a factor of 2.
Fig. 4 Four designs, each scored by its worst Ds-efficiency over a range spanning the stated factor either side of the guess.

At a factor of two the maximin design for the subset holds 80.8% and the local design for the subset holds 79.9% — a gap of one point, which is the honest statement that robustness is not worth having over a narrow range. Both use two settings there.

At a factor of four the gap is 61.2% against 42.1%; at eight, 51.6% against 16.3%; at sixteen, 41.8% against 5.2%. The local design falls off a cliff and the maximin design declines slowly, and where the two curves separate is where the third setting appears.

The design for the pair over the same widening range holds 61.4%, 41.9%, 33.4% and 26.3% on the subset criterion — always well below the design built for it, and always well above the local designs past a factor of four. It is a genuine hedge against the wrong thing.

Four designs, scored on the one parameter that was wantedEvery design scored by its Ds-efficiency for K at 13 true values across a 32-fold range. The peaked curve is the subset design built at the guess K = 1: 100% there and 28.2% at the worst point of the range. The flat curve is the maximin-Ds design, never above 64.3% and never below 57.7%. Between them is the maximin design for the pair — a robust design, protecting something else, and worth 38.5% at worst here. The lowest curve is the D-optimal design at the guess, which is what an experimenter who wanted K and looked up a design for the model would actually run: 29.1% at the worst point, against 57.7% available.00.2500.5000.7501-0.602-0.30100.4520.903the true value of K, on a log scaleDs-efficiency for the half-saturation constantthe guess, K = 1worst 57.7%four designs, 13 true values, ×32worst 29.1%, 28.2%, 38.5%, 57.7%
Fig. 5 Drag the top of the range. The maximin curve stays flat and drops; the local curves keep their peak at the guess and lose everything at the far end, which is what a guarantee about a range does and does not buy.

The third setting is what buys the protection

At a factor of two the maximin design uses two settings; past that it uses three. That is an integer that the arithmetic decides rather than the experimenter, and it is the same staircase the criterion field measures for a Bayesian design as the prior widens.

The reason is structural. A two-point design has two settings and one weight to spend, which is enough to be optimal at one value of K, and a range wide enough that the best settings at its two ends do not overlap cannot be covered by two points however they are placed. The third setting is where the runs go that would be right if the guess were wrong.

The worst case is a tie, and the prior says where. The maximin-Ds design's efficiency for the half-saturation constant at every true value across a 4-fold range. It is flat to within 19.2 points and its minimum, 80.77%, is attained at 2 values including both ends — a minimum attained at one point could be improved by moving towards it, so the tie is what says the search has finished rather than stalled. The bars along the bottom are the least favourable prior found by a completely different route: a search over weightings, with Nelder–Mead inside it rather than the multiplicative algorithm the pair's version uses, because the Ds sensitivity's fixed point is not the one that algorithm converges to. It reaches 80.75% and puts its weight at K = 0.50, 2.00, which is where the ties are.
Fig. 6 The tie over a narrower range, where two settings are enough and the design is nearly the local one. The equalisation property still holds; there is just less for it to equalise.

Two repairs that complement rather than overlap

The four worst cases on the subset criterion — 29.1%, 42.1%, 41.9% and 61.2% — are a complete two-by-two, one factor being whether the criterion was fixed and the other whether the range was, and reading them as one says something the list does not.

Fixing the criterion is worth 13.0 points on a local design and 19.3 on a maximin one. Fixing the range is worth 12.8 on the D criterion and 19.1 on the Ds one. Both readings give the same interaction, +6.3 points, and its sign is the finding: the two repairs together are worth 32.1 points where separately they are worth 25.8.

They are complements. That is the opposite of the pattern this site’s other paired repairs show — adjusting a covariate and conditioning on the allocation rule overlap almost completely, and the two reference-distribution repairs for a nested table cancel — and it has a structural reason. Those pairs address one deficiency from two sides. These address two independent deficiencies, so fixing one leaves the other’s cost intact and the second fix then acts on a design that has more to protect.

The practical form is unambiguous. A half-repaired design is worth about a fifth of the range in efficiency and a fully repaired one about a third, and the extra sixth is only available to somebody who does both.

The coincidence is a crossing, and it is at this width

Designing for the right question locally and for the wrong question over the range are worth 42.1% and 41.9%, which reads as a coincidence at one width. The width sweep says it is a crossing.

At a factor of two either side the local design for the subset holds 79.9% and the maximin design for the pair holds 61.4% — the criterion repair ahead by 18.5 points. At a factor of eight they are 16.3% and 33.4% — the range repair ahead by 17.1. At sixteen, 5.2% against 26.3%, ahead by 21.1.

So the two curves cross, they cross steeply, and they cross at very nearly the fourfold range the essay measures at. The equality is not a fact about either repair; it is the point where a repair whose value falls with width overtakes one whose value rises with it.

That gives a practitioner a rule with a number in it. Below about a fourfold uncertainty either side of the guess, fix the criterion; above it, fix the range — and the penalty for choosing wrongly is about twenty points at either end of the sweep, which is larger than the interaction and larger than either repair’s own contribution at the crossing.

It also says why the crossing is easy to miss. Measured at one width, and at that width, the two repairs look interchangeable and the choice between them looks like a matter of taste. Measured at two widths a factor of four apart, they look like answers to different questions — which is what the four-letter comparison says about criteria and what this field says about the pair of them at once.

The guarantee is a promise about a procedure, not about an experiment

One reading of these numbers is available and is wrong, and it is worth closing off because it is the natural one.

A worst-case efficiency of 61.2% does not mean the experiment will be 61.2% efficient. It means that whatever K turns out to be, within the stated range, the design will be at least 61.2% efficient for estimating it — and the whole of the flat curve says it will be between 61.2% and 64.8%. The number is a floor over a range of unknowns, which is a different kind of statement from an average and a different kind again from a realised efficiency.

That is why the comparison with the local design is not “42.1% against 61.2% on the same experiment”. The local design is 100% efficient if the guess is right and 42.1% at the worst point; the maximin design is 61.2% at its worst and never above 64.8%. An experimenter confident in the guess should run the local design and one who is not should not, and the whole content of the field is that the second case has an answer rather than only the first.

The same reading applies to the interval a design produces: a guarantee about efficiency is a guarantee about variance, and variance is what the interval’s width is made of, not what its coverage is made of.

Where this sits against the two fields it composes

The robust field’s version of this essay measures a worst case of 78.74% over the same sixteenfold range. This one measures 61.19%. The two numbers are not comparable and the temptation to compare them is exactly what this field exists to head off: the first is a guarantee about the pair of parameters and the second about one of them, and the second is smaller because it is a harder promise.

A design that is right once, and one that is never wrong by much. Three designs for the Michaelis–Menten model, scored at every true value of K across a 16-fold range. The peaked curve is the two-point local design built at K = 1: 100% there and 66.7% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 67.9% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 3 settings, never above 80.8% and never below 78.8%. Its worst case is 12.0 points better, and what it gives up is the 21.2 points at the one value the local design was built for.
Fig. 7 The robust field’s own picture: three designs scored on the pair criterion over the same range. The flat line there is a design this essay scores at 41.9% on the question the experimenter asked.

A design’s guarantee is a promise about a criterion over a range, and both halves have to be quoted. Neither field’s number is wrong; a paper that reported either one as “the worst-case efficiency of the robust design” would be.

What each design is worth on the other's criterion. Two designs for the Michaelis–Menten model — one chosen for both parameters, one for the half-saturation constant alone — each scored by both criteria. A design chosen for the pair is 84.93% efficient for the one parameter; a design chosen for the one parameter is 88.34% efficient for the pair. Neither number depends on K, V or T: measured at K = 0.25, 0.5, 1, 2, 4 they agree to nine decimals, because every quantity in either criterion is a function of g(t) = t/(K + t) alone and both designs sit at fixed points of that scale. The asymmetry runs the way nobody expects — narrowing the target costs the wider question less than the wider question costs the narrow one.
Fig. 8 The two cross-efficiencies at a single value of K, from the previous essay. Over a range they become the four worst cases this essay is about, and the asymmetry gets larger rather than smaller.
Robust about what?. Four designs over a range of K from 0.25 to 8, each shown by its worst case on both criteria: the bar is the worst efficiency for the half-saturation constant alone, the mark beside it the worst efficiency for the pair. The maximin design for the subset holds 57.7% on the subset and 55.6% on the pair; the maximin design for the pair holds 76.6% on the pair and 38.5% on the subset. Each is about twenty points better on its own criterion and about twenty points worse on the other, so a design described as robust over this range has said nothing until it says which of the two it means.
Fig. 9 The same four-way comparison over a range twice as wide. Every guarantee falls and the ordering does not change, which is what says the four designs are answering four different questions rather than approximating one.

What is claimed here, and what is not

This essay claims standardised maximin design for a subset of the parameters of a non-linear model: the design, its equalisation property, the least favourable prior that reaches it by a second route, and the four-way comparison that says the composition is a third problem.

What stays out and is named as a decision: maximin over a range of both parameters, which is available here only because V enters the settings not at all and would be a genuinely two-dimensional search in a model where it did; subsets of more than one parameter, where the criterion carries an s-th root and the efficiencies stop being ratios of variances; and the question of how a range should be chosen, which is the same judgement the robust field declines to make and for the same reason — it is an assumption about the answer, and the point of the whole construction is to need less of one.

The checks, and what they are checked against

Four claims are gated in this field’s library. The maximin design for the subset is required to beat the maximin design for the pair on the subset criterion by more than ten points, and to lose to it on the pair criterion by about as much, which is what makes the composition a third problem rather than a refinement. The worst case is required to be attained at more than one true value with both ends of the range among them. The least favourable prior is required to reach the same worst-case efficiency to within two points and to put its weight where the ties are, at efficiencies that are equal to within three. And the design is required to use more than two settings, because a two-point design cannot cover this range however its two points are placed.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

  • The design that hedges — both name bayesian optimal design, d-optimality, design measure, experimental design, locally optimal design, michaelis–menten, the non-linear model, optimal design
  • The family behind the letters — both name d-optimality, design measure, ds-optimality, equivalence theorem, experimental design, information matrix, optimal design
  • The guess with two numbers in it — both name d-optimality, design measure, ds-optimality, equivalence theorem, information matrix, the non-linear model, optimal design
  • Augmenting a design that has already run — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design
  • The criterion with no derivative — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design
  • The design that stops guessing — both name d-optimality, experimental design, locally optimal design, michaelis–menten, the non-linear model, optimal design

Named objects

A flat tag is an object no other essay names yet.

Bayesian optimal designD-optimalityDesign measureDs-optimalityEquivalence theoremExperimental designInformation matrixLeast favourable priorLocally optimal designMaximin designMichaelis–MentenThe non-linear modelOptimal designSubset of parameters