What the procedure may not read

The worst case in two directions

A design that protects a range of one parameter is robust. Protect the range of one parameter while holding the other at a guess and the design is still robust, still has a guarantee, and guarantees no more than a design that protects nothing at all.

Worth reading first: A design is a number · The design that needs the answer.

A design built at a guess is worth what it is worth at the truth, and the truth is somewhere else. The repair this site has measured twice is the maximin design: the design whose worst efficiency over a stated range of parameter values is as large as it can be made. It has been measured across a range of one parameter and across a range of one parameter for a criterion about a subset, and in both cases the answer had a property that made it recognisable as an answer: the worst case is attained more than once.

Both of those problems are one-dimensional, because the model they use has only one parameter that moves the settings. Here both do, so the range is a rectangle and the worst case is a minimum over a surface. Nothing about the one-dimensional argument for equalisation survives the change, and there is no reason a minimum over a two-dimensional set should be attained at more than one point.

It is attained at three.

The design, and what it guarantees

The rectangle is a fourfold range in each coordinate around the guess θ = (0.2, 1.2): θ₁ from 0.1 to 0.4, θ₂ from 0.6 to 2.4. The design that maximises the worst Ds-efficiency for θ₁ over it uses three settings — 1.006, 3.431 and 8.065, carrying 0.091, 0.419 and 0.490 of the runs — and its guarantee is 45.9%.

Four designs, and what each of them guarantees. The worst Ds-efficiency each design achieves anywhere in a rectangle of parameter values 4 times wide in each coordinate. The design built at the guess guarantees 7.3%: it is perfect where it was built and nearly useless at one corner. The D-optimal design at the same guess guarantees 25.6% — it answers the wrong question everywhere and is therefore not concentrated on being right anywhere. The third is the one this field exists to measure: maximin over the first parameter's whole range, with the second held at its guess. It guarantees 7.0%, which is no better than the design that protects nothing. Protecting both is worth 45.9%, and it needs 3 settings to do it.
Fig. 1 Four designs and one number each: the worst Ds-efficiency each achieves anywhere in the rectangle. Two of the four are hard to tell apart and one of those two was built to be robust.

The worst case is attained at (0.1, 0.6), (0.4, 0.6) and (0.4, 2.4) — three of the rectangle’s four corners, at efficiencies equal to within the search’s tolerance. That is the equalisation property, and it is what says the search has finished rather than stalled: at a design where the minimum is attained once, moving towards that point improves the worst case, so the optimum is exactly where improving one bad point costs another.

The property is worth a moment because it is doing a job here that no closed form can do. This model has no analytic optimum — the design is whatever a derivative-free search returned — so the question “is this the maximin design or the place the simplex stopped?” needs an answer that does not come from the same search. In the one-dimensional problems the answer came from two places: the equalisation condition, and a second search over least favourable priors. Here it comes from equalisation alone, and the reason the other route is missing is given below.

What robust in both guarantees, point by point. The Ds-efficiency of robust in both at every corner, edge and centre of a rectangle of parameter values 4 times wide in each coordinate, centred at the guess θ = (0.2, 1.2). Each cell is what the design delivers for the first parameter as a fraction of what a design built at that very point would deliver. The worst cell is 45.9% and it is the number the whole field is about: a design's guarantee is its worst cell, not its average, and not the 100% it scores where it was built.
Fig. 2 The robust design at every parameter point of the rectangle. Three cells are equal and lowest, and they are three of the four corners; the cell at the guess is 68.4%, which is what the protection costs where the guess happens to be right.

What a guarantee of 45.9% actually says

Efficiency here is a ratio of information, so it converts into runs directly: a design that is 45.9% efficient needs 1/0.459 = 2.18 times as many observations as the design somebody who knew the parameters would have run, to reach the same precision on θ₁. That is the price of the guess, at the worst point of the rectangle, once the design has been chosen as well as it can be.

The comparison that matters is not with the omniscient experiment, which nobody can run. It is with the alternatives available to the same experimenter. The design built at the guess needs 13.6 times the runs at that corner; the D-optimal design at the guess needs 3.90 times; the design robust in one coordinate needs 14.2 times. And at the guess itself, where the local design is by definition perfect, the robust design needs 1.46 times — which is the whole of what protection costs when the guess turns out to have been right.

Every one of those figures is computed at nine parameter points and checked against twenty-five, and the two grids agree to four decimal places on all four designs.

Those five numbers are the field in one paragraph. Every one of them is an efficiency read as a multiplier on the experiment, and the choice between the designs is a choice about which of them the experimenter would rather be exposed to.

The weight that was a constant, drawn as the function it is. The share of the experiment the Ds-optimal design spends at its early setting, at each of the nine parameter points of the rectangle. For the model one field back the corresponding number is exactly 1/√2 = 0.7071 at every parameter value, every scale and every horizon — a closed form with nothing in it about the guess. Here it runs from 2.6e-9 to 0.2508. Where the two rates are far apart it is driven to zero altogether: the criterion is a ratio in which the numerator and the denominator vanish together, so its supremum is approached rather than attained and the optimal design asks for a vanishing share of the runs at the setting that identifies the nuisance.
Fig. 3 The weights of the designs built at each parameter point of the rectangle, which is what a single design has to compromise between.

How many settings a guarantee needs

The support size is part of the answer rather than a detail of it. Every design built at a single parameter point in this model has two settings, because the model has two parameters and there is nothing for a third to do.

How many settings a guarantee needs. The best worst-case Ds-efficiency attainable over the rectangle by a design with each number of settings. Two cannot do it however they are placed: 36.8%. Three reach 45.9%. Four reach 46.3% and five 44.2%, which is the search reporting its own resolution rather than finding anything — a four-setting answer that is 0.4 points above a three-setting one, with a weight of a tenth on a setting close to another, is not a fourth setting. How many settings a robust design needs is part of the answer, and here it is three.
Fig. 4 The best worst case attainable with each number of settings. The step from two to three is the protection; everything after it is the search reporting its own resolution.

Two settings cannot protect the rectangle however they are placed: the best a two-setting design achieves is 36.8%. Three reach 45.9%. Four reach 46.3% and five 44.2%, which is not a sequence with anything in it — a four-setting answer four tenths of a point above a three-setting one, whose extra setting carries a tenth of the weight and sits one and a half units from another, is a search finding the same design twice. So the answer is three, and the third setting is what buys the protection, exactly as the third setting does in the one-parameter version of this problem.

What the third setting is for is visible in the design: 1.006, 3.431 and 8.065 spread across the whole horizon, where the design built at the guess has everything at 0.656 and 5.743. A design that must work at several ratios cannot put its runs where any one ratio wants them.

The second route to a design that has no closed form. The Ds-sensitivity of the design for the first parameter, at the guess, across the whole design space. The equivalence theorem says a design is optimal exactly when this curve stays at or below one and touches one at every setting the design uses. It does: the largest value anywhere is 1.000000, at t = 5.742, which is one of the design's own settings. There is no closed form for this model — the design was found by a search — so this curve is what distinguishes an optimum from wherever the search happened to stop.
Fig. 5 The equivalence theorem at the guess, which is the only second route this field has and is what says the designs being compared are optima rather than search results.

Robust in one coordinate

Now the measurement this field exists for.

Take the previous field’s construction exactly as it stands — maximin over a fourfold range of the parameter of interest — and apply it here with the nuisance held at its guess. This is not a straw man: it is what “protect the range of the uncertain parameter” produces when the thing somebody is uncertain about is named as one parameter, and it is the only version of the construction available to anybody who has not noticed that the guess has two numbers in it.

Over the range it was given, it does exactly what it should. Its worst efficiency across θ₁’s fourfold range, with θ₂ at 1.2, is 71.5% — a proper guarantee, better than anything else here manages over anything.

Over the rectangle, its worst efficiency is 7.0%.

The design built at a single point for both parameters guarantees 7.3%, and the D-optimal design at the same guess — which answers the wrong question everywhere and is therefore concentrated nowhere — guarantees 25.6%. The one-coordinate robust design guarantees very slightly less than the design that protects nothing, and a third of what a design built for the wrong criterion does.

What robust in one coordinate guarantees, point by point. The Ds-efficiency of robust in one coordinate at every corner, edge and centre of a rectangle of parameter values 4 times wide in each coordinate, centred at the guess θ = (0.2, 1.2). Each cell is what the design delivers for the first parameter as a fraction of what a design built at that very point would deliver. The worst cell is 7.0% and it is the number the whole field is about: a design's guarantee is its worst cell, not its average, and not the 100% it scores where it was built.
Fig. 6 The one-coordinate design evaluated at every parameter point. It is above 90% along the row it was protecting and falls off a cliff in the other direction, which is the direction nobody asked it about.

The reason is not subtle and that is the point. The worst case simply lives in the other coordinate. Averaged over the rectangle the one-coordinate design is fine — 55.3%, marginally better than the design at the guess — and it is fine along the whole line it was built to protect. It is the corners in θ₂ that kill it, and no amount of protection in θ₁ touches them.

A guarantee that names one parameter is not a guarantee. The word “robust” is a claim about a set, the set has to be the set somebody is actually uncertain about, and a design that is maximin over a subset of that set has a number attached to it which is true about the subset and false about the experiment.

One detail of the three attained corners is worth having, because it says what the design is trading off. Two of them — (0.4, 0.6) and (0.4, 2.4) — are the fast-θ₁ corners, where the response has decayed long before the horizon ends and the late runs are wasted. The third, (0.1, 0.6), is the slow corner, where the horizon binds and the design would like settings it cannot have. The design is balanced between running out of signal and running out of experiment, which are different failures, and the corner it is not worst at is (0.1, 2.4) — the one where the two rates are most separated and the nuisance is easiest to pin down.

Where the runs go, when the answer is not known in advance. The two designs as they would be run. The local design at K = 1 puts half its runs at 0.833 and half at 10.00, which is KT/(2K + T) and the end of the range, both from a closed form. The maximin design over K from 0.25 to 4 uses 3 settings: 23.4% at t = 0.298, 32.2% at t = 2.029, 44.4% at t = 9.992. The extra setting is the whole of the protection — a two-point design cannot cover the range however its two points are placed, because the point that is right for a small K is wrong for a large one and there is nowhere in between that is right for both.
Fig. 7 Where the one-dimensional maximin design puts its settings as the range widens, which is the same spreading this one does in two directions at once.

What it costs as the guess gets vaguer

The comparison changes shape as the rectangle widens, and the changes are worth reading because two of them are not what the sentence above would predict.

What the guarantee costs as the guess gets vaguer. Four designs' worst-case efficiencies over rectangles of increasing width, all centred at the same guess. The upper curve is the design that protects both coordinates; it falls from 76.7% to 36.0% as the rectangle goes from twice to sixteen times wide. The lower pair are the design built at the guess and the design that protects the first parameter's range with the second held fixed, and they are close to each other at every width — the second is robust in a coordinate where the worst case does not live. The D-optimal design at the guess is drawn as well, and above the two of them: a design for the wrong question beats a design robust in the wrong coordinate.
Fig. 8 Four designs’ guarantees against the width of the rectangle. The top curve falls slowly; the two built at a point fall off the bottom; the one-coordinate design does something in between that is worth reading twice.

At a rectangle twice as wide, every design is respectable — 76.7% for the maximin design, 56.3% for the design at the guess, 55.6% for the one-coordinate design — and the distinctions barely matter. A guess good to a factor of √2 in each coordinate is a good guess.

At four times wide, the design at the guess and the one-coordinate design have both collapsed to about 7% and the maximin design holds 45.9%.

At eight and sixteen times wide the one-coordinate design separates from the design at the guess and holds 17.3% and 12.7% against 1.2% and 0.4%. So protecting one coordinate does buy something once the range is wide enough — it is simply worth about a third of what protecting both is worth, at a width where the design at the guess has stopped being an experiment at all.

And the maximin design’s own guarantee falls slowly and then stops falling: 76.7%, 45.9%, 37.5%, 36.0%. Doubling the rectangle from eight times to sixteen costs it a point and a half, because past a certain width the worst corner stops being about the ratio and starts being about the horizon — which is the one thing a design cannot fix, since it cannot measure past the end of the experiment.

What the guarantee costs as the guess gets vaguer. Four designs' worst-case efficiencies over rectangles of increasing width, all centred at the same guess. The upper curve is the design that protects both coordinates; it falls from 76.7% to 36.0% as the rectangle goes from twice to sixteen times wide. The lower pair are the design built at the guess and the design that protects the first parameter's range with the second held fixed, and they are close to each other at every width — the second is robust in a coordinate where the worst case does not live. The D-optimal design at the guess is drawn as well, and above the two of them: a design for the wrong question beats a design robust in the wrong coordinate.
Fig. 9 The same four curves. Drag the rectangle: the ordering between the middle two changes with the width, which is why “robust in one coordinate” cannot be given a single verdict.

The insurance costs half an experiment and saves eleven

The five run-multipliers are the field in one paragraph, and differencing them says what the protection is actually being bought and sold for.

At the worst corner the robust design needs 2.18 experiments’ worth of runs and the design built at the guess needs 13.6 — so the protection is worth 11.4 experiments where it is needed. At the guess itself the robust design needs 1.46 against the local design’s 1.00, so it costs 0.46 of an experiment where it is not.

Twenty-five to one. That is the exchange rate a designer is being offered, and it is an unusual one by the standards of the rest of this site: most of the trades measured here run between one and five to one. The reason is that the local design does not degrade towards mediocrity at the corners, it collapses, and an insurance premium of half an experiment against a fourteen-fold loss is not a close decision at any plausible belief about the guess.

The D-optimal design’s showing is the other reading in the same column and it is the more uncomfortable one. Built at a point and for the wrong criterion, it needs 3.90 experiments at the worst corner — 1.8 times the robust design’s requirement and 3.5 times better than the design built at the same point for the right criterion.

So on this rectangle, answering the wrong question robustly beats answering the right question at a point. The D-optimal design is not robust by construction; it is robust by accident, because a criterion that wants both parameters spreads its runs and a criterion that wants one concentrates them, and concentration is what a wrong guess punishes.

The guarantee has a floor, and the one-parameter version does not

The maximin design’s worst case across the width sweep runs 76.7%, 45.9%, 37.5%, 36.0%, and the falls between them are 30.8, 8.4 and 1.5 points. Each doubling costs about a fifth of what the last one did, and by an eightfold rectangle the decline has effectively stopped.

That is not how the one-parameter problem behaves. There the guarantee falls 6.5 points per doubling and is still falling — accelerating, in fact — at the widest range measured. Here it flattens at about 36%.

The difference is the horizon. Past a certain width the worst corner stops being the one where the ratio of the two rates is most awkward and becomes the one where the response has not finished by the end of the experiment, and that constraint does not get worse as the rectangle widens — it is already fully binding. A guarantee limited by the length of the experiment is a guarantee that a vaguer guess cannot make worse.

The practical form is unusually clean. Past an eightfold rectangle, widening the stated range is free, so an experimenter unsure whether to claim a fourfold or a sixteenfold uncertainty should claim the sixteenfold one: it costs ten points at the first doubling and a point and a half at the last, and the wider claim is the one that is true.

The one-coordinate design crosses the do-nothing design somewhere in the same region — behind it at a fourfold rectangle, 7.0% against 7.3%, and ahead by a factor of fourteen at eightfold and thirty at sixteenfold. So half a repair is worth nothing exactly where a full one is worth most, and starts paying only where the design it is being compared with has stopped being an experiment.

The route that is not here

The one-dimensional versions of this problem are checked twice: once by equalisation, and once by finding the least favourable prior — the weighting of the parameter range that makes the best average efficiency as small as possible — and confirming that the design which is Bayes-optimal under it is the maximin design. The two routes share nothing below the efficiency itself, and they agree to a fraction of a point.

That check is not made here, and the omission is deliberate rather than an oversight.

Over a rectangle the outer problem has eight free weights instead of one, and every evaluation of its objective is itself a search over designs. Two implementations were tried — a derivative-free search over the prior with a fresh inner solve at each step, and a multiplicative update with warm starts — and both converge to worst-case values well above the direct search’s: 52% and 30% against 45.9%. Neither of them is finding a better design; both are failing to solve their own outer problem.

Shipping either as a second route would produce a check that reports a disagreement between two routes whenever its outer optimiser has a bad day, which is worse than having one route, because it would be a check whose failures are its own arithmetic. What is checked instead is the equalisation condition, which is what the least favourable prior would have been supported on.

What is claimed here, and what is not

This essay takes maximin design over a rectangle for a criterion about a subset, and the claims are three: the equalisation property survives in two dimensions and is attained at three corners; the guarantee needs three settings where every local design has two; and protecting one coordinate of a two-coordinate guess is worth nothing at moderate widths and about a third of the full repair at large ones.

What stays out and is named as a decision: the least favourable prior, for the reason above; a continuous parameter region rather than a nine-point grid, which would change the numbers by whatever the grid is missing and is a real limitation — the worst case is a minimum over a set and a grid can only report the worst grid point; and any weighting of the rectangle other than the flat one implied by “maximin”, which is the Bayesian version of this problem and a different question.

The grid is worth one more sentence because it is the honest weak point. Every worst case here is computed at nine parameter points, checked against twenty-five, and the two agree to within a tenth of a percentage point at every design measured. That is evidence and not a proof: a minimum over a surface can hide between grid points, and the only defence offered is that the surface is smooth and that refining the grid does not move the answer.

The boundary against the design fields before this one is the model. Everything about what maximin design is, why the optimum sits on a tie and why a derivative-free search is needed rather than a multiplicative one is established there and used here unchanged — the search itself is theirs, handed a list of parameter pairs instead of a list of numbers, which is the only change the whole composition required.

The checks, and the refusals that make them mean something

Three claims are gated in this field’s library. The worst case is required to be attained at more than one parameter point, and the maximin design is required to need more settings than the design built at the guess — the equalisation property and the support count, stated as things that would fail if the two-dimensional problem behaved differently. The protection is required to be a multiple of what a design built at the centre achieves rather than a margin. And the one-coordinate design is required to guarantee no more than a design built at a point, which is this essay’s central measurement written so that it fails if it ever stops being true.

The refusal beside them is that same design offered as a robust one. Handed a design that is maximin over the parameter of interest with the nuisance held at its guess, the check compares what it guarantees over the range it was given — 71.5% — with what it guarantees over the rectangle, and throws. The failure it is designed to catch is not a bad design: it is a correct design with a guarantee attached to it that is about the wrong set.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Compartmental modelDesign measureDs-optimalityEfficiencyEqualisationEquivalence theoremLeast favourableLocal optimalityMaximin designThe non-linear modelNuisance parameterOptimal designRobustnessSupport points