The criterion, and what it assumes

The two terms anybody wanted

D-optimality estimates all six parameters of a quadratic as precisely as possible. Nobody wants that. An experimenter looking for a maximum wants the two curvature terms, and the design that gives them is not the D-optimal one — it is a quarter of the runs at the centre, exactly, and the D-optimal design is 75.3% efficient for the question that was actually asked.

Worth reading first: A design is a number · Walking up the gradient.

The quadratic model in two factors has six parameters: an intercept, two main effects, an interaction and two squared terms. D-optimality maximises the determinant of the information matrix, which is a statement about all six of them at once — the volume of the confidence ellipsoid for the whole vector.

Almost nobody wants all six. An experimenter climbing towards a maximum wants the curvature, because that is what says where the maximum is and whether there is one. An experimenter screening factors wants the main effects and would be content to know nothing about the intercept. The intercept in particular is never the point: it is the response at the centre of a region that was chosen for convenience, and no decision in any experiment has ever turned on it.

Ds-optimality asks the question that was actually asked. Nominate a subset of the parameters, treat the rest as nuisance, and optimise for the subset.

Where a Ds-optimal design puts its runs. The Ds-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.2500, 0.1250, 0.0625, which nine equal runs cannot express.
Fig. 1 The Ds-optimal measure for the two curvature terms, over the same 121 candidate settings the whole field uses. Nine settings again, and a set of weights that is nothing like D’s: a quarter of the experiment at the centre and a sixteenth at each corner.

The criterion, and the theorem it brings with it

The definition is a ratio of determinants:

Ds = |M| / |M₂₂|

with M₂₂ the block of the information matrix belonging to the parameters that are not of interest. Equivalently it is 1/|(M⁻¹)₁₁| — the information the design carries about the s parameters of interest after the nuisance parameters have been allowed for, which is what a subset confidence region’s volume is built from. Larger is better, and the s-th root puts it on a per-parameter scale in the same way D’s p-th root does.

The reason this is not merely another entry in a list is that it brings its own equivalence theorem, and the theorem is an equality:

δ(x) = f(x)′M⁻¹f(x) − f₂(x)′M₂₂⁻¹f₂(x), and at the optimum max δ = s

The subtraction is the whole content. The first term is the prediction variance the D case uses; the second is the part of it the nuisance parameters are responsible for. What drives the search is the information a candidate carries about the subset that the rest of the model has not already supplied — so a setting that is highly informative overall but only through the intercept contributes nothing, and the algorithm discards it.

Because it is an equality it is gated at machine precision. The Ds-optimal measure’s largest directional derivative comes out at s to within 9 × 10⁻¹⁵. The optimality field learned the value of gating an exact claim exactly the expensive way: a figure there once claimed machine precision and delivered 1.4 × 10⁻⁴, with a tolerance loose enough to accept both, so two numbers stood for one claim in one essay.

What comes back, and it is exact

The nine settings are the 3² factorial again — the same nine every criterion in this field returns over 121 candidates. The weights are these:

  • the centre, at exactly 0.25,
  • each edge midpoint, at exactly 0.125,
  • each corner, at exactly 0.0625.

Not approximately. The search returns 0.249999999999982 and 0.124999999999995 and 0.062499999999999, on an eleven-point grid and again on a twenty-one-point grid, and the essay is free to write ¼, ⅛ and 1/16.

Those nine numbers are the product of the one-dimensional measure (¼, ½, ¼) with itself. A corner is ¼ × ¼, an edge midpoint is ¼ × ½, and the centre is ½ × ½. In other words, the design that answers how curved is the response in each factor is what an experimenter would get by treating the two factors separately and running the product — the arrangement anybody would write down without any of this machinery.

Where a D-optimal design puts its runs. The D-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.1458, 0.0962, 0.0802, which nine equal runs cannot express.
Fig. 2 The D-optimal measure for comparison, on the same nine settings. 0.1458 at a corner, 0.0802 at an edge, 0.0962 at the centre — and those do not factorise. A product measure with 0.1458 at the corners and 0.0962 at the centre would put 0.1184 at an edge midpoint, and the search puts 0.0802.

That failure to factorise is a real difference in kind rather than in degree, and it is worth naming what causes it. The interaction term x₁x₂ is the one part of the model that couples the two factors, and it is estimated best at the corners. D cares about it and buys corners; Ds for the pure quadratic terms does not care about it at all, and the design separates.

So the two criteria differ by more than their weights: one of them separates the factors and the other does not, and the interaction is where the difference lives. Both facts are gated, the first at 10⁻¹¹ and the second as a comparison a product measure fails by four percentage points of weight.

The product form gives every number of factors at once

The factorisation is reported as a curiosity about two factors and it is a construction for any number of them.

If the two-factor Ds-optimal measure for the pure quadratic terms is the one-dimensional (¼, ½, ¼) multiplied by itself, then the k-factor one is the same measure multiplied by itself k times. Every weight follows without any search: a setting at the extreme in j of the k factors and at the centre in the rest carries (¼)^j (½)^(k−j).

At three factors that gives the eight corners 1/64 each, the twelve edge centres 1/32, the six face centres 1/16 and the centre 1/8 — totalling 0.125, 0.375, 0.375 and 0.125, which sums to one exactly.

The reading worth carrying is what happens to the centre. Its weight is 2^(−k): a quarter of the experiment at two factors, an eighth at three, a sixteenth at four. The centre matters less the more factors there are, which is the opposite of the response-surface habit of replicating the centre point, and it happens because a k-factor product spreads the same one-dimensional structure across more coordinates rather than concentrating it.

None of that needed a search or a grid, and it is the practical yield of the exactness: a measure that comes back as ¼, ⅛ and 1/16 rather than as three decimals is a measure whose generalisation can be written down.

Where Ds sits between D and E

Totalling each measure by kind puts the three criteria on one line.

D spends 58.3% at the corners, 32.1% at the edge midpoints and 9.6% at the centre. Ds spends 25%, 50% and 25%. E — from the family’s far end — spends 21.0%, 40.1% and 38.9%.

So Ds sits between the two in the corners-to-centre deformation, and much nearer E than D. That is not a coincidence of arithmetic: the family’s own analysis identifies the worst-determined direction as the combination only a centre run separates — the intercept against the sum of the squared terms — and that is precisely the direction Ds is buying by declaring the intercept a nuisance.

E and Ds arrive at nearly the same design from opposite premises. One asks for the worst-determined direction whatever it turns out to be; the other names the parameters it wants and lets the criterion subtract what the nuisance already knows. On this model they name the same thing, which is why the D design loses 49.8 points to E and 24.7 to Ds — a run multiplier of 2.0 and 1.33 for the same underlying failure at two intensities.

What the subtraction is doing

The derivative is worth one more paragraph, because reading it is what makes the weights obvious in retrospect.

Take the centre of the region. Its model vector is (1, 0, 0, 0, 0, 0): it says nothing about either main effect, nothing about the interaction, and nothing about either squared term directly. Under D-optimality its directional derivative is f′M⁻¹f, which is small — the centre is a poor place to learn about five of the six parameters, and D’s weight there is correspondingly modest at 0.0962.

Under Ds the same setting is scored by f′M⁻¹f minus f₂′M₂₂⁻¹f₂, where the second term is what the nuisance parameters already know about it. The centre is enormously informative about the intercept, and the intercept is nuisance — so almost all of the first term is subtracted away, and what is left is the part that is about the curvature. That residue is large, because the centre is the one setting that separates the intercept from the sum of the squared terms.

So the criterion is not rewarding the centre for being informative. It is rewarding the centre for being informative about something the rest of the design cannot supply, which is the only sense in which a run is worth anything to a subset.

The same reading explains the corners going down. A corner is where the interaction is largest, the interaction is nuisance here, and most of what a corner contributes is subtracted. It keeps a sixteenth of the experiment rather than none because the corners also carry the squared terms — at a corner both x₁² and x₂² are one — and that part survives the subtraction.

The price of the wrong letter

Two efficiencies, each against its own optimum, which is the only honest way to compare them.

The D-optimal design is 75.3% efficient on Ds. An experimenter who asked for “the optimal design” and was given the letter that comes first has given away a quarter of the information about the two parameters they came for.

The Ds-optimal design is 83.6% efficient on D. The reverse loss is smaller, and the asymmetry has the same shape as the one in the family: a design chosen for the narrow question is nearly right for the broad one, and a design chosen for the broad question can be well wide of the narrow one.

Six criteria on 5 designs, on a square region. Each design scored as an efficiency — its value over the best attainable — so six criteria in as many different units sit on one scale where 1 is the optimum. Rows are ordered by DS. D picks 13-run exchange; A picks face-centred composite; G picks 13-run exchange; I picks face-centred composite; E picks face-centred composite; DS picks face-centred composite. Every design has been scaled to just fit the region first, because a design run at settings the region does not contain is not a competitor on it. The disagreement is the point: the letter is a choice, and it is almost never reported as one.
Fig. 3 Six criteria on the catalogue designs and the two found by search, ordered by Ds. The thirteen-run exchange design that wins D at 0.998 sits at 0.704 here; the face-centred composite that loses D by nearly two-tenths wins the column at 0.919.

The row worth reading twice is the thirteen-run exchange design. It is the best design in the table by D — 0.998, about as close to the theoretical optimum as an integer number of runs can reach — and it is 0.704 on the subset an experimenter looking for a maximum actually wants. Thirty per cent of the information about the curvature, given up by a search that did precisely what it was asked.

The catalogue’s own designs, scored on the question

The efficiencies above are against optimal measures, which is the right comparison and is not the one an experimenter faces. Nobody runs a measure. What is available is a catalogue, and the catalogue’s designs score as follows on the subset: the 3² factorial at 0.889, the face-centred composite at 0.919, the rotatable composite at 0.540.

Two of those are worth a sentence each.

The face-centred composite winning is not an accident and is not a general recommendation. It has axial runs at the faces rather than outside them, so it carries three levels of each factor with more of its runs near the middle of the region than a corner-heavy design does — which is what the subset wants, for the contrast reason above.

Where a A-optimal design puts its runs. The A-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.2332, 0.0978, 0.0940, which nine equal runs cannot express.
Fig. 4 The A-optimal measure, drawn for scale: it sits between D and Ds in how much it loads the centre, and its Ds efficiency sits between theirs too. The criteria are not a taxonomy of unrelated answers but a line, and every design in the catalogue lands somewhere along it.

The rotatable composite at 0.540 is the row this site has had to be careful with before. It is tabulated at axial distance √2, which is outside a square region, and scoring it as tabulated gives a D-efficiency of 1.1990 against a bound of one — arithmetic saying it was being run at settings the region does not contain. Every row here is coded to just fit the region first, which is why it scores 0.540 rather than something impossible, and the systems field records the same trap from the other side.

Why the centre gets a quarter

The weights are worth an explanation rather than only a table, because the reason is the same structural fact the response-surface field turns on.

A quadratic’s curvature is estimated by a contrast — roughly, the average of the outer runs minus the average of the centre runs. A contrast of two averages is estimated best when the two groups are about the same size, which is why the centre, a single setting, has to carry as much weight as the four corners put together.

That is not a new observation on this site. The response-surface field found the same thing when it compared two thirteen-run designs: replicating the factorial and adding five centre runs gives 86.6% power for the curvature test against 82.9% for four corners and nine centre runs, and wins on the contrast’s standard error too, because ȳ_f − ȳ_c wants its groups balanced. The essay there had to be rewritten around it. Here the same fact arrives as an exact weight rather than as a power comparison, and the two are consistent — which is the closest thing to a second route this particular claim has.

A quarter of the runs at one setting is a real instruction

It is worth pausing on what the answer tells an experimenter to do, because it is unusual advice and the arithmetic behind it is not the reason anybody would have arrived at it.

A twelve-run experiment built to these weights puts three runs at the centre, one at each edge midpoint, and — with eight corners’ worth of weight adding to half a run each — nothing reliable at the corners at all. A sixteen-run experiment puts four at the centre, two at each edge midpoint and one at each corner. Both are designs an experimenter would look at twice, because replicating one setting three or four times looks like waste to anyone trained to spread runs out.

What the Ds-optimal measure predicts along the diagonal. d(x) = f(x)′M⁻¹f(x) along the diagonal of a square region, for the Ds-optimal measure. The line at 6 is the number of parameters in the model. Kiefer and Wolfowitz's theorem says a design is D-optimal exactly when the largest d anywhere in the region is p — not approximately, equals — so the optimal curve is tangent to that line at its support points and below it everywhere else. Here it reaches 11.000, which is 83.3% above the bound and is exactly the amount by which this design is not D-optimal.
Fig. 5 The prediction variance along a cut through the region under the Ds-optimal measure. It is drawn here for the same reason the optimality field drew it under D: a measure is an abstraction, and what an experimenter feels is what the design predicts, which is a different shape under a criterion that has put a quarter of the runs in one place.

It is not waste, and the reason is the contrast. Curvature is estimated by comparing the outside of the region with the middle of it, and the middle is one point. Every run spent there sharpens one half of a difference of two averages, and a difference of two averages is limited by whichever half is noisier.

There is a second consequence that the criterion cannot see and that an experimenter should. Three runs at the identical setting also supply pure error — an estimate of σ² that owes nothing to the model being right — which is what a lack-of-fit test is built from. The optimality field’s first essay named this trade explicitly: a design chosen to estimate p parameters as precisely as possible is, by construction, a design that spends nothing on detecting that those p parameters are the wrong ones, and model adequacy is not a function of the information matrix. Here the criterion happens to buy replication for its own reasons, and the lack-of-fit degrees of freedom arrive free. That is luck rather than design, and it is worth taking.

Where a Ds-optimal design puts its runs. The Ds-optimal measure over 81 candidate settings on a circular region. It keeps 5 of them and discards the rest, and the 5 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.3333, 0.1667, which nine equal runs cannot express.
Fig. 6 The same subset criterion on a circular region, where the corners the interaction lives at are not available. The weights are different and the structure is the same: the centre carries the largest share, because the curvature contrast is measured from it whatever shape the region is.

Where the subset is a different subset

Nothing above is specific to the curvature terms. Any subset can be nominated, and the design changes with it.

Where a E-optimal design puts its runs. The E-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.3887, 0.1003, 0.0525, which nine equal runs cannot express.
Fig. 7 The E-optimal measure, drawn here because it is the closest neighbour to Ds among the criteria in this field: it also loads the centre heavily, at 0.3887, for a related reason — the worst-determined direction in a quadratic model is the one only a centre run separates.

Two subsets are worth naming and neither is measured here. Nominating the main effects gives the screening problem, where the answer is close to a two-level factorial and the quadratic terms become nuisance. Nominating a single parameter gives what is usually called c-optimality, whose optimal designs can be supported on as few as two settings and which is the criterion behind the interval for the location of an optimum — a quantity that is a ratio of parameters rather than one of them, which puts it outside the family entirely.

What is being claimed here, and what is not

This essay claims Ds-optimality on one subset of one model: its equivalence theorem gated as an equality, the exact support and weights, the factorisation of the subset optimum against the non-factorisation of the whole-model one, and the cross-efficiencies both ways.

What stays out: subsets chosen adaptively, which is a selection problem and would need the machinery the adaptive field uses rather than this one; the general theory of when a subset optimum is supported on fewer points than the full one; and any suggestion that an experimenter can nominate the subset after seeing the data. That last is not a technical omission. A criterion is a statement about what the experiment is for, and it has the same standing as a stopping rule: it is part of the design, it is decided before the data exists, and choosing it afterwards is the same error under a different name.

Six criteria on 5 designs, on a square region. Each design scored as an efficiency — its value over the best attainable — so six criteria in as many different units sit on one scale where 1 is the optimum. Rows are ordered by E. D picks 13-run exchange; A picks face-centred composite; G picks 13-run exchange; I picks face-centred composite; E picks face-centred composite; DS picks face-centred composite. Every design has been scaled to just fit the region first, because a design run at settings the region does not contain is not a competitor on it. The disagreement is the point: the letter is a choice, and it is almost never reported as one.
Fig. 8 The six-criterion table ordered by E rather than by Ds, which is nearly the reverse ordering. Two columns of one table, computed from one information matrix, disagreeing about every design in it.

The checks

Three claims are gated in this field’s library.

The equivalence theorem holds as an equality, with max δ − s below 10⁻⁹ — the same standard the D case is held to and for the same reason.

The subset optimum factorises, with the centre, edge and corner weights required to be ¼, ⅛ and 1/16 to within 10⁻¹¹, and the support required to be exactly nine settings. A search that drifted onto a different set of nine, or onto weights that were merely close, fails.

And the whole-model optimum does not, asserted by reconstructing the product measure the D-optimal weights would imply and requiring it to miss the edge weight by more than a hundredth. That half exists because “Ds factorises” is only interesting if something else does not, and a check that only ever confirmed the first half would prove nothing about the difference between the criteria.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CurvatureD-optimalityDesign measureDs-optimalityE-optimalityEquivalence theoremInformation matrixNuisance parameterOptimal designThe Φₚ familyResponse-surface