Four letters and two camps
Worth reading first: A design is a number · The theorem that says when to stop.
The information matrix has six rows and six columns. A criterion has to turn that into one number, and there is no canonical way to do it — a matrix is not larger or smaller than another matrix, it is only larger or smaller in particular directions.
So there are several criteria, they are named by letters, and the letters are usually presented as a menu: pick the one that matches the goal in hand. That framing is correct and it leaves out the measurement, which is how much the choice actually changes. If the four criteria mostly agreed, the letter would be a notational preference. They do not.
What each letter is asking
The four are not variations on a theme. They are four different questions, and the differences are substantive.
D — the volume of the information. Maximise |M|^(1/p). This is a statement about the whole parameter vector at once: the confidence ellipsoid for all six coefficients has volume proportional to |M|^(−1/2), so D shrinks that ellipsoid. It is the criterion of somebody who wants to estimate the model.
A — the average variance of a coefficient. Minimise tr(M⁻¹), which is the sum of the coefficients’ variances. Where D cares about the volume of an ellipsoid, A cares about the sum of the lengths of its axes — a different thing, because a very elongated ellipsoid can have a small volume and a large trace. It is the criterion of somebody who wants every coefficient estimated well, and it is far more sensitive than D to one parameter being badly determined.
G — the worst prediction in the region. Minimise the largest d(x). This is a minimax criterion, and the only one of the four that never asks about a coefficient at all. It is the criterion of somebody who will use the fitted surface somewhere in the region and does not know where.
I — the average prediction over the region. Minimise the average of d(x) across the region. Same concern as G, different attitude to risk: G protects the worst point and I-optimality the typical one. Unlike the other three it needs to know what the region is rather than merely where its edges are, because an average is over a measure.
Two of those are about parameters and two are about predictions, and that is the division the numbers turn out to respect.
There is a fifth letter worth naming to say it is not here. E-optimality maximises the smallest eigenvalue of M, which protects the single worst-determined direction in parameter space — a criterion for somebody who fears one particular contrast being unidentifiable rather than one particular coefficient being imprecise. It has the same shape as the other three and it is not measured here, so it is not claimed here. The same goes for the whole family of Φₚ criteria that D, A and E are the limiting cases of, and for Ds-optimality, which optimises a subset of the parameters and is the criterion an experimenter who only cares about the quadratic terms actually wants.
The table
Every design in this field, scaled to just fit the square region, scored on all four. Reading each design as D, A, G, I:
- the thirteen-run exchange design, found by searching for D-optimality — 0.9977, 0.7925, 0.8718, 0.7452;
- the 3² factorial at nine equal runs — 0.9740, 0.9295, 0.8276, 0.9029;
- the face-centred composite at thirteen — 0.8195, 0.9300, 0.5841, 0.9492;
- the rotatable composite at thirteen — 0.4758, 0.4829, 0.2098, 0.5833.
Reading the first of those: the design found by searching for D-optimality gets 99.77% of the attainable D, and 74.52% of the attainable I. It is the best design here on two criteria and the worst of the three usable ones on the other two.
That is not a small effect. Losing a quarter of the attainable average-prediction quality is the sort of difference that would justify a larger experiment, and it was incurred by choosing a letter.
How far apart two criteria can be in principle
Before reading how much the four disagree here, it is worth knowing how much they are allowed to disagree, because for one of the pairs the answer is unbounded.
D is a function of the product of the information matrix’s eigenvalues and A is a function of the sum of their reciprocals. Fix the product and the sum of reciprocals is free to move: by the arithmetic–geometric mean inequality it is smallest when every eigenvalue is equal, and it has no upper bound at all.
A concrete pair makes the size of it visible. Six eigenvalues at (2, 2, 2, 2, 2, 2) and six at (64, 1, 1, 1, 1, 1) have the same product, so identical D. Their traces of are 3.00 and 5.02 — a factor of 1.67 in A, between two designs a D-criterion cannot tell apart.
So D places no ceiling on A whatsoever. A design can be D-optimal and have one coefficient estimated arbitrarily badly, provided the others are estimated well enough to keep the volume down. That is not a quirk of a particular region; it is what the two functions are.
And one pair that is the same criterion in disguise
The other pair is the opposite case, and it cuts across the camps this essay is about.
For continuous designs — where the design is a measure over the region rather than a fixed number of runs at fixed points — the general equivalence theorem says the D-optimal design is also the G-optimal one, and that at that design the maximum of over the region equals , the number of coefficients. Here that is 6.
So D and G are not two questions at all in the limit. They are one design characterised two ways, and a G-efficiency can be read directly as .
Which sharpens what any D-versus-G disagreement in this table can be about. It cannot be the criteria being fundamentally different concerns, because in the continuous limit they are the same concern. It has to come from the exact-design constraint: a fixed budget of runs, placed at points a practitioner can actually set, cannot in general reach the measure the theorem is about.
That is worth holding while reading the section on why the D-optimal design predicts badly. The surprising thing is not that a parameter criterion and a prediction criterion disagree; it is that these two, which coincide when the runs may be split arbitrarily, come apart once the runs have to be whole.
Two camps rather than four
The disagreement has a shape, and it is more informative than the disagreement itself.
D and G agree. Both pick the exchange design. That is not luck — it is the equivalence theorem, which says the D-optimal measure is exactly the G-optimal measure. The two criteria have one answer, so designs that approximate one approximate the other.
A and I-optimality agree. Both pick the face-centred composite. That is not a theorem, and it is not luck either: both criteria are dominated by how well the centre of the region is served, and the composite design spends five of its thirteen runs there.
So four letters produce two camps, and a reader choosing among them is really making one binary decision: is the model being estimated, or used?
The three optimal measures use the same nine settings. They differ only in how much of the experiment each setting gets, and the differences are large: the centre’s share runs from 9.6% to 25.7%, a factor of 2.7 across three criteria that all claim to be optimising the same experiment.
The reversal that is not about the letters at all
Here is the effect that turns out to be larger than any disagreement between criteria.
Change the region from a square to a disc — same model, same criteria, same designs, each scaled to just fit whichever region is in force — and the ranking inverts. The 3² factorial’s D-efficiency goes from 0.9740 on the square to 0.7253 on the disc. The rotatable composite’s goes from 0.4758 to 0.8929, in the opposite direction, over the same change.
The 3² factorial is the second-best design here on a square and the worse of the two on a disc. The rotatable composite is the worst design on a square and comfortably better on a disc.
And it reverses on all four criteria at once. On the disc the rotatable composite scores 0.8929, 0.9504, 0.7385 and 0.9634 against the factorial’s 0.7253, 0.6288, 0.4286 and 0.7419 — better on every one. On the square every one of those comparisons goes the other way. Whatever is happening here is not a disagreement between criteria; it is the region deciding, and the criteria all following.
The reason is geometric and it is worth stating plainly, because the response-surface field spent an essay on the rotatable composite’s virtues without ever naming the assumption underneath them. A design’s prediction variance being constant around any ring is a property that is only valuable if the region is a ring. On a circular region, where every direction from the centre reaches the boundary at the same distance, treating all directions alike is exactly right. On a square, the corners are √2 further out than the edge midpoints, and a design that refuses to distinguish directions is spending precision evenly on a region that is not even.
Why the D-optimal design predicts badly
The top row of the list above deserves an explanation rather than only a number, because “the D-optimal design is bad at prediction” sounds like a contradiction and is not.
D maximises the volume of information about the coefficients, and the way to get information about a quadratic’s coefficients is to observe at the extremes: the corners, where every term in the model is at its largest. A design chasing D pushes runs outward. The thirteen-run exchange design puts eight of its thirteen runs on the four corners and one at the centre.
Prediction at a typical setting is a different demand. Most of a square is not near a corner, and a design with almost nothing in the middle predicts the middle by extrapolating inward from its edges — which works, and works less precisely than having observed there. The I-optimal measure puts 25.68% of the experiment at the centre; the D-optimal one puts 9.62%.
Written as weights rather than percentages, the I-optimal measure is a single number and then eight copies of another: 0.2568 at the centre, and 0.0929 at each of the eight settings around it. Those eight are equal to the last digit, and the equality is the criterion’s own rather than something imposed on it — nothing in the optimisation was told that the square has a symmetry group, only that average prediction variance over the square is what to minimise, and the measure that minimises it inherits every symmetry the region has. The eight weights sum to 0.7432, which is what is left of the experiment once the centre has taken its quarter, and the arithmetic is exact: eight times 0.0929 is 0.7432.
So the two are not in tension by accident. They are asking for opposite things: estimate the surface’s shape wants leverage, and leverage means extremes; predict well across the region wants coverage, and coverage means the middle. The information matrix cannot serve both maximally, and each criterion is a statement about which to sacrifice.
The same trade is visible in a place this site has already measured it. A regression’s slope is estimated most precisely by observations far from the mean — which is exactly why one point can own a slope when it sits far enough out. Leverage buys precision about a coefficient and spends it on robustness, and here it buys precision about coefficients and spends it on prediction near the middle. It is one phenomenon appearing under two names in two fields.
The curve above is worth one more sentence, because it is the sharpest version of the trade available. The thirteen-run D-optimal design’s worst-predicted setting in the entire region is the centre, at d = 6.882, where its best is at (−0.75, −0.75) at d = 3.402. The three catalogue designs are the exact opposite: every one of them is worst at a corner. A design chosen to estimate coefficients is worst where a design chosen to predict would be best, and the reversal happens inside a single square with a single model on it.
Coding, which is the trap underneath the reversal
There is a step in the table above that has to be stated or the comparison is meaningless, and it is the kind of step that has bitten this site before.
The rotatable composite puts axial runs at ±√2. A square region stops at ±1. So the design as tabulated cannot be run on that region at all, and scoring it as it stands gives a D-efficiency of 1.1990 — above 1, on a scale whose maximum is 1.
That number is not a design beating the optimum. It is a design being allowed to turn the dials further than the region permits, and beating a competitor that was not. Every design in this field’s tables is therefore scaled so that it just fits the region before anything is computed, which puts the rotatable composite’s axial runs at ±1 and its corners at ±0.707.
The response-surface field records the same trap from the other side: a textbook result about how many centre runs give uniform precision came out one way in the standardised coding and another way in the coding this site writes designs in, and both numbers were computed here before anybody noticed they were two different languages. A comparison between designs at different codings is not a comparison, and an efficiency above 1 is the arithmetic saying so.
What the disagreement means for choosing
The practical residue is short, and the first item is the one this site keeps arriving at.
The letter is a choice and it is not reported. A design described as “optimal” has had a criterion picked for it, and the table above says that choice is worth up to 25 points of efficiency on the criteria that were not picked. This is the same structure as Bonferroni and Benjamini–Hochberg being described interchangeably as “correcting for multiple comparisons” when they bound different things — two procedures under one word, with a measurable gap between what each delivers.
The region is a choice and it is reported even less. It moved the ranking further than any criterion did, it never appears in a design’s name, and half the designs in common use were derived for a region whose shape is not the shape most experiments have.
Two of the four are nearly redundant. D and G have one optimum by theorem, and A and I-optimality are close enough here — 0.9300 against 0.9492 on their common favourite — that the distinction between them is unlikely to change a decision. The real menu has two items.
And the honest answer is often to score a design on all four. They cost nothing once M is formed, an experimenter who cannot decide can look at whether the disagreement matters for the design they were going to run anyway, and a design that scores well on all four — the 3² factorial does, at 0.9740, 0.9295, 0.8276 and 0.9029 — is a defensible choice under any of the questions.
Where the disagreement does not reach
One boundary on all of this, because a field about criteria disagreeing can leave the impression that the choice is arbitrary and it is not.
Every design compared here is above 20% efficient on every criterion, and the three usable ones are above 58% on all four. The disagreement is about the last third of the available quality, not about whether a design works. An experimenter who picks the wrong letter runs a slightly larger experiment than they needed to; an experimenter who picks a design whose information matrix is singular learns nothing about curvature at all, whatever letter they had in mind.
So the ordering of importance is: can the design fit the model, is the region the right region, and only then which functional. The field’s own table puts the singular design in it for exactly that reason — it scores zero in every column, and the distance from zero to 0.4758 is much larger than the distance from 0.4758 to 0.9977.
The check, and the refusal that makes it mean something
The disagreement is a comparison across designs and across criteria, so no figure can assert it and it is gated in this field’s library in three parts: the D winner is not the winner on I-optimality, the four criteria fall into exactly two camps with D pairing with G and A with I-optimality, and the gap between the two winners’ I-efficiencies is more than five points rather than merely non-zero.
The comparison is restricted to designs that can actually be run on the region being scored — the singular one, which cannot fit the model, and the rotatable composite in its uncoded form, which is outside the region. Including either would have won the argument by arithmetic rather than by measurement.
The refusal is the uncoded design itself, asserted rather than quietly fixed: scored without being brought into the region first it must exceed a bound whose maximum is 1, because an efficiency above 1 is the only signal available that a comparison has been made between two things that were not comparable. A version of this field that silently rescaled everything would have passed every other check and lost the one number that says the scaling was necessary.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The design for the worst case — both name d-optimality, experimental design, information matrix, optimal design
- The design that refuses the corners — both name central composite design, experimental design, prediction variance, rotatability
- Balancing more than one number — both name experimental design, information matrix, optimal design
- The design that hedges — both name d-optimality, experimental design, optimal design
- The design that stops guessing — both name d-optimality, experimental design, optimal design
- The guess with two numbers in it — both name d-optimality, information matrix, optimal design
Named objects
A flat tag is an object no other essay names yet.
A-optimalityCentral composite designD-optimalityExperimental designG optimalityInformation matrixOptimal designPrediction varianceRotatability