Field

The criterion, and what it assumes

D-optimality gave the catalogue back, which was reassuring, and left two things unsaid. A, D and E are one family and the letter is a choice — the design that wins the first is nearly worst at the last. And for a model whose information depends on its own parameters, a design is optimal only at a guess about the answer, which makes the cost of guessing wrong a quantity with a closed form.
The end of the family is the member the algorithm cannot reach. Every member of the Φₚ family optimised on the same 11² candidates, and two eigenvalues of each answer. The upper curve is the smallest eigenvalue of the information matrix — the quantity E-optimality maximises — which rises from 0.0993 at the D end to 0.1994 at p = 64. The lower curve is the gap between that eigenvalue and the next one up, which falls from 0.0611 to 0.0011. A smallest eigenvalue is not differentiable where it is repeated, and the family is driving the gap to zero: the one criterion here whose meaning fits in a sentence is the one whose optimum sits on a corner of its own surface. The candidate grid is on a slider and it answers a narrower question than it looks. On this square region, refining an odd grid from seven to eleven moves nothing at all — the optimum's support is the corners, the edge midpoints and the centre, and every odd grid from five up contains all of them. An even grid has no centre point and cannot reach the answer at any member of the family. On a disc, where the boundary passes through no grid point, refinement does move it.

The family behind the letters

A, D and E are not three ideas. They are three points of one family with a single dial, and running the dial from one end to the other doubles the smallest eigenvalue of the information matrix while closing the gap above it fifty-three-fold — which is the family driving its own last member to the place where it stops being differentiable.

Where a Ds-optimal design puts its runs. The Ds-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.2500, 0.1250, 0.0625, which nine equal runs cannot express.

The two terms anybody wanted

D-optimality estimates all six parameters of a quadratic as precisely as possible. Nobody wants that. An experimenter looking for a maximum wants the two curvature terms, and the design that gives them is not the D-optimal one — it is a quarter of the runs at the centre, exactly, and the D-optimal design is 75.3% efficient for the question that was actually asked.

Where to look depends on the answer. The information a single run at time t carries about the rate of an exponential decay, (∂η/∂θ)² = t²·exp(−2θt), at three values of θ. Each curve has one maximum and it is at t = 1/θ exactly — marked, and found by a search over 8,001 settings that was never told the formula. Nothing in a linear model behaves this way: there the information matrix is X′X and the parameters are not in it, so a design can be chosen once and used whatever the answer turns out to be. Here the design is optimal at a guess, and the three curves are three different experiments for one model.

The design that needs the answer

Every design this site has computed is optimal whatever the experiment turns out to say, because X′X does not contain the parameters. For a non-linear model it does, so the best place to take a measurement is a function of the number the measurement exists to find — and guessing it three times too low costs two and a half times more than guessing it three times too high.

A design that is right once, against one that is never wrong by much. Two designs for the same two-parameter model, scored at every true value of K across a range of 16-fold. The local design is the two-point optimum for a guess of K = 1: it reaches 100% there and 66.7% at the worst point of the range. The hedged design maximises the average of log|M| over a prior spanning a factor of 4 either side, uses 3 settings rather than two, and is never below 75.4%. What it costs is 10.4 points at the one value the local design was built for — which is the whole trade, and it is only available to somebody willing to say how wrong the guess might be.

The design that hedges

A locally optimal design is right at one value of the unknown and 23.9% efficient at the edge of a sixteenfold range. Averaging the criterion over a prior instead buys the worst case back to 56.3% — and buys it by adding support points, at spreads the arithmetic decides rather than the experimenter — a third setting at a factor of 3.36 and a fourth at 8.86.

All essays