Concept

Optimal design — where it appears

An arrangement of experimental runs chosen by maximising a stated functional of the information matrix, rather than taken from a catalogue. Which functional is chosen is a statement about which parameters the experiment is for, and two reasonable choices can rank the same designs oppositely.

Named by 20 essays across 7 fields — each of them below, with the objects they name alongside it.

Where a D-optimal design puts its runs. The D-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.1458, 0.0962, 0.0802, which nine equal runs cannot express.

A design is a number

A standard design is taken from a catalogue and then measured. Turn the arithmetic round and a design becomes the answer to an optimisation — and over 121 candidate settings the search keeps nine of them, which are exactly the nine a catalogue would have offered, at weights nine equal runs cannot express.

optimality · Criterion
Two designs for one model, and the weights are not equal. Where the runs go, for the same two-parameter model at K = 1 and a ceiling of T = 10. The D-optimal design for both parameters is the familiar one: half the runs at 0.8333 and half at the ceiling. The design for the half-saturation constant alone moves the lower setting down to 0.6040 — where the response curve is still bending, which is where K is visible — and, unlike every design in the two fields before this one, it does not split the runs evenly: the weights are exactly 1/√2 and 1 − 1/√2, 0.7071 and 0.2929, at every K, V and T. The equal weights of D-optimality were a consequence of asking about both parameters at once, and nobody had to notice while that was the only question being asked.

An efficiency that is a ratio

A design chosen for a model is not a design chosen for the parameter somebody wanted. Asking for one of two parameters moves the runs, unbalances the weights, and costs the other question exactly 15.07% — at every setting, because it is algebra.

guarantee · Criterion
A design that is right once, and one that is never wrong by much. Three designs for the Michaelis–Menten model, scored at every true value of K across a 16-fold range. The peaked curve is the two-point local design built at K = 1: 100% there and 66.7% at the worst point of the range. The curve just under it is the design that averages the criterion over a uniform prior on the same range, which is barely different — 67.9% at worst — because averaging is dominated by the middle of the range where the local design is already good. The flat line is the maximin design: 3 settings, never above 80.8% and never below 78.8%. Its worst case is 12.0 points better, and what it gives up is the 21.2 points at the one value the local design was built for.

The design for the worst case

A design for a non-linear model is optimal at a guess about the answer. Averaging over a prior repairs that on average; protecting the worst value in a range is a different problem, with a different answer, and it needs a third setting to reach it.

robust · Local design
The end of the family is the member the algorithm cannot reach. Every member of the Φₚ family optimised on the same 11² candidates, and two eigenvalues of each answer. The upper curve is the smallest eigenvalue of the information matrix — the quantity E-optimality maximises — which rises from 0.0993 at the D end to 0.1994 at p = 64. The lower curve is the gap between that eigenvalue and the next one up, which falls from 0.0611 to 0.0011. A smallest eigenvalue is not differentiable where it is repeated, and the family is driving the gap to zero: the one criterion here whose meaning fits in a sentence is the one whose optimum sits on a corner of its own surface. The candidate grid is on a slider and it answers a narrower question than it looks. On this square region, refining an odd grid from seven to eleven moves nothing at all — the optimum's support is the corners, the edge midpoints and the centre, and every odd grid from five up contains all of them. An even grid has no centre point and cannot reach the answer at any member of the family. On a disc, where the boundary passes through no grid point, refinement does move it.

The family behind the letters

A, D and E are not three ideas. They are three points of one family with a single dial, and running the dial from one end to the other doubles the smallest eigenvalue of the information matrix while closing the gap above it fifty-three-fold — which is the family driving its own last member to the place where it stops being differentiable.

criteria · Criterion
One curve, three designs, and the disagreement is the weights. The compartmental response exp(−θ₁t) − exp(−θ₂t) at the guess θ = (0.2, 1.2), with three designs underneath it. Each mark is a setting and its height is the share of the experiment spent there. The D-optimal design splits the runs equally between two settings — that is the determinant's answer and it is equal for every model of this kind. The design for the first parameter alone puts 1.9% of the experiment at its early setting and the rest at its late one, because the early runs are there to identify the nuisance and nothing more. The design that protects a 4-fold rectangle needs 3 settings and spends 9.1% at the earliest of them.

The guess with two numbers in it

Every optimal design for a non-linear model is optimal at a guess. Where the model has one parameter that moves the settings, that guess is a number and everything about it comes out in closed form; where it has two, three constants become functions and one of them becomes zero.

blind · Local design
Four designs, scored on the one parameter that was wanted. Every design scored by its Ds-efficiency for K at 13 true values across a 16-fold range. The peaked curve is the subset design built at the guess K = 1: 100% there and 42.1% at the worst point of the range. The flat curve is the maximin-Ds design, never above 64.8% and never below 61.2%. Between them is the maximin design for the pair — a robust design, protecting something else, and worth 41.9% at worst here. The lowest curve is the D-optimal design at the guess, which is what an experimenter who wanted K and looked up a design for the model would actually run: 29.1% at the worst point, against 61.2% available.

Protecting one parameter over a range

A design for a non-linear model is optimal at a guess. A design for one of its parameters over a range of guesses is a worst case of a ratio of two determinants, and it is not a special case of either problem it is made of.

guarantee · Local design
Two rates, not a factor. The standard deviation of the covariate imbalance under three rules, at five trial sizes, 260 trials each, on log axes. The upper line is a coin: its slope is -0.489, against a closed form of exactly −½. The middle line is minimisation on a median split; its slope is -0.519 — the same rate — because inside a category the assignment is still a coin, and what it buys is the constant, 0.654 of a coin's at n = 200. The lower line is the rule that reads x and maximises the information about the treatment effect: slope -0.987, nearly twice as steep. Its advantage is therefore not a number that can be quoted — it is 0.258 of a coin's at n = 50 and 0.065 at n = 800, and it keeps going.

The rule that reads the number

Stop categorising and let the rule read the covariate itself. What it should minimise is not an invented distance but the variance of the effect being estimated — and what comes back is not a better constant but a different rate.

continuous · Assignment
The variance touches p and never crosses it. d(x) = f(x)′M⁻¹f(x) along the diagonal of a square region, for the D-optimal measure. The line at 6 is the number of parameters in the model. Kiefer and Wolfowitz's theorem says a design is D-optimal exactly when the largest d anywhere in the region is p — not approximately, equals — so the optimal curve is tangent to that line at its support points and below it everywhere else. Here the largest value anywhere on a 41×41 grid is 6.000000000.

The theorem that says when to stop

A search that maximises the volume of the information has no way of knowing it has finished, because nothing tells it what the maximum is. Kiefer and Wolfowitz's equality does — a design is D-optimal exactly when the worst prediction anywhere in the region equals the number of parameters, which is 6.000000000059 here, gated at machine precision.

optimality · Equivalence
Where a Ds-optimal design puts its runs. The Ds-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.2500, 0.1250, 0.0625, which nine equal runs cannot express.

The two terms anybody wanted

D-optimality estimates all six parameters of a quadratic as precisely as possible. Nobody wants that. An experimenter looking for a maximum wants the two curvature terms, and the design that gives them is not the D-optimal one — it is a quarter of the runs at the centre, exactly, and the D-optimal design is 75.3% efficient for the question that was actually asked.

criteria · Criterion
Four designs, and what each of them guarantees. The worst Ds-efficiency each design achieves anywhere in a rectangle of parameter values 4 times wide in each coordinate. The design built at the guess guarantees 7.3%: it is perfect where it was built and nearly useless at one corner. The D-optimal design at the same guess guarantees 25.6% — it answers the wrong question everywhere and is therefore not concentrated on being right anywhere. The third is the one this field exists to measure: maximin over the first parameter's whole range, with the second held at its guess. It guarantees 7.0%, which is no better than the design that protects nothing. Protecting both is worth 45.9%, and it needs 3 settings to do it.

The worst case in two directions

A design that protects a range of one parameter is robust. Protect the range of one parameter while holding the other at a guess and the design is still robust, still has a guarantee, and guarantees no more than a design that protects nothing at all.

blind · Criterion
One curve is a binomial coefficient and the other is a line. The number of subsets a maximin over this dictionary would have to score, against the number the exchange algorithm actually scores. At three functions the walk is 2,024 subsets and is the honest answer; at eight it is 735,471 and the exchange algorithm has looked at 421. The warrant for the second curve is the four sizes where both exist and agree, which is a weak warrant — it says the algorithm has not yet been wrong, not that it cannot be — and it is the only one available past the point the first curve leaves the page.

Where the enumeration stops

A maximin over an eight-function dictionary is a walk over seventy subsets. Over twenty-four it is 735,471 at eight functions, and the exchange algorithm that replaces the walk scores 421. What licenses the second curve is four sizes where both exist and agree, which is a weaker warrant than it looks.

product · Optimum
Four criteria on 5 designs, on a square region. Each design scored as an efficiency — its value over the best attainable — so four criteria in as many different units sit on one scale where 1 is the optimum. Rows are ordered by D. D picks 13-run exchange; A picks face-centred composite; G picks 13-run exchange; I picks face-centred composite. Every design has been scaled to just fit the region first, because a design run at settings the region does not contain is not a competitor on it. The disagreement is the point: the letter is a choice, and it is almost never reported as one.

Four letters and two camps

D, A, G and I are four ways of turning one matrix into one number, and they do not agree. The design that wins on D is the worst thing here on I. And the same two designs swap places entirely when the region changes from a square to a disc — on all four criteria at once.

optimality · Criterion
Where to look depends on the answer. The information a single run at time t carries about the rate of an exponential decay, (∂η/∂θ)² = t²·exp(−2θt), at three values of θ. Each curve has one maximum and it is at t = 1/θ exactly — marked, and found by a search over 8,001 settings that was never told the formula. Nothing in a linear model behaves this way: there the information matrix is X′X and the parameters are not in it, so a design can be chosen once and used whatever the answer turns out to be. Here the design is optimal at a guess, and the three curves are three different experiments for one model.

The design that needs the answer

Every design this site has computed is optimal whatever the experiment turns out to say, because X′X does not contain the parameters. For a non-linear model it does, so the best place to take a measurement is a function of the number the measurement exists to find — and guessing it three times too low costs two and a half times more than guessing it three times too high.

criteria · Local design
One experiment finding out where to look. A single run of the fully sequential design: 40 runs, the first 8 placed at the guess K = 1, then the model refitted and the design revised after every 2. The marks are the settings the runs were made at. The horizontal lines are where a design built at the truth K = 3 would have put them — 1.875 and 10.00 — and the rule walks onto them without being told: its estimate of K after the first eight runs was 2.694, and by the end 2.765 against a truth of 3. The whole experiment is 96.5% as efficient as the design that knew the answer, where running all 40 at the guess would have been 81.1%.

The design that stops guessing

Every repair so far protects a guess. The alternative is to run part of the experiment, estimate the parameter from it, and design the rest at the estimate — which recovers most of what a threefold wrong guess costs, and has a best moment to stop guessing that is earlier than anyone expects.

robust · Local design
What balancing several numbers at once costs each of them. The criterion generalises without a word changing — the covariate imbalance becomes a vector and the correction a quadratic form — so the question is what it is worth rather than whether it can be done. At n = 200 with 200 trials per point, a rule balancing one covariate leaves 12.7% of a coin's imbalance in it; balancing eight leaves 23.2% in each. The assignment has a fixed amount of freedom and every covariate added takes a share of it. The rule degrades rather than failing: at eight covariates it is still four times better balanced than a coin, and the eight are being held simultaneously rather than in turn.

Balancing more than one number

The criterion generalises to several covariates without a word changing, which makes the question what it is worth rather than whether it can be done. Each one added takes a share of the assignment's freedom, and the imbalance left in every one of them rises.

continuous · Blocking
Adding a run can make the design worse. The D-efficiency of the best N-run design at each size, against the optimal measure. It is not a rising curve. 13 runs reaches 99.77% and 14 falls to 99.44%, because the optimal weights are real numbers and N runs is an integer approximation to them, so how good a design can be depends on how well N divides. A Wald interval behaves the same way: a larger sample sometimes makes its coverage worse, for exactly this reason.

The design that has to be integers

The optimal design is a set of real weights and an experiment is a set of runs, so the theory's answer is never available. Thirteen runs reach 99.77% of it and fourteen reach 99.44% — adding a run makes the design worse per run, and the search that finds it does not always find the same one.

optimality · Criterion
A design that is right once, against one that is never wrong by much. Two designs for the same two-parameter model, scored at every true value of K across a range of 16-fold. The local design is the two-point optimum for a guess of K = 1: it reaches 100% there and 66.7% at the worst point of the range. The hedged design maximises the average of log|M| over a prior spanning a factor of 4 either side, uses 3 settings rather than two, and is never below 75.4%. What it costs is 10.4 points at the one value the local design was built for — which is the whole trade, and it is only available to somebody willing to say how wrong the guess might be.

The design that hedges

A locally optimal design is right at one value of the unknown and 23.9% efficient at the edge of a sixteenfold range. Averaging the criterion over a prior instead buys the worst case back to 56.3% — and buys it by adding support points, at spreads the arithmetic decides rather than the experimenter — a third setting at a factor of 3.36 and a fourth at 8.86.

criteria · Local design
What visiting fewer settings costs, 6 parameters. Carathéodory's bound puts the support of an optimal measure between 6 and 21. 6 settings: D-efficiency 88.90%, G-efficiency 57.18%, 0 degrees of freedom for lack of fit; 7 settings: D-efficiency 94.54%, G-efficiency 61.22%, 1 degrees of freedom for lack of fit; 8 settings: D-efficiency 95.99%, G-efficiency 64.60%, 2 degrees of freedom for lack of fit; 9 settings: D-efficiency 97.40%, G-efficiency 82.76%, 3 degrees of freedom for lack of fit. The saturated design has none, and buying the first one costs about five points of efficiency to get back.

How many places a design goes

Carathéodory's bound puts an optimal design's support between six and twenty-one settings, and every design in this field that can fit the model visits exactly nine. The count is not a choice anybody makes, it decides how many degrees of freedom are left to check the model with, and the first spare setting costs six points of efficiency to get back.

optimality · Equivalence
The certificate when 4 runs are already spent. a 2² factorial has already been run and 2 further runs are to be placed. The stationarity condition is no longer max d = p; it is max d = (p − λ·tr(M⁻¹M_fixed))/(1 − λ) with λ = 0.6667 the share of runs already spent, which is 7.0985 here. The search reaches 7.098508790 against it, and the largest value anywhere on a 41×41 grid is 7.098508790.

Augmenting a design that has already run

The equivalence theorem still certifies when some runs are already spent, and one number in it changes: the bound is no longer p but (p − λ·tr(M⁻¹M_fixed))/(1 − λ). It equals p again exactly when the runs already made can still be absorbed into the design that would have been chosen — so the certificate says whether the experiment is still recoverable.

optimality · Equivalence
Where each criterion's optimum puts the information. the D-optimal design's smallest eigenvalue is 0.09927, attained once; the A-optimal design's smallest eigenvalue is 0.16516, attained once; the I-optimal design's smallest eigenvalue is 0.17541, attained once; the E-optimal design's smallest eigenvalue is 0.19999, attained 2 times. A criterion that reads the smallest eigenvalue has no derivative where that eigenvalue is repeated, and the E-optimal design is exactly there.

The criterion with no derivative

E-optimality maximises the smallest eigenvalue of the information matrix, and at its own optimum that eigenvalue is attained twice — which is exactly where the function has a corner. The multiplicative search this field's other three criteria use assumes a derivative that is not there, and stops at 37.2% of the optimum.

optimality · Equivalence

Named alongside it

The objects these essays reach for when they reach for this one.

D-optimalityExperimental designInformation matrixDesign measureEquivalence theoremThe non-linear modelDs-optimalityLocally optimal designMichaelis–MentenPrediction varianceBayesian optimal designNuisance parameter

All concepts