An efficiency that is a ratio
Worth reading first: A design is a number · The design that needs the answer.
Every design in the two optimality fields is chosen by a functional of the whole information matrix. D-optimality maximises its determinant, A-optimality the trace of its inverse, E-optimality its smallest eigenvalue — and all three are statements about the parameter vector, standing in for the questions an experimenter might ask.
An experimenter usually wants one of them. The half-saturation constant of a reaction, the rate of a decay, the point at which a response reaches half its maximum: these are the quantity, and the others are apparatus that has to be estimated because it is in the way. The criterion field builds that distinction for a linear model and gates its theorem at nine parts in a quadrillion. What it leaves is the composition with the field beside it — a subset of the parameters of a model whose information depends on those parameters — and that turns out to be exactly solvable.
The criterion, written so the nuisance is visible
The model is the one both non-linear fields use: η(t) = Vt/(K + t), two parameters, a response that rises and saturates. Write M for the information matrix at a design and split it by parameter. The variance of K̂ when V is estimated alongside it is proportional to
so the thing to maximise is |M| / Mᵥᵥ, the determinant of the whole matrix divided by the information about the parameter nobody asked for. That ratio is the Ds criterion, and writing it this way makes the structure visible: the numerator is what D-optimality maximises, and the denominator is what the nuisance parameter takes away.
The two criteria therefore disagree about exactly one thing — whether information spent on V counts as information — and the disagreement is not small.
The design, in closed form
For a two-point design the criterion collapses. Everything about a setting that either criterion can see enters through g(t) = t/(K + t), and with weights w and 1 − w on settings t₁ and T the criterion reduces to a function of g₁ = g(t₁) and g₂ = g(T) alone. Maximising over the weight first gives w = g₂/(g₁ + g₂); substituting it back leaves
√Ds ∝ g₁g₂(g₂ − g₁)/(g₁ + g₂)
whose maximum over g₁ is the positive root of x² + 2g₂x − g₂² = 0. That root is (√2 − 1)g₂, and at it the weight is g₂/(√2 g₂), which is exactly 1/√2.
So the whole design is a closed form and one of its two halves does not depend on anything:
g(t₁) = (√2 − 1)·g(T) for the setting, and weights 1/√2 and 1 − 1/√2 — whatever K, V and T are.
The equal weights of D-optimality were a consequence of asking about both parameters at once. Every design in the two fields before this one splits its runs evenly between two settings, and it reads as a fact about the model. It is a fact about the question: ask about one parameter and the weights come out 0.7071 and 0.2929, at every setting of everything.
The lower setting moves too. At K = 1 the D-optimal design puts half its runs at 0.8333 and the subset design puts 70.7% of them at 0.6040 — down the curve, where the response is still bending, which is where K is visible. The ratio between the two settings runs from 0.712 at K = 0.25 to 0.756 at K = 4, so it is not itself a constant; the weight is.
The theorem that says it is the answer
A closed form derived by maximising over a two-point family is an answer only if the optimum is a two-point design, and nothing above establishes that. The equivalence theorem does, and it is the same instrument the criterion field uses on a linear model’s subset.
A design is Ds-optimal if and only if
for every setting t, with s the number of parameters of interest — here 1 — and with equality at the support points of an optimal design. The function d says what a further run at t would buy relative to what the design’s own runs are buying; if it exceeded 1 anywhere, a run could be moved there.
Evaluated at the closed form, the maximum of d over four thousand settings is
1.000000000000000 — the theorem holding as an equality at machine precision, which is the same
standard assertTheSubsetTheoremIsExact holds the linear case to. Evaluated at the D-optimal
design, the maximum is 1.6402, at t = 0.6923: there is a setting where a further run would buy
64% more than the design’s own runs, which is what “not optimal for this question” looks like as a
number rather than as a comparison of two efficiencies.
A grid search over settings and weights that was told neither formula lands on the same design, to 0.03 in the setting and 0.005 in the weight, which is the resolution of the grid.
What the two criteria cost each other
Both designs exist; each can be scored by both criteria. The four numbers that come back are the useful summary of the whole field, and two of them are constants.
The D-optimal design is 84.93% efficient for K alone. The subset design is 88.34% efficient for the pair. Measured at K = 0.25, 0.5, 1, 2 and 4 and at two different ceilings, those numbers agree to nine decimals, because every quantity in either criterion is a function of g and both designs sit at fixed points of that scale. They are algebra, not measurements, and this site’s habit of checking a claim at several settings is what makes that visible: a number that is identical at five settings is a number that was never about the settings.
The asymmetry runs the way nobody expects. Narrowing the question costs the wider one 11.7%; widening it costs the narrow one 15.1%. An experimenter who wants K and runs the catalogue’s design for the model gives up more than one who wants K, designs for K, and is later asked about V as well.
The weight is the surprising half
Of the two closed forms, the setting is the one that could have been guessed and the weight is not.
That the design for K moves its lower point down the curve is almost predictable: the response bends near t = K and flattens above it, so information about where the bend is comes from settings near the bend, and a criterion that cares only about K should want more of them. The ratio 0.72 between the two lower settings is a number rather than an intuition, but the direction is not a surprise.
The weight is. Every design in the two optimality fields before this one splits its runs evenly between two support points, and the pattern is so consistent that it reads as a property of two-parameter models. It is not. It is a property of D-optimality: the determinant of a 2 × 2 information matrix built from two rank-one pieces is w(1 − w) times a factor the weights do not enter, and w(1 − w) is maximised at a half whatever the two settings are. The moment the criterion acquires a denominator — Mᵥᵥ, which the weights do enter — the balance goes.
And it goes to a number with no free parameters in it. 1/√2 = 0.7071 is not an approximation, is not a function of the model’s parameters, and does not move when the ceiling moves; it is what maximising x(g₂ − x)/(x + g₂) does. A design whose two settings depend on the unknown and whose two weights do not is an odd object, and it is the one this criterion returns.
What the number means in runs
Fifteen per cent of efficiency is an abstraction until it is converted. Ds-efficiency for one parameter is a ratio of variances, so an efficiency of 0.8493 means the variance of K̂ is 1/0.8493 = 1.177 times what it could have been, and reaching the same precision takes 17.7% more runs. Forty runs at the catalogue’s design do the work of thirty-four at the design for the question.
That is a large number to lose to a choice nobody records having made. It is not lost to ignorance of optimal design — the experimenter who looks up the D-optimal design for the Michaelis–Menten model has done everything the literature asks — but to the gap between the criterion in the catalogue and the question in the laboratory — which is the same gap the design that needs the answer measures for a guess about a parameter rather than for a choice of criterion.
Where the runs go, and why 71% of them
The two settings are doing different jobs and the weights say which is which.
The upper setting, at the ceiling, is the only place the response is near its plateau, so it is where V is pinned down. The lower setting is where the curve is bending, which is where K shows. A criterion that wants both parameters needs both jobs done and splits its runs evenly. A criterion that wants only K still needs V pinned — the nuisance has to be estimated or K̂ inherits its error — but it does not need V pinned as well, only well enough that the estimate of K is not contaminated.
So the subset design keeps the ceiling and starves it: 29% of the runs rather than 50%. What it does with the other 21% is put them where K is visible, and it moves that setting down at the same time, because with more runs there the best place for them is further into the bend.
That is the whole mechanism, and it is why the answer is not “put everything where K is visible”. There is a real trade — a nuisance parameter estimated badly is a parameter of interest estimated badly — and 1/√2 is where it balances.
Both cross-efficiencies in closed form
The two numbers are described above as algebra rather than measurement, on the evidence that they repeat to nine decimals across five values of K and two ceilings. That evidence is strong and the algebra itself is short enough to finish, which turns “they are constants” into “they are these constants”.
Everything either criterion sees enters through g = t/(K + t), and the gradient of the model is proportional to (g, −c·g(1 − g)) with c = V/K. For a two-point design the determinant is therefore w(1 − w)·c²·g₁²g₂²(g₂ − g₁)², from which D-optimality’s answer falls out immediately — w = ½ because w(1 − w) is maximised there, and g₁ = g₂/2 because g₁(g₂ − g₁) is. Dividing by M_VV = w g₁² + (1 − w)g₂² gives the subset criterion, and with a = √2 − 1 its optimum simplifies to a²(1 − a)²/2 = 17 − 12√2 while the D-optimal design scores exactly 1/40.
So the efficiency of the catalogue’s design for the question nobody asked it is
(1/40)/(17 − 12√2) = (17 + 12√2)/40 = 0.8492641
and the D-efficiency of the subset design is the square root of 64(29√2 − 41) = 0.8833862. No K, no V, no ceiling, and nothing left to measure at another setting.
Two things follow that the repeated decimals only hinted at. The constants are properties of the pair of criteria, not of the Michaelis–Menten model — any two-parameter model whose gradient is proportional to (g, h(g)) for a monotone g has the same structure, and the arithmetic above never uses the specific form of h beyond its being g(1 − g). And the 15.07% is exactly reciprocal to a statement about runs: 40/(17 + 12√2) = 1.1775, so the catalogue’s design needs 17.75% more runs, which is the number an experimenter budgets in.
The ratio that is constant is not the one in t
The essay notes that the ratio between the two designs’ lower settings is not itself a constant — 0.712 at K = 0.25 and 0.756 at K = 4 — and the closed forms say why, and say what the constant actually is.
In g the two lower settings are g₂/2 and (√2 − 1)g₂, so their ratio is 2(√2 − 1) = 0.8284, at every K, every V and every ceiling. At K = 1 with a ceiling of 10 that is 0.37656 against 0.45455, which recovers t₁ = 0.6040 and 0.8333 exactly.
The drift from 0.712 to 0.756 is therefore not a feature of the design problem at all. It is the map back from g to t = Kg/(1 − g), which is non-linear, so a fixed ratio in g becomes a varying one in t — and which of the two an experimenter sees depends only on which axis the design is reported on.
That is worth a sentence beyond the arithmetic, because it is a general hazard in reading design tables. A design’s settings are reported in the units the experiment is run in, and the criterion lives in a transformed coordinate. A quantity that looks like it depends on the parameters may be a constant seen through the transformation, and the only way to tell is to write the criterion in its own variable — which here takes one substitution and turns a drifting ratio into 2(√2 − 1).
Which parameter, and the answer changes
Nothing above is special to K. The same criterion with the roles reversed — |M|/Mₖₖ, protecting V and treating K as the nuisance — is a different design, and the fact that the two are different is the whole content of the subset idea. A design cannot be described as “for this model”; it is for a stated question about the model.
The general lesson survives the model. For any two-parameter non-linear model, the subset criterion weights the settings by how much each one distinguishes the parameter of interest from the nuisance rather than by how much information it carries overall — and it is that distinguishing that the equal weights of D-optimality were never doing.
The nuisance cannot be dropped, only conditioned on
There is an obvious shortcut and it does not work, which is worth measuring because the shortcut is what an experimenter reaches for when told to design for one parameter.
The shortcut is to maximise the information about K on its own: Mₖₖ, the entry of the matrix that is about K. That drops the nuisance block instead of conditioning on it, and the design it produces puts every run at t = K, which is where the derivative with respect to K is largest.
At that design the two parameters’ derivatives are proportional, the information matrix is singular, its determinant is −1.7 × 10⁻¹⁸, and the variance of K̂ is infinite. The criterion that chose the design reports its largest possible value, and the design is 0.00% efficient for the parameter it was chosen for. Treating a nuisance parameter as known is not a conservative simplification: it is a different problem whose answer happens to be a design that cannot answer the original one.
What is claimed here, and what is not
This essay claims Ds-optimal design for one parameter of a non-linear model: the closed form, the equivalence theorem that certifies it, and the two cross-efficiencies, which are constants.
What stays out and is named as a decision: subsets of more than one parameter in a model with more than two, where the criterion carries an s-th root and the weights stop being a single number; c-optimality, which asks for a linear combination of parameters rather than a subset of them and is the natural neighbour of this criterion; and the whole question of a design for K over a range of values of K, which is the next essay and is not a special case of this one. What a design that must also decide when to stop costs is further along.
The checks, and the refusal that makes them mean something
Three claims are gated in this field’s library. The weight on the lower setting is required to be 1/√2 to twelve decimals and the setting to solve g(t₁) = (√2 − 1)g(T) to the same, at three values of K and two ceilings. The equivalence theorem is required to hold as an equality at each of them, and a grid search told neither formula is required to find the same design. And the two cross-efficiencies are required to be identical across five values of K and two ceilings, which is the check that turns a measurement into a statement about algebra.
The refusal is the shortcut: a design chosen by maximising the information about K alone, which puts every run at one setting and makes the variance of the parameter it was chosen for infinite. The check requires that failure, because a criterion that could not tell that design apart from a good one would not be measuring anything.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The design for the worst case — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, locally optimal design, michaelis–menten, the non-linear model, optimal design
- The design that hedges — both name d-optimality, design measure, experimental design, locally optimal design, michaelis–menten, the non-linear model, optimal design
- The family behind the letters — both name d-optimality, design measure, ds-optimality, equivalence theorem, experimental design, information matrix, optimal design
- Augmenting a design that has already run — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design
- The criterion with no derivative — both name d-optimality, design measure, equivalence theorem, experimental design, information matrix, optimal design
- The design that stops guessing — both name d-optimality, experimental design, locally optimal design, michaelis–menten, the non-linear model, optimal design
Named objects
A flat tag is an object no other essay names yet.
Closed formD-optimalityDesign measureDs-optimalityEquivalence theoremExperimental designInformation matrixLocally optimal designMichaelis–MentenThe non-linear modelNuisance parameterOptimal designSubset of parameters