The criterion, and what it assumes

The design that hedges

A locally optimal design is right at one value of the unknown and 23.9% efficient at the edge of a sixteenfold range. Averaging the criterion over a prior instead buys the worst case back to 56.3% — and buys it by adding support points, at spreads the arithmetic decides rather than the experimenter — a third setting at a factor of 3.36 and a fourth at 8.86.

Worth reading first: A design is a number · What a prior is worth.

The previous essay ended with a design that is perfect at one value of a number nobody has and 23.9% efficient at the edge of a sixteenfold range around it. That is not a defect in the design; it is what “locally optimal” means, and the local design is doing exactly what it was asked.

The question is what to ask instead. And the honest answer is that the experimenter has more information than the guess: they have a sense of how wrong the guess might be. A design can use that, and the two standard ways of using it are the two ways of summarising a range — average over it, or protect against its worst point.

A design that is right once, against one that is never wrong by much. Two designs for the same two-parameter model, scored at every true value of K across a range of 16-fold. The local design is the two-point optimum for a guess of K = 1: it reaches 100% there and 66.7% at the worst point of the range. The hedged design maximises the average of log|M| over a prior spanning a factor of 4 either side, uses 3 settings rather than two, and is never below 75.4%. What it costs is 10.4 points at the one value the local design was built for — which is the whole trade, and it is only available to somebody willing to say how wrong the guess might be.
Fig. 1 Two designs for one model scored across a range of true values. The peaked curve reaches 100% at exactly one place. The flatter one never reaches it and never falls far, and what it gives up at the guess is what it buys everywhere else.

Averaging the criterion

The construction is one line. Instead of maximising log|M(ξ, K)| at a guessed K, maximise its average over a prior:

maximise Σₖ πₖ · log|M(ξ, Kₖ)|

The directional derivative averages the same way — δ(t) = Σₖ πₖ·f(t,Kₖ)′M(Kₖ)⁻¹f(t,Kₖ) — so the same multiplicative algorithm runs unchanged with the derivative summed over the prior. Nothing new is required and nothing about the equivalence theorem changes: the search still stops when max δ reaches the number of parameters.

The choice of log|M| rather than |M| is not cosmetic. Averaging determinants would let one value of K with a large determinant dominate the average; averaging their logarithms is averaging on a scale where a factor of two counts the same wherever it happens, which is what a criterion that is trying to be uniformly acceptable ought to do. It is the same reason D-efficiency is reported as a p-th root rather than as a raw determinant.

What averaging the logarithm actually costs the alternative

“Not cosmetic” is worth a number, because Jensen’s inequality says exactly how far apart the two criteria can be.

Take two values of K at which a design’s determinant is 1 and 100. Their arithmetic mean is 50.5; the mean of their logarithms corresponds to a geometric mean of 10. A factor of five, produced by a design that is superb at one value and useless at the other.

So maximising an average of determinants would hand the prize to exactly that design, and maximising an average of logarithms would not. Averaging on the log scale is averaging ratios rather than levels, and since M1/p|M|^{1/p} is the quantity a D-efficiency is a ratio of, the averaged criterion is literally an average of log-efficiencies across the prior.

The stopping rule is also a certificate

One more property of the equivalence theorem is worth stating because it is free and the algorithm has already computed it.

The search stops when maxtδ(t)\max_t \delta(t) reaches p. Before it stops, the same quantity bounds how far from optimal the current design is: any design satisfies

D-efficiency    pmaxtδ(t).\text{D-efficiency} \;\ge\; \frac{p}{\max_t \delta(t)} .

So on six parameters a design whose maximum directional derivative has fallen to 6.3 is guaranteed at least 95% D-efficient, without the optimum ever being computed.

That matters more for the averaged criterion than for the plain one, because the averaged optimum is not in any catalogue and there is nothing to check the answer against. The bound is the check, it holds at every iteration, and it holds for the averaged criterion for the same reason the theorem does — the derivative averages, and the stopping value stays p.

What comes out is a wider design, and the width is an integer

The interesting result is not the efficiency. It is the shape.

A locally optimal design for this two-parameter model has two support points. The hedged design has more — and how many more is an integer that changes at particular prior widths.

A wider prior buys another setting, and the arithmetic says when. The number of distinct settings in the design that maximises the average of log|M| over a prior on the unknown K, against how wide that prior is. It is a staircase because the answer is an integer: a prior reaching a factor of three either side of the guess is still answered by the two settings a local design uses, the third arrives at a spread of 3.36 and the fourth at a spread of 8.86. Neither threshold was put in — both are found by bisecting on the design the algorithm returns. This is what "hedge the guess" means concretely: runs have to be spent at settings that would be right if the guess were wrong, and a two-point design has nowhere to put them.
Fig. 2 The number of distinct settings the hedged design uses, against how wide the prior is. It is a staircase because the answer is an integer, and where the steps fall was found by bisection rather than put in.
  • A prior reaching a factor of three either side of the guess is still answered by two settings — the same two the local design uses, shifted.
  • The third setting arrives at a spread of 3.36.
  • The fourth at 8.86.

Nothing put those thresholds in. They were found by bisecting on the design the algorithm returns, and they are properties of the model and the range rather than choices.

This is the concrete content of “hedge the guess”, and it is more specific than the phrase suggests. Hedging does not mean spreading runs out generally. It means spending runs at settings that would be right if the guess were wrong, and a two-point design has nowhere to put them — so the first thing a prior buys is not better weights, it is another place to stand.

At a spread of four the design is 0.246 of the runs at t = 0.42, 0.283 at t = 1.53, and 0.471 at the end of the range. The middle point is close to where a local design at the guess would have put its lower run; the new point at 0.42 is where a local design would have put it had K been four times smaller. The design is, quite literally, hedging: it is running part of somebody else’s experiment.

What it buys and what it costs

Both designs scored across the range the prior claims, which is the only fair comparison and is a correction this field had to make — scoring a design hedged against a factor of sixteen on a range that only reaches four charges the hedge for runs spent where the comparison never looks, and makes hedging appear worse the more of it is bought.

prior reach local, worst case hedged, worst case hedged, at the guess
×2 90.2% 90.5% 100.0%
×4 66.7% 75.4% 89.6%
×8 42.1% 67.0% 72.2%
×16 23.9% 56.3% 68.7%

Three readings.

At a factor of two, hedging buys nothing. The local design’s worst case is already 90.2% and the hedged design’s is 90.5%, and it costs nothing at the guess because the design has not changed. An experimenter who knows the answer to within a factor of two should stop reading here and run the local design.

At a factor of sixteen it more than doubles the worst case, 23.9% to 56.3%, and costs 31.3 points at the guess. That is a real trade and it is a trade an experimenter can make: the local design is better if the guess is right and catastrophic if it is not, and which of those matters is a question about the consequences rather than about the arithmetic.

And what it costs at the guess grows with the hedge. 0.0, 10.4, 27.8, 31.3 points as the prior widens. Hedging is not free and it is not cheap at the widths where it is most needed.

A design that is right once, against one that is never wrong by muchTwo designs for the same two-parameter model, scored at every true value of K across a range of 256-fold. The local design is the two-point optimum for a guess of K = 1: it reaches 100% there and 23.9% at the worst point of the range. The hedged design maximises the average of log|M| over a prior spanning a factor of 16 either side, uses 4 settings rather than two, and is never below 56.3%. What it costs is 31.3 points at the one value the local design was built for — which is the whole trade, and it is only available to somebody willing to say how wrong the guess might be.00.2500.5000.7501-1-0.50000.5001the true value of K, on a log scaleD-efficiency of the design that was runthe guess, K = 1worst 56%peaked: the design for one guess · flatter: the hedged oneboth designs scored at 7 true values of K, against the local optimum at eachworst case 56.3% against 23.9%
Fig. 3 The two designs across a sixteenfold range. The local design’s curve has collapsed at both ends and the hedged design’s is nearly flat. Drag the prior’s reach to watch the flat curve pay for its flatness.

The comparison this essay had to fix

There is a correction worth recording, because it changed the sign of the conclusion and it was made by the arithmetic rather than by review.

The first version of the comparison scored every design across a fixed range of true values — the same seven values of K whatever the prior said. That is the obvious thing to write and it produces a table in which hedging looks worse the more of it is bought: at a spread of sixteen the hedged design came out at 53.8% worst case against the local design’s 66.7%, which reads as a clear recommendation not to hedge.

It is not a comparison. A design hedged against a factor of sixteen has spent runs at settings that would be right if K were sixteen times smaller, and a scoring range reaching only a factor of four never looks there — so the hedge is charged for the runs and credited with none of their purpose. The range scored has to be the range the prior claims, and once it is, the same numbers run the other way: 23.9% against 56.3%, and hedging wins by more the wider the prior gets.

The general form is one this site keeps meeting. A comparison between two procedures has to be scored on the situation both of them were built for, and a fixed evaluation range is a silent third assumption that belongs to one of them. It is the same error as comparing an adaptive design’s power against a fixed design’s without holding the two to the same size, which the exact field had to make explicit before its power comparison meant anything.

What a wrong guess costs, and which direction is the cheap one. The design for the exponential model is a single run at t = 1/θ, so the efficiency of a design built at a guess θ₀ and used where the truth is θ has a closed form with no design in it: ρ²·exp(2(1−ρ)) at ρ = θ/θ₀. It is 1 at ρ = 1 and falls away on both sides, and not at the same rate — 16.5% at a threefold underestimate of the rate against 42.2% at a threefold overestimate. Underestimating the rate means measuring too late, which is the direction that feels cautious, and it is the expensive one: information about a decay is destroyed by the decay itself, so a run placed past the optimum is measuring something that has mostly happened.
Fig. 4 The one-parameter efficiency curve, for what a prior is being asked to cover. A prior that spans a factor of four either side of the guess is a prior over most of the interesting part of this curve, which is why the hedged design has to change shape rather than only reweight.

The prior is a real input and it is not a small one

A prior that is wrong is worse than no prior, and the arithmetic above says how.

The hedged design is optimal for the range it was told about, and the range is a claim. Told a range of four when the truth is at a factor of ten out, it is a locally optimal design for the wrong neighbourhood plus some scatter — better than a point guess, worse than a design that had been told the truth. There is no width that is safe, because widening the prior is not free either.

This site has measured the analogous thing directly in the Bayesian hierarchical field, and the result there is the one worth carrying: three proper priors on the population spread span 54% of the widest answer at small τ and still 15% at large, so the prior never stops mattering at eight groups. The reassuring story — the data takes over and the prior washes out — is true about whether a difference exists and false about how large it is.

Here it is worse in one respect and better in another. Worse, because the prior is used before any data exists, so there is no data to take over from it; the design is a function of the prior and nothing else. Better, because the consequence is an efficiency rather than a coverage — a design built on a poor prior wastes runs and does not produce an interval that lies about its own coverage.

There is a second difference and it decides how much of this an experimenter should worry about. In the hierarchical field the prior is part of the analysis, so a reader of the published interval inherits it whether or not they agree with it, and the only defence is to report the sensitivity. Here the prior is part of the design, and it is discharged the moment the runs are taken: the analysis of the resulting data need not mention it, the estimate is whatever least squares says, and a reader who thinks the prior was wrong is entitled to say so and has lost nothing but efficiency.

That makes a design prior a considerably softer commitment than an analysis prior, and it is the strongest argument for using one. The cost of a prior in an analysis is measured on this site in observations — a prior is worth a stated number of them, and it moves the answer. The cost of a design prior is measured in efficiency and moves nothing at all, because by the time anybody reads the result the design has already happened.

A design that is right once, against one that is never wrong by much. Two designs for the same two-parameter model, scored at every true value of K across a range of 64-fold. The local design is the two-point optimum for a guess of K = 1: it reaches 100% there and 42.1% at the worst point of the range. The hedged design maximises the average of log|M| over a prior spanning a factor of 8 either side, uses 3 settings rather than two, and is never below 67.0%. What it costs is 27.8 points at the one value the local design was built for — which is the whole trade, and it is only available to somebody willing to say how wrong the guess might be.
Fig. 5 The two designs at a factor of eight, which is the width where the trade is sharpest: the local design has fallen to 42.1% at the edges and the hedged one holds 67.0%, for 27.8 points given up at the guess. Every number in this essay is one point on this picture at one width.
A design that is right once, against one that is never wrong by much. Two designs for the same two-parameter model, scored at every true value of K across a range of 4-fold. The local design is the two-point optimum for a guess of K = 1: it reaches 100% there and 90.2% at the worst point of the range. The hedged design maximises the average of log|M| over a prior spanning a factor of 2 either side, uses 2 settings rather than two, and is never below 90.5%. What it costs is 0.0 points at the one value the local design was built for — which is the whole trade, and it is only available to somebody willing to say how wrong the guess might be.
Fig. 6 The narrowest prior of the four, where hedging buys nothing: the two curves are on top of each other and both are above 90% everywhere. An experimenter who knows the answer to within a factor of two is not in this essay.

What a support point is worth, and why the answer is discrete

The staircase deserves more than a paragraph, because a quantity that changes in integer steps is unusual on this site and the reason it does is instructive.

A design measure is a continuous object: weights are real numbers and can be moved by arbitrarily small amounts. So it is not obvious that anything about the answer should be discrete. What makes the support count an integer is the equivalence theorem, which says the optimum puts weight only where the directional derivative reaches its maximum. As the prior widens, the derivative’s shape changes continuously — and a second local maximum rises continuously towards the ceiling until, at one particular width, it touches. Before that moment the setting gets zero weight; after it, positive weight. The weight is continuous in the width; the support is not.

That also says what happens either side of a threshold, and it is the reassuring part. Just past 3.36 the third support point exists and carries almost nothing, so the design there is barely different from the two-point design just before it. The staircase is a discontinuity in the description rather than in the experiment, and an experimenter who mis-specifies the prior slightly near a threshold gets a design that is slightly different rather than a different design.

A wider prior buys another setting, and the arithmetic says when. The number of distinct settings in the design that maximises the average of log|M| over a prior on the unknown K, against how wide that prior is. It is a staircase because the answer is an integer: a prior reaching a factor of three either side of the guess is still answered by the two settings a local design uses, the third arrives at a spread of 3.15 and the fourth at a spread of 6.09. Neither threshold was put in — both are found by bisecting on the design the algorithm returns. This is what "hedge the guess" means concretely: runs have to be spent at settings that would be right if the guess were wrong, and a two-point design has nowhere to put them.
Fig. 7 The same staircase computed on a wider experimental range. The steps move, because where the design can put runs is part of the problem — a longer range gives the hedge somewhere further out to stand, and the widths at which it needs another setting change accordingly.

Where the runs actually go

It is worth reading the hedged design’s settings rather than only its count, because they say what it is protecting against.

At a spread of four, the three settings carry 0.246, 0.283 and 0.471 of the runs, at t = 0.42, 1.53 and the end of the range. The local design at the guess would have run at 0.83 and the end of the range, half and half. So the hedge has done three things: it has kept the run at the end of the range and given it more weight, it has moved the lower run outward, and it has added a run well inside where the lower one was.

The run at the end of the range surviving is the least surprising and the most useful. It is there for the asymptote, which is the parameter whose best measurement point does not depend on K at all — so it is the one part of the design that needs no hedging, and the criterion recognises that by loading it. The hedge is concentrated entirely in the part of the design that was sensitive to the unknown, which is what a well-behaved answer should look like and is not something the algorithm was told.

The two lower points bracket the local design’s single lower point, which is the geometric statement of what averaging did: one setting that is right at K = 1 has been replaced by two that are right at K = 0.25 and K = 1.5 respectively, and the average of the criterion over the prior prefers the pair.

Protecting the worst point instead

Averaging is one of the two summaries and the other is a maximin criterion: maximise the minimum efficiency across the range, rather than the average of the criterion.

It is the natural thing to want, and it is the harder of the two for the same reason E-optimality is harder than D: a minimum over a range is not a differentiable function of the design, so the multiplicative algorithm has nothing to follow. The two problems are the same problem — a maximin design is the E of this field — and the family’s lesson transfers unchanged: the criterion whose statement is simplest is the one whose optimum sits on a corner.

What is measured here is the averaged design’s worst-case efficiency, which is a lower bound on what a maximin design would achieve. The 56.3% at a spread of sixteen is therefore a floor rather than the answer, and how much a maximin design would improve it is not measured.

What is being claimed here, and what is not

This essay claims Bayesian D-optimality on one two-parameter non-linear model: the averaged criterion, the support count as a staircase in the prior’s width with the two thresholds located by bisection, and the worst-case efficiencies against the local design at four widths.

What stays out: maximin designs, named above and not built; priors that are not three-point discrete approximations, which change the numbers and not the shape; and sequential design, where the prior is updated between runs and the whole problem becomes a different one. Sequential design is the obvious neighbour of everything in these two essays and is the reason the boundary needs stating: it is not a better hedge, it is a different experiment, with an error-rate question of its own that the adaptive field would recognise immediately.

The end of the family is the member the algorithm cannot reach. Every member of the Φₚ family optimised on the same 11² candidates, and two eigenvalues of each answer. The upper curve is the smallest eigenvalue of the information matrix — the quantity E-optimality maximises — which rises from 0.0993 at the D end to 0.1994 at p = 64. The lower curve is the gap between that eigenvalue and the next one up, which falls from 0.0611 to 0.0011. A smallest eigenvalue is not differentiable where it is repeated, and the family is driving the gap to zero: the one criterion here whose meaning fits in a sentence is the one whose optimum sits on a corner of its own surface. The candidate grid is on a slider and it answers a narrower question than it looks. On this square region, refining an odd grid from seven to eleven moves nothing at all — the optimum's support is the corners, the edge midpoints and the centre, and every odd grid from five up contains all of them. An even grid has no centre point and cannot reach the answer at any member of the family. On a disc, where the boundary passes through no grid point, refinement does move it.
Fig. 8 The family curve from this field’s first essay, drawn here for the parallel the section above makes: a maximin design is the E of this problem, and a minimum over a range is not differentiable for the same reason a minimal eigenvalue is not.

The checks

One claim is gated in this field’s library and it has three parts, because any one of them alone would pass for the wrong reason.

The hedged design has more support points than the local one — which fails if the prior is too narrow, and is therefore checked at a width where the design is known to have three.

Its worst efficiency over the range is materially better, by more than five points, which is the part that says the extra points are doing something rather than merely existing.

And it gives something up at the guess, asserted as an inequality in the other direction: the hedged design has to be below the local one at the guess and above 85%. A hedge that cost nothing would be a hedge that had not changed the design, and a check that only ever confirmed the first two parts would pass for a design that had simply been given more runs.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Bayesian optimal designD-optimalityDesign measureExperimental designLocally optimal designMichaelis–MentenThe non-linear modelOptimal designPriorPrior sensitivity