The surface between the corners

The optimum is a ratio, and its interval is sometimes the whole line

The best setting is −b₁/2b₂: a ratio of two estimates whose denominator is a curvature the design can often barely see. The delta method reports a finite interval every time and covers 68.8% where the curvature is weak; Fieller's set covers 95% and says so by being unbounded.

Worth reading first: Walking up the gradient · One factor at a time · What the 95% refers to.

An experiment run to find the best setting of a factor ends with a fitted quadratic and one number read off it. Differentiate b₀ + b₁x + b₂x², set the derivative to zero, and the maximum is at

x* = −b₁ / 2b₂

which is a ratio of two estimates. Everything unpleasant in this essay follows from that and from nothing else.

Where the maximum is, from 15 runsOne dataset, one fitted quadratic, and two answers to "where is the best setting". The delta method reports 0.80 ± 0.46, a finite interval it will report whatever the data does. Fieller's set is 0.49 to 1.76, because the curvature here has t = -4.04. The true optimum is at 0.75.506070-2-1012the factor, in coded unitsresponsethe true optimum, 0.75delta methodFieller15 runs at 5 levels, σ = 2curvature t = -4.04
Fig. 1 Fifteen runs at five levels, one fitted quadratic, and two answers to “where is the best setting”. One of them is an interval and the other is everything outside one.

Why a ratio is different

The two coefficients are well behaved. Each is a linear combination of the observations, each is normally distributed if the errors are, each has a standard error the design fixes in advance and a t interval that covers exactly what it claims. Nothing about b₁ or b₂ is difficult.

Their ratio is a different object. Divide by a quantity whose distribution puts mass near zero and the result has no mean, no variance, and tails that decay like a Cauchy’s. That is not an approximation breaking down at the edges; it is the shape of the distribution.

And the denominator here is not an arbitrary parameter. It is the curvature — the same quantity the curvature test has 21% power to detect against one standard deviation at the design most people run. A design that has just barely established that there is a maximum is a design whose estimate of where it is has a denominator indistinguishable from zero.

The delta method, and what it delivers

The standard answer is to linearise. Expand the ratio to first order around the true coefficients, and the variance of x̂* comes out as a combination of the variances of b₁ and b₂ and their covariance, divided by 4b₂² and 4b₂⁴. Take the square root, multiply by a t quantile, and report an interval.

It is easy, it is what software does, and it has one property worth stating precisely: it always returns a finite interval. Whatever the data, whatever the curvature, however close the denominator came to zero, out comes a centre and a width.

Counted over twelve thousand datasets of fifteen runs at five levels, at four true curvatures:

curvature β₂ true optimum delta-method coverage Fieller coverage
−8 0.38 95.1% 95.0%
−4 0.75 93.0% 95.2%
−2 1.50 88.6% 95.2%
−1 3.00 79.4% 95.2%
−0.6 5.00 68.8% 95.1%

Where the curvature is sharp the delta method is fine. Where it is weak — where the experiment has barely established that there is an optimum at all — an interval labelled 95% covers 68.8% of the time.

Two intervals for the same optimum, counted. 6,000 datasets at each curvature, with both constructions applied to each. Fieller's set covers 95.1% on average and never departs from its claim; the delta method covers 68.4% where the curvature is barely visible and 95.3% where it is sharp. It fails in the direction that matters — too narrow, so it understates how little the experiment has settled.
Fig. 2 Both constructions counted on the same datasets, against how curved the truth is. The flat line is Fieller. The one that falls away is the delta method, and it falls away exactly where the experiment knows least.

Fieller’s construction, which asks a different question

There is an exact answer, and it comes from refusing to make a ratio at all.

The optimum is the x at which the derivative is zero, and the derivative at any fixed x is

b₁ + 2b₂x

which is a linear combination of the coefficients. Linear combinations are easy: it has a normal distribution, a standard error the design supplies, and an exact t test at every x.

So define the interval as the set of x at which that test does not reject. Square both sides of the comparison, and the condition becomes a quadratic inequality in x:

Ax² + Bx + C ≤ 0, with A = 4b₂² − 4t²Var(b₂)

which is solvable in closed form. Counted on the same datasets, the resulting set covers 95.0% to 95.2% at every curvature in the table above.

The sign of A is the whole of it

Look at what A is. It compares b₂² with t²Var(b₂) — which is exactly the curvature’s own t statistic compared with its critical value. So:

A > 0 means the curvature is significant at this level, and the set is an interval between the two roots. An ordinary answer.

A < 0 means it is not, and the solution set of the inequality is the outside of an interval: everything below one root and everything above the other. The experiment is saying the optimum is not in the middle stretch, and beyond that it does not know.

And when there are no real roots at all, the set is the whole line, or empty.

None of those is a numerical failure. They are the honest answers to a question the data cannot narrow, and the reason the construction covers what it claims is precisely that it is allowed to give them.

Where the maximum is, from 15 runs. One dataset, one fitted quadratic, and two answers to "where is the best setting". The delta method reports 2.20 ± 3.31, a finite interval it will report whatever the data does. Fieller's set is everything outside -4.48 to 0.86, because the curvature here has t = -1.46. The true optimum is at 3.00.
Fig. 3 The same fifteen runs against a truth curved four times more gently. The delta-method interval is narrower than the region the design covers; Fieller’s answer is everything outside a stretch in the middle.

How often the honest answer is unbounded

That property has a cost and the cost should be printed rather than glossed.

Two intervals for the same optimum, counted. 6,000 datasets at each curvature, with both constructions applied to each. Fieller's set covers 95.1% on average and never departs from its claim; the delta method covers 68.4% where the curvature is barely visible and 95.3% where it is sharp. It fails in the direction that matters — too narrow, so it understates how little the experiment has settled.
Fig. 4 The share of datasets on which Fieller’s set is unbounded, against the curvature. The line along the bottom is the delta method, which is unbounded on none of them.

At a curvature of 4 — where the delta method covers 93.0% and looks nearly respectable — the set is unbounded on 16.0% of datasets. At 2, on 68.5%. At 0.6, on 93.0%.

So on the great majority of experiments in the weak-curvature regime, the exact answer to “where is the best setting” is not in this stretch, and otherwise unknown. That is unsatisfying, and it is what the experiment established.

The delta method’s finite interval on those same datasets is not additional information. It is the same information with the uncertainty deleted: a median width of 4.16 in coded units at β₂ = −2, on a design whose runs span −1 to +1. An interval four design-widths across is not a location; it is a statement that the location is unknown, printed in a format that looks like a location.

The unbounded share is the curvature test, complemented

The sign of A is the curvature test’s verdict, so the share of datasets on which Fieller’s set is unbounded is one minus that test’s power — and the field therefore measures the power of the curvature test without setting out to.

At a curvature of 4 the set is unbounded on 16.0% of datasets, so the curvature test rejects on 84.0%. At 2 it rejects on 31.5%. At 0.6, on 7.0%.

Those three numbers are worth having on their own, because they are the design’s ability to establish that there is a maximum at all, measured on the same fifteen runs at five levels that everything else here is measured on. And they sit beside the delta method’s coverage in a way that is easy to read: 93.0% coverage where the design detects curvature five times in six, 88.6% where it detects it once in three, 68.8% where it detects it once in fourteen.

The delta method is trustworthy exactly where the curvature test is nearly certain, and nowhere else — which is a sharper condition than “where the curvature is sharp”, because it is checkable on the data in hand rather than against a truth nobody has.

One thing the sweep cannot separate

The five rows vary the curvature and they move something else at the same time, and it is worth naming because it bounds what the coverage column can be read as.

Multiplying each true optimum by its curvature gives 3.04, 3.00, 3.00, 3.00 and 3.00. So the linear coefficient is held fixed across the sweep and the optimum is 3/|β₂| — which is why it runs from 0.38 at the sharpest curvature to 5.00 at the weakest.

Five units is far outside any five-level design on this scale. So the weak-curvature rows are not only rows where the denominator is near zero; they are rows where the quantity being estimated is not inside the experimental region at all, and any interval for it is an extrapolation from a quadratic fitted somewhere else entirely.

That does not damage the comparison — both constructions face the same problem on the same datasets, and Fieller covers 95% at every row while the delta method does not. It does mean the sweep cannot say which of the two difficulties the delta method is failing on. A design that moved the curvature while holding the optimum fixed, by scaling the linear coefficient with it, would separate them, and this one does not try.

The honest summary is therefore narrower than the coverage column looks. Where an experiment has weak curvature, the optimum is usually outside the region it explored, and the delta method’s interval is both too narrow and pointed at somewhere the data has no information about. Those two failures arrive together in practice as well as in this sweep, which is a defence of the design rather than of the reading.

Where the point estimate lands

The interval is the second thing to go wrong. The first is the estimate itself.

Across the same datasets at β₂ = −2, whose true optimum is at 1.50, the middle 90% of the fitted optima run from −1.96 to 7.02, with a median of 1.41. 78.5% of them land outside the region the design actually explored, which means the reported best setting is an extrapolation from a quadratic fitted to data that says nothing about the response there.

At β₂ = −0.6 the middle 90% runs from −17.60 to 18.59. Those are coded units on a design whose runs go from −1 to +1.

A distribution with that shape has no useful centre, which is why the median is quoted rather than the mean: the mean does not exist. And a sampling distribution with no mean is not a technicality about moments — it is the reason the delta method fails, since the delta method’s whole content is a first-order approximation to a variance that is infinite.

Two factors, and a third way to fail

In two factors the stationary point is −½B̂⁻¹b̂, and a matrix inverse in the denominator adds a failure the one-factor case does not have: the fitted surface can have the wrong shape.

400 fits of the same surface, σ = 3. Each point is the stationary point of one fitted quadratic, from one central composite design run on the same true surface — which has a maximum at (0.91, 0.35), marked. 8% of the fits are saddles rather than maxima, so their stationary point is not an optimum of anything, and 39% land outside the region the design explored. The median distance from the centre is 1.18 against a true 0.98.
Fig. 5 Four hundred central composite designs run on the same true surface, which has a maximum inside the region. Each point is one fitted stationary point, coloured by whether the surface it came from is a maximum, a saddle or a minimum.

The eigenvalues of B̂ decide it: both negative is a maximum, mixed signs a saddle. On a true surface with a genuine interior maximum, the share of fits that come out as saddles is 0% at σ = 1, 2% at σ = 2, 10% at σ = 3 and 22% at σ = 4.

A saddle’s stationary point is a real number, computed correctly, and it is not an optimum of anything. It is the point where the surface stops rising in one direction and stops falling in another. Software reports it, an experimenter reads it as a recommended setting, and nothing about the number says which of the two shapes produced it.

Meanwhile the share of fitted optima landing outside the design region goes 6%, 27%, 40%, 45%, and the median distance from the centre at σ = 4 is 1.28 against a true 0.98 — the estimate wanders outward, because the ratio’s denominator is small more often than it is large.

400 fits of the same surface, σ = 1. Each point is the stationary point of one fitted quadratic, from one central composite design run on the same true surface — which has a maximum at (0.91, 0.35), marked. 0% of the fits are saddles rather than maxima, so their stationary point is not an optimum of anything, and 6% land outside the region the design explored. The median distance from the centre is 1.01 against a true 0.98.
Fig. 6 The same measurement at low noise, where every fit is a maximum and the estimates cluster on the truth. Nothing about the method changed; the ratio’s denominator simply stopped coming near zero.

There is a reading of those two figures that is worth resisting. They look like a noise problem, and noise problems have a standard remedy — collect more — which would be the wrong lesson. What changes between them is not how much data there is but the ratio of the curvature to its own standard error, and that ratio is moved by the design as much as by the sample size: by the width of the region, by how the runs are spread across levels, and by whether the experiment was run where the surface is actually curved. An experiment sitting on a flat stretch will produce this picture at any sample size a budget allows.

Two intervals for the same optimum, counted. 6,000 datasets at each curvature, with both constructions applied to each. Fieller's set covers 95.1% on average and never departs from its claim; the delta method covers 52.7% where the curvature is barely visible and 94.5% where it is sharp. It fails in the direction that matters — too narrow, so it understates how little the experiment has settled.
Fig. 7 The same coverage comparison at twice the noise. The delta method’s curve slides down and to the right; Fieller’s stays where it is and pays for it in unbounded answers instead.

The design can move the denominator

Everything above is downstream of one quantity: how well the design estimates b₂. That is a design decision, taken before any data exists, and it can be measured the same way everything else here is.

Fifteen runs spent three ways on the same true surface, at β₂ = −2:

the levels used runs Fieller unbounded delta-method coverage
three, five runs at each 15 61.3% 89.1%
five, three runs at each 15 68.5% 88.6%
seven, two runs at each 14 73.3% 87.4%

Fewer distinct levels, more replication at each: the curvature is estimated better, the denominator comes near zero less often, and the share of experiments that cannot bracket their own optimum falls by twelve points for no extra runs.

The reason is the allocation argument arriving in an unexpected place. A quadratic coefficient is estimated from the contrast between the ends and the middle, and spreading runs over seven levels spends most of them at settings that contribute little to that contrast. The design that estimates a quadratic best puts its runs where the quadratic differs most, which is three levels and nowhere else.

What the extra levels do buy is the ability to notice that the quadratic is wrong — a lack-of-fit test needs more distinct settings than the model has parameters, and a three-level design in one factor has exactly none spare. So the trade is real and it is not the trade it looks like: the extra levels are not buying precision about the optimum, they are buying the right to check the model that locates it.

The same shape, elsewhere on this site

A ratio whose denominator is an estimate that might be zero is not peculiar to response surfaces, and three other essays here are about the same object in different clothes.

The population spread estimated at exactly zero on a third of eight-group datasets, which is a denominator hitting the boundary of its range and being read as an instruction to pool completely.

The cointegrating slope read against a t table, where the statistic and the table were each correct and belonged to different distributions — the same failure as using a delta-method standard error for a quantity whose sampling distribution has no variance.

And the winner’s curse, which is what happens when a quantity is reported conditional on having cleared a threshold. The fitted optimum is not conditioned on anything, but its distribution is heavy-tailed for the same underlying reason: a small denominator is not rare, and the values it produces are extreme.

The common repair, in all four cases, is to stop treating the estimate as a number with an error bar and to invert a test that is exact for the quantity actually estimated. That is what Fieller’s construction does here, what the integrated interval does for a population spread, and what a simulated critical value does for a residual test with no table.

How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups.
Fig. 8 The same failure in the pooling field: a population spread estimated at exactly zero on a third of eight-group datasets. A denominator that reaches the boundary of its own range there; a denominator that passes through zero here.

What to report instead

Four things follow, and the first three cost nothing.

Report the interval for the location, not the location. A design with a well-determined curvature gives a tight one; a design without gives an unbounded one, and that is the design’s answer rather than a defect in the reporting.

Use Fieller’s set rather than a standard error. It is closed-form, it needs the same three variance components the delta method needs, and it covers what it says. The only thing it costs is the willingness to print an answer that is not an interval.

Check the shape before quoting the point. In two or more factors, the eigenvalues of B̂ are computed in the same breath as the stationary point and decide whether the stationary point is an optimum at all. A canonical analysis is one line of arithmetic and it is the difference between a maximum and a saddle.

And treat an optimum outside the region as a direction rather than a destination. If the fitted optimum is beyond the runs, the experiment has not located it; it has said which way to move, which is what a first-order design says more cheaply. The second-order design is worth its runs when the optimum is inside, and whether it was is knowable after the fact.

The check, and the refusal that makes it mean something

The coverage claim above is asserted at four curvatures against the simulation’s own standard error, so Fieller’s 95% has to hold where the delta method’s does not.

The refusal beside it is the one that matters, and it required admitting that the first version of this machinery was wrong. Fieller’s set was originally implemented as the whole line whenever the leading coefficient went negative — which is a superset of the correct answer, and covers more often than it claims: 97.4% against a nominal 95%, measured. It passed a coverage check, because the check only asked whether the truth was inside.

An interval construction that rounds its own answer outwards is not exact, and the way that was caught is worth recording: the coverage came out above nominal and stayed there at every setting. A construction that over-covers uniformly is not conservative by design unless it was designed to be; it is usually a bug that makes the answer bigger.

There is a second reason the bug survived as long as it did, and it is the reason this site prints the unbounded share beside every coverage number in this field. An unbounded set covers by definition, so a construction that returns unbounded sets often enough will pass any coverage check that is run alone. The pair — how often does it cover and how often does it decline to answer — is the only honest report, and either number without the other can be made to look good by a method that is useless.

That is the same discipline the multiple-comparison field needed for a different reason: a procedure that rejects nothing controls every error rate there is, so a familywise rate is meaningless without the power beside it. Here the pair is coverage and boundedness, and the shape of the mistake is identical.

The check now requires the exterior case to be an exterior — two roots, everything outside them — and requires the whole-line case to be rare. At the weakest curvature measured, 93.0% of the unbounded answers are exteriors and essentially none are the whole line. The distinction matters to a reader: the optimum is not between 0.4 and 3.1 is a finding, and nothing is known is not.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Central composite designConfidence intervalCoverageCurvatureDelta methodFieller intervalMonte CarloRatio estimatorSaddle pointStationary point