The surface between the corners

The sign the curvature has

A fitted surface reports a maximum, a minimum or a saddle, and the report is a comparison of two estimated eigenvalues against zero. At a true second eigenvalue of −0.25 the fit calls a genuine maximum a saddle on 26.4% of studies, and at +0.25 it calls a genuine saddle a maximum on 25.1%.

Worth reading first: Walking up the gradient.

The optimum is a ratio is about where the stationary point is. Before that question can be asked there is a prior one, and it is usually answered in a sentence of the output: what kind of stationary point is it.

A fitted quadratic surface has a stationary point wherever its gradient vanishes, and that point is a maximum, a minimum or a saddle depending on the signs of two eigenvalues. The software prints the word. The word is the sign of two estimates.

What the fit calls the shape, against what it isOne eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 26.4% of studies; at +0.25 — a genuine saddle — it reports a maximum on 25.1%. The standard error of a squared coefficient under this design is 0.3791, and the region of confusion is about that wide either side of zero.00.2500.5000.7501-2-1012the second eigenvalue of the true surfaceshare of studies reporting each shapewhere the truth changes shapereported as a maximumreported as a saddle2,500 studies a point, σ = 1, 5 centre runsone standard error of a squared coefficient is 0.379
Fig. 1 One eigenvalue of the true surface held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left, a saddle on the right, and changes at exactly zero. At a true second eigenvalue of −0.25 the fit reports a saddle on 26.4% of studies; at +0.25 it reports a maximum on 25.1%.

What canonical analysis does

The fitted surface is y^=b^0+g^x+xB^x\hat y = \hat b_0 + \hat{\mathbf g}'\mathbf x + \mathbf x'\hat{\mathbf B}\mathbf x with B^\hat{\mathbf B} symmetric, and the classification is one eigendecomposition.

Rotating to the eigenvectors of B^\hat{\mathbf B} turns the surface into a sum of independent squared terms, y^=y^+λ1w12+λ2w22\hat y = \hat y^{*} + \lambda_1 w_1^{2} + \lambda_2 w_2^{2}, with λ1\lambda_1 and λ2\lambda_2 the eigenvalues and ww the coordinates in the rotated frame. Then:

  • both λ\lambda negative — the surface falls away in every direction, and the stationary point is a maximum;
  • both positive — a minimum;
  • opposite signs — a saddle, rising in one direction and falling in the other.

That is the whole of canonical analysis and it is exact given B^\hat{\mathbf B}. The difficulty is that B^\hat{\mathbf B} is three estimated coefficients, and the classification is a pair of inequalities applied to functions of them.

A central composite design, 13 runs. Adding 4 axial runs at ±√2 gives every factor three levels, which is the least that can estimate a squared term. The normal matrix now inverts, so each βᵢᵢ has an estimate of its own — and at exactly this axial distance the design is rotatable, which the next figure measures.
Fig. 2 The design every figure here is fitted with: the thirteen-run central composite of three levels and a ring, four corners at ±1\pm 1, four axial runs at ±2\pm\sqrt{2} and five at the centre. Its arrangement is what fixes the standard error of a squared coefficient, and therefore what fixes how small a curvature can be told from zero.

Why zero is the hard place

A hypothesis test compares an estimate against a threshold and reports uncertainty. This comparison reports a word, and there is no interval on a word.

The eigenvalues are continuous functions of the three quadratic coefficients, each of which has a standard error the design determines. Under the thirteen-run composite design at σ=1\sigma = 1, the standard error of a squared coefficient is 0.3791 — a number available before any data, from the design alone.

So a true eigenvalue of 0.25-0.25 is two-thirds of a standard error below zero. An estimate of it lands above zero about a quarter of the time, and when it does, the word printed changes from “maximum” to “saddle”.

That is what the figure’s shaded band is: one standard error either side of zero, and it covers the whole of the region where the reported shape is unreliable. The width of the confusion is a property of the design, and the design knows it in advance.

The comparison with a test is worth making explicit, because a reader who is used to hypothesis tests will assume one is happening here and none is.

A test of “is this eigenvalue zero” would have a level, a power curve and a stated error rate, and the 27% would be its size or its type II error depending on which side of zero the truth sat. What is happening instead is a point estimate compared with zero and reported categorically — the statistical equivalent of rounding. Its error rate is a half at the boundary by construction, which is worse than any test would accept, and nothing about the output flags that a comparison was made.

So the difficulty is not that the classification is an under-powered test. It is that it is not a test, and nothing that would tell a reader how much to trust it was ever attached.

An exact ridge is a coin

The middle of the sweep is worth isolating because it is the case the whole subject is about.

At λ2=0\lambda_2 = 0 the true surface is a ridge: it falls away in one direction and is exactly flat in the other, so every point along a line through the stationary point has the same response. There is no unique optimum — there is a line of them — and the honest report is that the surface is flat in that direction.

The fit reports a maximum on 48.7% of studies and a saddle on 51.3%. It never reports a ridge, because “flat” is not one of the three words, and a coefficient estimated as exactly zero has probability zero.

A ridge is the one case where the right answer is not on the menu, and it is the case an experimenter most needs to know about: a flat direction means the setting can be moved along it for free, which is usually the most actionable thing an experiment can find.

The other reading of the same number is a sample-size statement. The standard error of a squared coefficient scales as σ\sigma over the square root of the number of runs, so halving the confusion width means quadrupling the experiment. Telling a curvature of 0.1-0.1 from zero at this noise would take about sixteen times the thirteen runs — two hundred and eight of them — which is not a decision anybody makes to settle a classification.

So the practical answer is not a bigger design. It is to stop asking the question in a form whose answer costs that much, which is what the last section of this essay is about.

Both directions of error, and they are not the same mistake

The confusion runs both ways and the consequences do not match.

A true maximum reported as a saddle. The experimenter concludes there is no optimum in the region, that the surface rises in some direction, and that further runs should be made outside it. They are walking away from an optimum they had.

A true saddle reported as a maximum. The experimenter concludes the stationary point is the best setting, stops, and recommends it. In fact the surface rises away from it in one direction, and the recommended setting is the worst point along that direction.

The second is the more expensive because it stops the search. The first is recoverable — further runs are made and eventually find the maximum — and the second produces a confident recommendation that is a minimum along one axis.

What the fit calls the shape, against what it is. One eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 47.2% of studies; at +0.25 — a genuine saddle — it reports a maximum on 36.8%. The standard error of a squared coefficient under this design is 1.1374, and the region of confusion is about that wide either side of zero.
Fig. 3 The same sweep at three times the noise. The standard error of a squared coefficient is now 1.1374 and the confusion covers the whole sweep: at a true eigenvalue of −1 the shape is read correctly on 74.8% of studies and at −0.25 on 52.6%, which is barely better than the coin at the ridge.
What the fit calls the shape, against what it is. One eigenvalue held at −3 and the other swept from −2 to 2, so the truth is a maximum on the left and a saddle on the right and the change happens at exactly zero. At an eigenvalue of −0.25 — a genuine maximum — the fit reports a saddle on 10.4% of studies; at +0.25 — a genuine saddle — it reports a maximum on 9.8%. The standard error of a squared coefficient under this design is 0.1896, and the region of confusion is about that wide either side of zero.
Fig. 4 And at half the noise, where the standard error is 0.1896. Every truth but the two nearest zero is read correctly on essentially every study — 99.6% at −0.5, 100% at −1 — and the ridge is still a coin at 49.7%. The confusion narrows with the noise and never closes at the ridge, because no amount of precision makes an estimated eigenvalue exactly zero.

The width scales, and the ridge does not

The three noise levels give the scaling law, and it is the one the standard error predicts.

σ\sigma standard error of a squared coefficient shape read correctly at λ2=0.5\lambda_2 = -0.5 at 0.25-0.25 at 00
0.5 0.1896 99.6% 89.6% 50.3%
1 0.3791 89.2% 73.6% 51.3%
3 1.1374 61.2% 52.6% 55.6%

The first two columns move together: doubling the noise doubles the standard error and roughly doubles the eigenvalue that can be told from zero. That is the ordinary 1/n1/\sqrt{n} arithmetic arriving as a σ\sigma arithmetic, since the design is fixed.

The last column does not move at all. At a true ridge the answer is a coin at every noise level, and it will be a coin at any sample size, because the fit is being asked which side of zero an estimate of zero fell on. That is not a precision problem and no design fixes it; it is the classification having no word for the truth.

What the design could do about it

Since the confusion width is a property of the design, the natural question is which design feature controls it, and the answer is not the one the usual design advice would suggest.

The standard error of a squared coefficient comes from the diagonal of (XX)1(\mathbf X'\mathbf X)^{-1}, and for a composite design that entry is decided by how far the axial runs reach and how many runs sit at the centre. Pushing the axial runs out increases the spread of x2x^{2} across the design, which is what a squared coefficient is estimated from, and so shrinks its standard error.

That is a different objective from the ones the essay that chose a ring’s radius picked α\alpha by. Rotatability puts the axial runs at F1/4F^{1/4}; uniform precision decides the centre runs by the prediction variance at the middle. Neither is aimed at the precision of an individual squared coefficient, and an experimenter who cares about the shape rather than about prediction wants the third objective.

The honest version of that recommendation is narrow. The gains available from moving α are tens of per cent in a standard error, so they shift the confusion width by tens of per cent — real, and not the order of magnitude that would make a ridge readable. No arrangement of thirteen runs distinguishes a curvature of a tenth from zero, and that is a statement about the number of runs rather than about where they go.

The stationary point moves with the shape

The classification and the location are not independent, and the joint behaviour is worse than either alone.

400 fits of the same surface, σ = 3. Each point is the stationary point of one fitted quadratic, from one central composite design run on the same true surface — which has a maximum at (0.91, 0.35), marked. 8% of the fits are saddles rather than maxima, so their stationary point is not an optimum of anything, and 39% land outside the region the design explored. The median distance from the centre is 1.18 against a true 0.98.
Fig. 5 Where the fitted stationary point lands, over four hundred studies at three times the noise. The cloud is not a cloud around the truth: it has a long tail, because the stationary point is a ratio whose denominator is the curvature and the curvature is close to zero on the studies where the shape is misread.

The stationary point is 12B^1g^-\tfrac12\hat{\mathbf B}^{-1}\hat{\mathbf g}, and B^1\hat{\mathbf B}^{-1} blows up as B^\hat{\mathbf B} approaches singularity — which is exactly where an eigenvalue is near zero, which is exactly where the shape is misread.

So the studies that get the shape wrong are the studies whose stationary point is furthest away. The two failures are the same failure, and the ratio essay was measuring one face of it: the denominator whose interval sometimes has to be the whole line is the same eigenvalue whose sign decides the word.

What a reader of a report cannot recover

The practical shape of this is about what gets written down, and it is worth being concrete because the loss is at the reporting step rather than at the fitting one.

A response-surface analysis is usually reported as: the fitted coefficients, the stationary point, and the shape. From the coefficients a reader can recompute the eigenvalues and see that one is 0.30-0.30 — and from the coefficients’ standard errors, if they are given, they can see that 0.380.38 is the scale. Almost nothing about this is hidden if the coefficients and their standard errors are both printed.

What defeats that is the summary. A report that says “the analysis identified a saddle point at (1.2, −0.4)” has thrown away the two numbers that say how nearly it was a maximum, and the site’s standing complaint about a summary applies exactly: the sentence is correct, the reader cannot tell which of two quite different situations produced it, and the missing information was computed and discarded.

What to report instead

The measurements support a short list, and the first item is the one that costs nothing.

Report the eigenvalues with their standard errors, not the word. Two numbers with intervals say everything the word says and also say how close to zero they are, and the standard errors come from the same fit. An eigenvalue of 0.30±0.38-0.30 \pm 0.38 and an eigenvalue of 3.10±0.38-3.10 \pm 0.38 is a complete description of a surface that is steep in one direction and flat in the other, and “maximum” is not.

Treat a near-zero eigenvalue as a ridge. The three-word classification has no entry for a flat direction, and a flat direction is the finding most worth having — it means the corresponding setting can be chosen on cost, or convenience, or anything else, at no loss of response.

Say which direction is flat. When one eigenvalue is near zero, the corresponding eigenvector is the direction along which the response barely changes, and it is a vector the fit already computed. Naming it — the response is insensitive to increasing temperature and pressure together in a two-to-one ratio — is a finding an engineer can use, and it survives the shape being misclassified because it does not depend on the sign.

And compute the confusion width before the experiment. The standard error of a squared coefficient is a property of the design, so an experimenter can say in advance: this design will not be able to tell a curvature smaller than about 0.4 from zero. If the curvatures of interest are smaller than that, the design is the wrong size.

Power against centre runs, at Σβᵢᵢ = -2. Every design here has the same four corners and differs only in how many runs sit at the centre. The line is the non-central t on one fewer degrees of freedom than there are centre runs, at non-centrality Σβᵢᵢ divided by σ√(1/factorial runs + 1/centre runs); the points are 6,000 simulated experiments each. Power goes from 32% at 3 centre runs to 92% at 16, and none of that came from the factorial.
Fig. 6 The design-side version of the same statement, measured earlier: the power of the centre-point test to detect curvature at all, against how much curvature there is. The shape question is the same question one level of detail further in — not is there curvature but what sign does each of its two components have.

The same shape, three fields over

The classification problem here has a form that recurs under other names, and naming it is worth a paragraph because the repair is the same in all of them.

A continuous quantity is estimated, thresholded, and the threshold’s output is reported instead of the quantity. An eigenvalue’s sign becomes a word. A p-value becomes significant or not. A variance estimate becomes zero or not, and a whole level of a model disappears with it. A count becomes a decision.

In every case the thresholding is defensible on its own terms and the information lost is the distance from the threshold, which is the only thing that says how much to trust the output. And in every case the repair is identical: report the quantity and its standard error beside the word, which costs a line and is almost never done.

What is particular here is that the word is not a decision anybody consciously took. Nobody chose to threshold; the classification into three shapes is how canonical analysis is described in every textbook, and the thresholding is inside the description rather than applied to its output.

What is claimed here, and what is not

The claim is how reliably a second-order fit reports the shape of the surface it fitted: that at a true second eigenvalue of −0.25 it calls a genuine maximum a saddle on 26.4% of studies and at +0.25 calls a genuine saddle a maximum on 25.1%; that at an exact ridge it splits 48.7% to 51.3% and never reports a ridge; and that the width of the confusion is about one standard error of a squared coefficient, which is 0.3791 under this design and is known before any data exists.

Every rate is two and a half thousand studies at each of nine truths, all fitted with the same thirteen-run composite design. The truths are built by rotating a diagonal of stated eigenvalues, so the surface is not axis-aligned and the interaction coefficient is doing work.

What stays out: the sampling distribution of the eigenvalues themselves, which is not normal near a repeated root and which is the reason the standard errors quoted for them by software should be treated carefully; formal tests of the shape, which exist and require the same near-zero case to be handled; three and more factors, where there are kk eigenvalues and 2k2^{k} sign patterns rather than three shapes; and the rising ridge, where one eigenvalue is near zero and the gradient has a component along it, which is a fourth case the three words do not cover either.

Still open: a stationary point that cannot be run

Half of the difficulty here is that a near-zero eigenvalue makes the stationary point’s location unstable, and the figure above shows the cloud without saying what to do about it.

The question an experimenter faces is concrete. The fit has returned a stationary point at a radius of 2.29 in coded units, the region runs to 1.41, and the recommended setting cannot be run. It happens on 24.9% of studies at twice the noise and on 11.8% the point lands more than three units out. The answer is not to report it, and it is not to ignore it either: there is a path of best settings at each radius, it has a closed form, and where it ends is what should be recommended. That is what to do when the best setting is outside the region.

The check, and the refusal

Three claims are gated. That the shape is read correctly at both ends of the sweep, where the truth is unambiguous — the control, without which the whole figure could be an artefact of the fit. That it is read less well at the ridge in the middle than at either end, which is the figure’s shape stated as an ordering rather than as a value. And that at an exact ridge the fit calls it a maximum about half the time, required within twelve points of a half, because a coin is the specific claim.

The refusals are the two directions of error and both are required to occur: there must be a true maximum the fit calls a saddle on more than one study in twenty, and a true saddle it calls a maximum at the same rate. A sweep in which only one direction of confusion appeared would suggest a bias in the fit rather than a resolution limit in the design, and the pair is what says it is the second.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Canonical analysisCentral composite designCurvatureEigenvalueExperimental designLeast squaresResponse-surfaceSaddle pointStationary point