The surface between the corners

Walking up the gradient

The fitted gradient is wrong by an angle with a closed form, σ/(|β|√N), and what that angle costs is its squared cosine — twelve per cent at twenty degrees. What costs a third of the gain is not the direction at all. It is deciding where to stop.

Worth reading first: One factor at a time.

A first-order design cannot say where the best setting is. What it can say is which way is up: fit a plane to a factorial, take its gradient, and walk. Run the experiment again at the stopping point, fit another plane, walk again. The method is nearly a century old, it is what industrial experimenters actually do, and it survives for a reason worth measuring rather than asserting.

The obvious worry is that the fitted gradient is noisy, so the direction is wrong. It is wrong, by an amount with a closed form. The worry is misplaced.

Twenty walks up the same hill, σ = 2Each walk fits a plane to the same four-corner factorial, takes its gradient as a direction, and steps along it until a run comes in below the one before. The true optimum is the cross. 80% of the walks stop before the best point on their own path — not because the direction was wrong, but because one noisy run is enough to stop them, and the direction error costs only 3.9% of the available gain.the true optimum20 walksσ = 2, step 0.35, one run per stept* = 1.00 along the true gradient
Fig. 1 Twenty walks up the same hill from the same starting point, each fitting its own plane to its own factorial and stepping along its own gradient until a run comes in below the one before. The cross is the true optimum.

How wrong the direction is

On a factorial coded ±1 the columns are orthogonal, so the fitted coefficients are independent with variance σ²/N — no correlation to unpick, and no dependence on which coefficient. The fitted gradient is therefore the true one plus a spherical error of scale σ/√N, and the angle between them, for small errors, is the perpendicular component divided by the length of the true gradient:

θ ≈ σ / (|β|√N) radians.

Measured across four thousand walks on a surface with |β| = 5 and a four-run factorial, the root-mean-square angle comes out at 0.102 radians at σ = 1 against a predicted 0.100, and at 0.208 at σ = 2 against 0.200. At σ = 4 the measurement is 0.456 against a predicted 0.400 — the closed form is a small-angle approximation and at 26 degrees it is 14% low, which is the sort of gap worth reporting rather than absorbing into a tolerance.

Twelve degrees of error at σ = 2 sounds bad. It is not, and the reason is the second closed form.

What a wrong direction costs is its squared cosine

Move along a direction u from the centre of the design. The response along that line is

y(t) = y₀ + t(β·u) + t²(u′Bu)

which is a quadratic in the distance travelled, with its maximum at t* = −(β·u)/(2u′Bu) and a gain there of −(β·u)²/(4u′Bu).

On a surface curving equally in every direction — a circular hill — the denominator is the same for every u, so the ratio of what a wrong direction reaches to what the best one reaches is

cos²θ.

Twenty degrees off costs 12% of the gain. Ten degrees costs 3%. The loss is second order in the error, which is why a method that estimates a direction from four noisy runs works at all: the quantity being maximised is flat at its maximum, so being wrong about where the maximum is barely matters.

Measured, the two agree to four decimal places at every noise level: the share of the available gain the fitted direction offers is 99.0% at σ = 1, 95.9% at σ = 2 and 84.5% at σ = 4, against squared cosines of the measured angular errors that match them to within 0.005.

Two ways to lose the gain, and only one of them is the direction. The upper curve is what the fitted direction offers — cos²θ of the best available, measured against the closed form at every point. The lower one is what the walk reaches once it also has to decide where to stop. At σ = 6 the direction costs 27% and the stopping rule costs a further 23%, with 68% of walks stopping before the best point on their own path.
Fig. 2 The upper curve is what the fitted direction offers, with the closed form cos²θ drawn as points over it. The lower curve is what the walk actually reaches.

The gap between the two curves is the whole essay

The lower curve is what happens when the direction is followed rather than merely computed, and it sits far below the upper one: 83.3% of the available gain at σ = 1, 77.4% at σ = 2, 65.3% at σ = 4.

So at σ = 1 the direction error costs one point of the available gain and something else costs sixteen.

That something else is the stopping rule. The walk takes a run at each step and stops the first time a run comes in below the one before, which is what an experimenter does, and each run carries the same σ as the runs that fitted the plane. A step that gains less than the noise is a step that will look like a decline about half the time — so the walk stops not when the response starts falling but when it stops rising faster than the noise.

At σ = 1, 43% of walks stop before the best point on their own path. At σ = 2, 64%. At σ = 4, 73%. The mean stopping distance at σ = 4 is 0.67 against a t* of 0.90 — the walk gives up about a quarter of the way short.

Twenty walks up the same hill, σ = 6. Each walk fits a plane to the same four-corner factorial, takes its gradient as a direction, and steps along it until a run comes in below the one before. The true optimum is the cross. 90% of the walks stop before the best point on their own path — not because the direction was wrong, but because one noisy run is enough to stop them, and the direction error costs only 24.9% of the available gain.
Fig. 3 The same twenty walks at three times the noise. The directions are visibly scattered and the walks are visibly short, and the second of those costs far more than the first.

What a bigger factorial buys, and what it does not

If the direction is the part with a closed form, the obvious lever is the one the closed form names: the angle falls as 1/√N, so replicate the factorial and the direction improves.

It does, and it is worth knowing how little that is worth. At σ = 4, the noisiest setting measured here:

runs in the factorial angular error gain the direction offers gain the walk reaches
4 26.1° 84.5% 65.3%
8 17.2° 91.8% 71.9%
16 11.9° 95.9% 75.3%

Quadrupling the factorial — four runs to sixteen, before any stepping has happened — halves the angular error and buys eleven points of direction quality. It buys ten points of outcome. The share of walks that stop short of the best point on their own path is 73%, 75% and 74%: unmoved, because nothing about the plane affects the runs taken along the path.

That is the practical form of the whole argument. The part of steepest ascent with a clean theory behind it is the part that is already good enough and can be improved further at a cost nobody should pay. The part with no theory attached — one run per step, stop when it falls — is the part doing the damage, and it is the part every account of the method describes in a subordinate clause.

The step size is a design decision with an interior optimum

If stopping early is the problem, the obvious fix is to take smaller steps, which places more observations along the path and gives a finer picture of where it turns.

It makes things worse, and the measurement is unambiguous. At σ = 2, holding everything else fixed:

step share of the available gain reached walks stopping short
0.15 46.7% 100%
0.25 67.9% 88%
0.35 77.4% 64%
0.50 73.2% 25%
0.80 25.9% 20%

A small step is a large number of chances for one unlucky run to end the walk, and at 0.15 essentially every walk stops short of the best point it would have reached. A large step overshoots: it sails past the optimum, and the first run that comes in low is already well down the far side, so the walk stops at a point worse than where it started from two steps back.

Two ways to lose the gain, and only one of them is the direction. The upper curve is what the fitted direction offers — cos²θ of the best available, measured against the closed form at every point. The lower one is what the walk reaches once it also has to decide where to stop. At σ = 6 the direction costs 27% and the stopping rule costs a further 41%, with 92% of walks stopping before the best point on their own path.
Fig. 4 A step of 0.15, where the walk reaches under half the available gain at every noise level. The upper curve is unchanged — the direction is estimated from the same factorial — so the entire difference is stopping.
Two ways to lose the gain, and only one of them is the direction. The upper curve is what the fitted direction offers — cos²θ of the best available, measured against the closed form at every point. The lower one is what the walk reaches once it also has to decide where to stop. At σ = 6 the direction costs 27% and the stopping rule costs a further 78%, with 30% of walks stopping before the best point on their own path.
Fig. 5 And a step of 0.8, which fails the other way. The walk stops short less often and reaches less, because what it stops at is past the optimum rather than short of it.

The optimum in the middle is the part that makes this a design decision. A quantity that improves as a setting is made smaller can be tuned by an experimenter with no theory at all; a quantity with a maximum in the middle of its range has to be chosen, and choosing it requires knowing roughly how far away the optimum is and roughly how large σ is — the two things the experiment is being run to find out.

The noise that matters is measured against the slope

The closed form σ/(|β|√N) has the gradient’s length in the denominator, and that is the term which decides whether any of this works.

It is a scale-free statement: what matters is not how noisy the runs are but how noisy they are relative to how much the response changes across the design. Doubling the size of the design region doubles |β| in coded units and halves the angular error, at no cost in runs — which is the one lever in this field that is genuinely free, and the reason experienced experimenters make the first factorial as wide as the process will tolerate.

It is also the term that ends the first-order phase. Approaching the optimum, the surface flattens, |β| falls towards zero and the angular error grows without bound; at the optimum the gradient is zero and the fitted direction is uniformly distributed over the circle. The method does not degrade gracefully into a good answer near the top. It degrades into a random direction, and the only thing that says so is the curvature test — whose power, in the same units, is rising exactly as the gradient’s is falling.

Those two facts are the same fact seen twice: the quantity σ1/nf+1/nc\sigma\sqrt{1/n_f + 1/n_c} that sets the curvature test’s non-centrality does not depend on |β| at all, so as the walk climbs, one measurement gets worse and the other does not. The switch between the two designs is not a matter of judgement about when the surface has “become curved”. It is the point at which the ratio of two knowable quantities crosses one.

One loss moves with the noise and the other does not

Splitting the gap between the two curves at each noise level says which half is a property of the setting and which is a property of the method.

The direction costs 1.0, 4.1 and 15.5 points of the available gain at σ = 1, 2 and 4. The stopping costs 15.7, 18.5 and 19.2. One of those grows fifteenfold across the sweep and the other moves by a fifth.

So the stopping loss is very nearly a constant of the procedure — about eighteen points, whatever the noise — and the direction loss is the only term the noise moves. The two are equal somewhere around σ = 4.5, and below that the stopping rule is the larger cost by a factor that reaches sixteen at the quiet end.

The replication table says the same thing from the other side. At σ = 4, going from four runs to sixteen takes the direction cost from 15.5 points to 4.1 — a 74% reduction — and takes the stopping cost from 19.2 to 20.6, which is if anything slightly worse. Replication is a repair aimed at the smaller term and it does not touch the larger one.

That puts a ceiling on the whole lever. As the factorial grows the direction cost goes to zero and the walk’s gain approaches 100% minus the stopping loss: about 80% of what was available, at any noise level, however many runs are spent before stepping. An experimenter replicating the factorial is buying their way towards a bound they will not cross, and the last four points of the twenty are the only ones a better plane can reach.

Where the best step is, and what it is a fraction of

The step sweep has an interior optimum and the five readings locate it. Fitting a parabola through the three around the peak — 67.9% at 0.25, 77.4% at 0.35 and 73.2% at 0.50 — puts the maximum at a step of 0.40, reaching about 78.5%.

That is a point above the best step measured, and the curve is flat enough there that the difference is not worth chasing: 0.35 gives 77.4% and 0.40 gives 78.5%, against 46.7% at 0.15 and 25.9% at 0.80. The penalty is nearly symmetric in the logarithm — halving the step costs 31 points and doubling it costs 52 — so the safe side is the large one, which is the opposite of what the stopping-short diagnosis suggests.

Expressed as a fraction of what the walk is trying to cover it is a rule rather than a number. The optimal distance along the fitted direction is t* ≈ 0.90 on this surface, so the best step is about 0.44 of the distance to the optimum — a walk of two to three steps.

That is worth carrying because it says what has to be guessed. Not the step size in the units of the design, which depends on how the factorial was scaled, but the number of steps: aim to arrive in two or three. A walk planned to take ten steps is a walk with ten chances for one unlucky run to end it, which is what the 0.15 row costs; a walk planned to take one is the 0.80 row.

Why the method is still right

Three measurements, taken together, are an argument for steepest ascent rather than against it.

The direction is cheap and good enough. Four runs and a plane give a direction whose error costs a few per cent of the available gain. Nothing in this field improves on that per run, and the alternative — a second-order design at every stage — costs thirteen runs instead of four to answer a question the plane answers well enough to move on.

The losses are recoverable and the walk is repeated. Stopping a quarter short is not the end of the experiment; it is the starting point for the next factorial, which is fitted at the new centre and walks again. A method that captures 77% of the available gain per stage and is then repeated is a method that arrives, and the residual error is not cumulative, because each stage re-estimates the gradient where it stands.

And the thing it is bad at is the thing the next design is for. As the walk approaches the optimum the gradient flattens, |β| falls, and the angular error σ/(|β|√N) grows without bound — the direction becomes worthless exactly where the curvature becomes detectable. That is not a failure of the method: it is the signal to stop walking and build a second-order design, and the curvature test is what reads it.

The case where cos²θ is not the loss

The law above is a statement about a circular hill, and it was checked on one. On an elliptical surface — curving hard in one direction and gently in another — the denominator u′Bu depends on the direction too, and the loss is not the squared cosine of anything.

It goes the wrong way from the way one would guess. On a surface with B = [−2, −0.5; −0.5, −3], the fitted direction reaches 97.3% of the available gain where its own squared cosine is 95.9%: the error helps, on average.

The reason is that the steepest direction is not the direction the surface rises furthest along. Along a ridge, the gradient points across the ridge — up the steep face — while the largest available gain is along it, where the surface falls away slowly. An error that rotates the direction towards the ridge buys more distance than it loses in slope, and it does so often enough to raise the average.

That is worth stating in the check as well as in the prose, and it is: the assertion requires the elliptical surface to beat its own squared cosine, so the law cannot be applied where it does not hold without something failing. Which is the standing gotcha on this site arriving in a new field — an assertion conditioned on a parameter is usually a fact about the default — caught this time by writing the check for the case the essay was not about.

Four runs, and the term they cannot reach. Every run sits at a corner, so x₁² and x₂² are 1 at every run and both columns are copies of the intercept. The normal matrix is singular: the design has no information about curvature at all, and no analysis can recover it.
Fig. 6 The design the whole walk is built on: four runs, one plane, one gradient. Everything measured above is what those four runs can and cannot buy.

Why the figure has twenty paths in it

The hero of this essay draws twenty walks and not one, and the reason is the site’s rather than the subject’s. A single walk is a single draw from a distribution of walks: it has a direction, a stopping point and a gain, all three of which are random variables, and a reader shown one of them has been shown an anecdote with a picture attached.

What the twenty show that one cannot is the shape of the failure. The directions fan out symmetrically around the true gradient — that part looks like noise and is — while the stopping points are not symmetric at all. They pile up short of the optimum and thin out beyond it, because the stopping rule can only trigger early: a walk that has passed the optimum has a genuinely falling response and stops almost at once, and a walk that has not can be stopped at any point by one unlucky run.

An asymmetric distribution of endpoints is not something a mean angular error can express, and it is the reason the two curves in the cost figure are as far apart as they are.

Twenty walks up the same hill, σ = 1. Each walk fits a plane to the same four-corner factorial, takes its gradient as a direction, and steps along it until a run comes in below the one before. The true optimum is the cross. 40% of the walks stop before the best point on their own path — not because the direction was wrong, but because one noisy run is enough to stop them, and the direction error costs only 1.0% of the available gain.
Fig. 7 The same twenty walks at σ = 1, where the directions are within six degrees of the truth and 43% of the walks still stop short. The scatter of endpoints is what a single-path illustration of this method leaves out.

What a walk is actually reporting

An experimenter who follows this procedure ends up at a setting, and that setting is not an estimate of the optimum. It is the last point at which the response was still visibly rising, which is a different quantity, biased short by an amount the measurements above put at a quarter of the distance at moderate noise.

Nothing in the walk says so. The path is a sequence of runs, each correctly measured; the stopping decision was made honestly; the reported setting is better than the one it started from. There is no residual, no diagnostic and no interval — and the number that ought to accompany it, how far the true optimum might be from here, cannot be computed from a first-order design at all, because a plane has no optimum.

Computing it is what the second-order design is for, and what the location of an optimum turns out to be — a ratio of two estimates, one of which is a curvature that is often indistinguishable from zero — is the last essay in this field.

Which leaves one thing worth reporting that costs nothing: the path itself. Every run along it was measured, and the sequence of them is an estimate of the response along a line, with a quadratic to fit and a maximum to locate. That is a one-factor version of the problem the whole field is about, it uses data the experiment already has, and it turns “the walk stopped here” into a statement with an interval attached — one which, as the interval for a ratio turns out to behave, will often be unbounded and will be right to be.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CurvatureExperimental designFactorial designGradientMonte CarloStandard deviationSteepest ascent