The surface between the corners

The design that cannot see a curve

A two-level factorial has every run at a corner, where every squared term equals one — so the column that would estimate curvature is a copy of the intercept, and the design has no information about it at all. A few runs at the centre buy one number back, and only one.

Worth reading first: One factor at a time · The variance removed before the data.

A factorial design answers which factors matter. It is very good at it: every run contributes to every effect, the estimates are orthogonal, and the arrangement is decided before a unit is measured. What it cannot do — not badly, not imprecisely, but at all — is say what setting of those factors is best.

The reason is arithmetic rather than statistics, and it is worth stating before anything is measured.

Four runs, and the term they cannot reach. Every run sits at a corner, so x₁² and x₂² are 1 at every run and both columns are copies of the intercept. The normal matrix is singular: the design has no information about curvature at all, and no analysis can recover it.
Fig. 1 Four runs at four corners. Every one of them has x₁² = 1 and x₂² = 1, so the columns that would estimate the squared terms are copies of the intercept column, and the normal matrix is singular.

The column that is the intercept

Fit a second-order model to a two-factor experiment and the design matrix has six columns: an intercept, two main effects, an interaction, and two squared terms. On a two-level factorial each factor takes the values −1 and +1 and nothing else, so x₁² is 1 at every run. So is x₂². So is the intercept.

Three of the six columns are therefore identical, and X′X — the matrix whose inverse the least-squares estimates are built from — is singular. The design does not estimate curvature badly. It contains no information about it whatever, and no analysis, no software and no amount of replication can recover what was never collected.

That is a different kind of failure from most on this site. A confidence interval that covers 87.6% is a procedure making a claim it does not meet; a p-value read without its sample size is a number being asked a question it cannot answer. Here the quantity has no estimate at all, and the honest report is a matrix that does not invert.

It matters because the surfaces experimenters care about are curved. A yield that rises with temperature and then falls; a dose that helps and then harms; a setting with a best value rather than a direction. Fit a plane to any of them and the plane’s message is always the same: keep going. A plane has no interior maximum, so a first-order design can recommend a direction and can never recommend a destination.

What a handful of runs at the centre buys

The repair is small and it is not free. Add runs at the centre of the design — all factors at zero, replicated a few times — and two things arrive at once.

The first is an estimate of σ that assumes nothing at all. Replicates at one setting differ only by noise, whatever the surface between the corners is doing, so the spread of the centre runs is a pure-error estimate: it cannot be inflated by the model being wrong, because no model was used to compute it.

The second is one contrast that sees curvature. Every corner has xᵢ² = 1 and every centre run has xᵢ² = 0, so the difference between the corner mean and the centre mean carries the squared terms and nothing else — the main effects cancel across the corners, and so does the interaction.

Four corners and 5 runs at the centre. The centre runs add two things at once. They estimate σ from replicates at one setting, which assumes nothing about the surface, and they supply the one contrast that sees curvature — the corner mean minus the centre mean, which estimates Σβᵢᵢ. It is one number: the design still cannot say which factor the curvature is in. 5 centre runs give 4 degrees of freedom for the pure-error estimate, and that is what sets the test's power.
Fig. 2 The same four corners with five runs at the centre. The design now estimates σ from replicates and carries one contrast that sees curvature — and the normal matrix is still singular, because one contrast is not two coefficients.

Its expectation is exactly Σβᵢᵢ, the sum of the pure quadratic coefficients. Measured across four thousand experiments on a surface with β₁₁ = β₂₂ = −2.5, the contrast comes out at −5.00 against a true −5. On a surface with main effects of 6 and 5, an interaction of 3, and no pure quadratic terms at all, the same contrast comes out at 0.0056: large effects and a large interaction move it by nothing, which is what makes it a measurement of curvature rather than a symptom of something else.

One number, and the design still does not invert

Notice what the centre runs did not buy. The contrast estimates the sum of the quadratic coefficients, and a sum is one number where the model has two. Adding centre points to a two-level factorial leaves X′X singular: the design can now say whether the surface is curved and cannot say which factor the curvature is in.

That is not a technicality, and the case where it bites is not exotic.

Take a surface curving upward in the first factor exactly as hard as it curves downward in the second — β₁₁ = +3 and β₂₂ = −3. That is a saddle: a surface with no maximum anywhere in the region, rising without limit in one direction and falling in the other, and about as far from a plane as a second-order surface gets.

Σβᵢᵢ is zero. Run the curvature test on it across eight thousand experiments and it rejects 4.89% of the time, which is its nominal rate. The same design and the same seeds, given a surface curving downward in both factors by the same amount, reject 100.0% of the time.

So the test is not weak on a saddle. It is correct, and it is answering a narrower question than the one being asked. This site more often finds machinery that returns the boundary of its own range and means “not there” — an estimated population spread of exactly zero, a quantile that returns its own bracket. This is the mirror image: exactly the right answer to a question one word narrower than the question that was asked.

Where the power comes from

If the curvature test is the gate that decides whether to spend runs on a second-order design, its power is worth knowing, and it has a closed form. The contrast has standard error σ1/nf+1/nc\sigma\sqrt{1/n_f + 1/n_c}, the pure-error estimate has nc1n_c - 1 degrees of freedom, and the statistic is a non-central t with non-centrality Σβii/(σ1/nf+1/nc)\Sigma\beta_{ii}/(\sigma\sqrt{1/n_f + 1/n_c}).

Both routes exist, so this site computes both.

Power against centre runs, at Σβᵢᵢ = -2Every design here has the same four corners and differs only in how many runs sit at the centre. The line is the non-central t on one fewer degrees of freedom than there are centre runs, at non-centrality Σβᵢᵢ divided by σ√(1/factorial runs + 1/centre runs); the points are 6,000 simulated experiments each. Power goes from 32% at 3 centre runs to 92% at 16, and none of that came from the factorial.00.2500.5000.750151015runs at the centre of the designpower of the curvature testthe non-central t6,000 experiments eachfour corners throughout, Σβᵢᵢ = -2, σ = 161% at five centre runs
Fig. 3 Every design on this curve has the same four corners and differs only in how many runs sit at the centre. The line is the non-central t; the points are six thousand simulated experiments each.

At a curvature of two standard deviations the test has 32.0% power with three centre runs, 61.4% with five and 82.9% with nine. The counted values are 32.2%, 61.4% and 83.0%, which is agreement to within the simulation’s own error at every point.

The shape of that result is the part worth carrying, and it needs stating carefully. Every design on the curve has the same four factorial points, so everything the curve shows was bought with centre runs — which invites the summary the power comes from the centre and not from the corners. That summary is false as stated, and the next section measures how, but the half of it that survives is worth having first: the contrast’s precision is limited by the number of centre runs on one side and by the degrees of freedom in the pure-error estimate on the other, and a bigger factorial with no replication improves neither. Sixteen corners run once, with five at the centre, has the same four pure-error degrees of freedom as four corners run once.

The degrees of freedom are the sharper of the two constraints and the easier to overlook. Three centre runs give two degrees of freedom and a critical value of 4.303. Five give four and 2.776. Nine give eight and 2.306. Going from three centre runs to five moves the critical value by more than a third, and no amount of unreplicated factorial buys any of it.

That is worth dwelling on, because it inverts the intuition a screening design trains. There, precision comes from runs, and more factors at the same two levels cost runs and buy information. Here what the test needs is repeated measurements at some setting — any setting — so that σ can be estimated without believing the model, and distinct settings do not supply that however many of them there are.

A 2⁴ factorial with three runs at the centre is a nineteen-run experiment carrying a two-degree-of-freedom estimate of σ, and its curvature test reaches 42.0% against a curvature of two standard deviations. Two more centre runs — twenty-one runs rather than nineteen — take it to 82.7%. Two runs, and the power doubles, because they are the two runs that move the critical value from 4.303 to 2.776.

Power against curvature, 3 centre runs. The two routes to the same number: the non-central t on 2 degrees of freedom, and 6,000 simulated experiments at each setting. A curvature of one standard deviation is found 13% of the time by this design, and one of three 55%. The test that decides whether to spend runs on a second-order design is itself a low-powered test.
Fig. 4 Power against the size of the curvature with three centre runs. The critical value is 4.303 and the test needs a curvature of nearly three standard deviations before it is more likely than not to fire.
Power against curvature, 9 centre runs. The two routes to the same number: the non-central t on 8 degrees of freedom, and 6,000 simulated experiments at each setting. A curvature of one standard deviation is found 31% of the time by this design, and one of three 99%. The test that decides whether to spend runs on a second-order design is itself a low-powered test.
Fig. 5 The same sweep with nine centre runs. The corners are unchanged and the whole curve has moved left.

Which of the two constraints the two runs actually relaxed

Two extra centre runs take the four-corner design from 32.0% to 61.4%, and the section above names two mechanisms without saying how the twenty-nine points divide between them. The closed form can be asked, by moving one at a time.

The contrast’s standard error is σ1/nf+1/nc\sigma\sqrt{1/n_f + 1/n_c}, which is 0.764σ at three centre runs and 0.671σ at five — a non-centrality of 2.62 rising to 2.98. The pure-error estimate goes from two degrees of freedom to four, and the critical value from 4.303 to 2.776.

  • The precision alone. Keep two degrees of freedom and the 4.303, and give the test the better non-centrality: 38.4%.
  • The critical value alone. Keep the worse non-centrality of 2.62, and give the test four degrees of freedom and the 2.776: 51.2%.

So of the 29.4 points gained, about 6 are the contrast being measured more precisely and about 19 are the critical value falling, with the remaining 4 in the interaction between them. Three quarters of what two centre runs buy is bought by the degrees of freedom, not by the precision.

The same split on the sixteen-corner design is more lopsided still. There the contrast is already precise — 0.629σ at three centre runs, against 0.764σ with four corners — so the precision half has less left to give: holding the critical value at 4.303 and improving only the non-centrality takes 41.9% to 54.8%, and the remaining twenty-eight points to 82.7% are the critical value alone.

That is the arithmetic behind the essay’s claim that a bigger factorial buys neither constraint. It buys a little of the first — sixteen corners do measure the corner mean better than four — and none at all of the second, and the second is where three quarters of the power is.

Where the half-power point is

The same closed form locates the curvature at which the test becomes more likely than not to fire, which is the honest way to describe a design’s reach.

With four corners and three centre runs it is 2.77 standard deviations: a surface whose two pure quadratic coefficients sum to less than that is more often missed than caught, by a design that has spent seven runs. With five centre runs the half-power curvature falls to 1.73, and with nine to 1.34.

Those three numbers are the design decision stated in the units of the thing being looked for. A practitioner who believes the surface’s total curvature is comparable to the noise needs nine centre runs to have an even chance of seeing it, and one who believes it is three times the noise can afford three.

The corners buy it too, if they are replicated

The sentence above is the one this essay was going to be built on, and measuring it properly makes it narrower.

What the test is short of is pure-error degrees of freedom, and replicates supply them wherever they sit. Run the four corners twice instead of adding four more centre runs — thirteen runs either way, and eight degrees of freedom either way — and the curvature test has 86.6% power against 82.9%, at a curvature of two standard deviations.

The replicated factorial wins on the standard error as well, and the reason is one this site has met in every other difference of two means: the contrast is ȳ_corners − ȳ_centre, so it wants its two groups the same size. Thirteen runs split eight and five give 1/nf+1/nc1/n_f + 1/n_c = 0.325. Split four and nine they give 0.361. Four corners and nine centre runs is a lopsided design for the one contrast the centre runs were added to supply.

And the replicated factorial estimates every main effect twice as precisely, which the extra centre runs do not touch at all.

So the honest version of the claim is: the curvature test’s power is bought with replication, and replication anywhere buys it. The centre runs’ real advantage is narrower and still worth having — they are the only runs that add degrees of freedom without changing the factorial, which matters when the factorial is a fraction chosen for its aliasing and cannot be doubled without doubling the experiment. The version that survives contact with a design catalogue is the narrow one.

Power against centre runs, at Σβᵢᵢ = -1. Every design here has the same four corners and differs only in how many runs sit at the centre. The line is the non-central t on one fewer degrees of freedom than there are centre runs, at non-centrality Σβᵢᵢ divided by σ√(1/factorial runs + 1/centre runs); the points are 6,000 simulated experiments each. Power goes from 13% at 3 centre runs to 39% at 16, and none of that came from the factorial.
Fig. 6 The same sweep at half the curvature. Everything is harder: nine centre runs on an unreplicated factorial reach 31.1% against a curvature of one standard deviation, where thirteen runs spent as two replicates of the factorial and five at the centre reach 34.0%.

A low-powered gate on an expensive decision

Put the two together and the picture is uncomfortable. The decision the curvature test governs is whether to run a second-order design, which typically doubles the experiment. The test itself, at the design most people actually run — four corners, five centre points — has 21.1% power against a curvature of one standard deviation and 61.4% against two.

A test with 21% power that does not fire is not evidence that the surface is flat. It is the ordinary failure to reject being read as a finding, in a place where the consequence is a design decision rather than a published claim: the experiment proceeds on a plane, recommends a direction, and the direction is followed until the response stops improving — which is the next essay’s subject and where the curvature that was not detected turns up again.

There is a cheaper reading of the same numbers, and it is the useful one. The curvature test is not really a hypothesis test about the world; it is a check on whether the first-order phase is still working. A first-order design is a tool for moving, and it stops being the right tool when the surface stops being locally planar. Read that way, 21% power against a small curvature is fine — a curvature small enough to be missed is a curvature small enough that the plane is still a good enough local description to walk on.

What it is not is a licence to report that the response is linear.

What it costs to fix, and why the fix is a third level

The design does not invert because two levels cannot estimate a squared term. The remedy is a third level, and where to put it is the question the next design answers.

A central composite design, 13 runs. Adding 4 axial runs at ±√2 gives every factor three levels, which is the least that can estimate a squared term. The normal matrix now inverts, so each βᵢᵢ has an estimate of its own — and at exactly this axial distance the design is rotatable, which the next figure measures.
Fig. 7 The central composite design: the four corners, four axial runs at ±√2, and five at the centre. Every factor now takes three values, the normal matrix inverts, and each squared coefficient has an estimate of its own.

Adding four axial runs takes the two-factor design from nine runs to thirteen and buys three things. The normal matrix inverts. Each βᵢᵢ has an estimate of its own, so the saddle above is no longer invisible. And the prediction variance acquires a property — the same in every direction, at the axial distance √2 and no other — which the next essay measures and which is exact rather than approximate.

The cost is worth stating plainly, because the sequence in which these designs are run is the whole of the practical advice. A second-order design in k factors needs 2k+2k+nc2^k + 2k + n_c runs: thirteen at two factors, twenty at three, thirty-one at four. A screening factorial in the same four factors, at half fraction, needs eight. The order — screen with a fraction, move with a first-order design, and build a second-order design only where the response has stopped rising — is not tradition. It is what the run counts make sensible, and the curvature test is the switch between the second step and the third.

The check this essay is built on

Three claims here have refusals behind them rather than tolerances.

The singularity is required rather than observed. The check calls the same inverse on the corners, on the corners plus centre runs, and on the full central composite design, and demands that the first two return nothing and the third return a matrix. A later change that quietly regularises the solve and returns a number for a design that cannot estimate one would be caught by the first two.

The contrast has to be blind. It is shown a surface with main effects of 6 and 5 and an interaction of 3, and required to report nothing.

And the saddle has to be missed. A test that fired on it would mean the contrast was picking up something other than Σβᵢᵢ, which would be a bug in either the design or the arithmetic. It rejects at 4.89% against a nominal 5%, on a surface that curves as hard as any in this field.

A fourth was written because the essay’s first draft was wrong. The claim that power comes from centre runs rather than corners is a fact about a fixed unreplicated factorial and reads as a fact about designs, and the check now compares the two thirteen-run designs directly and requires the replicated factorial to win. An essay-shaped generalisation that holds at one point of a design space is the gotcha this site has recorded in three previous phases; this is the first time it was caught in the prose rather than in an assertion.

The saddle check is the one that would be least natural to write, because it asserts that the machinery fails. It is also the only one of the three that says what the design is for: an experiment that reports “no curvature detected” has not said the surface is flat, and the one number it did estimate cannot tell the difference between a plane and a saddle.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Centre pointCurvatureDegrees of freedomExperimental designFactorial designThe non-central tPure-errorSaddle pointStatistical power