Three runs at the end of the line
Worth reading first: What normal actually looks like · The line that one point drew.
Residuals are not the errors found that a residual carries only part of its error’s variance, and the part it loses is largest at the points with the most leverage: at a leverage of 0.663 the residual keeps 34% of the error’s variance and the fit absorbs the rest. It ended on the one thing standardising cannot repair — the information the point never gave — and on a claim about design: a design with replication at its extreme points can tell a bad observation from a wrong model, and a design without it cannot, whatever residual is plotted.
This essay makes that claim exact. It is stronger than it sounds. Without replication the two explanations are not merely hard to separate; they produce data with the same distribution, so there is nothing in the data to separate them by.
Two explanations, one set of data
Take twelve runs, one at each of , and a straight line with normal errors of known spread . Something goes wrong at the top of the range, where the response comes out four above the line. There are two ways that can happen.
One observation is bad. The instrument slipped, a sample was mislabelled, the run was recorded wrong. The line is fine; one error is not.
The model is wrong at that end. The response saturates, or a second mechanism switches on at high doses, and the true mean at is off the line. Every run made there would be off it.
With one run at , the two explanations change the data in exactly the same way: they add four to the one number recorded there. The upper panels of the hero figure are not similar; they are identical, point for point, because the same errors were drawn and the same four was added to the same observation. A bad observation of size at a single point and a bend of size confined to that point are two stories about one vector of numbers. Their likelihoods are equal for every dataset, so every statistic has the same distribution under both, and no rule — residual plot, studentised residual, Cook’s distance, lack-of-fit test, a statistician’s eye — can call one more often under the first than under the second.
That is a statement about identifiability, not about power. More data at the other eleven points does not help, and a more careful diagnostic does not help. The design asked the top of the range one question, and one answer cannot say whether the question was misheard or the answer is surprising.
Three runs at the top
Move six of the twelve runs: three at , three at , and six spread over to . The same two explanations now leave different traces in the lower panels.
A bad observation moves one of the three runs at the top. The other two sit where the line predicts, and the bad one stands apart from its own replicates. A bend moves all three together, so they agree with one another and disagree with the line.
The two traces have two statistics, and under normal errors the statistics are independent of each other:
- the spread of the replicates about their own mean, , which is if the three share a mean — whatever that mean is — and noncentral with if one of them is off by ;
- the top mean against the line fitted to the other nine runs, which is shifted by under a bend and by only under a bad observation.
A bend cannot move at all, because it moves the replicates together. That is the whole of the separation. is the classical pure-error sum of squares, the part of the residual variation that no choice of model can explain because it is variation among runs at the same setting; the rest of the residual sum of squares is lack of fit. With one run at each setting, pure error has zero degrees of freedom and the split does not exist. In the replicated design it has four — two from the three runs at each end — and those four are what remain available when the spread is not known in advance: an estimate of from the replicates alone, which a bad run at the top inflates and a bend does not touch. The known-spread arithmetic below is the clean version of the argument, and the unknown- spread version keeps its structure while paying for the four degrees of freedom in power.
How well the two are separated
A simple rule reads first — significant at 5%, call it a bad run — and otherwise calls a significant top mean a bend.
| discrepancy | bad run named as bad | bend named as bend | bend called bad | bad run called a bend | one run at each x: flagged, either cause |
|---|---|---|---|---|---|
| 2σ | 28.95% | 44.41% | 5.00% | 6.82% | 38.97% |
| 3σ | 58.40% | 76.42% | 5.00% | 6.47% | 71.20% |
| 4σ | 84.10% | 91.56% | 5.00% | 3.82% | 91.91% |
| 5σ | 96.36% | 94.70% | 5.00% | 1.26% | 98.74% |
At four the replicated design names a bad run correctly 84.10% of the time and a bend 91.56%, and it mistakes one for the other 5.00% and 3.82% of the time. The spread design flags a four- discrepancy 91.91% of the time — slightly more often than the replicated design flags a bad run, 87.92%, because three runs at the top dilute one bad one — and it says nothing about which explanation holds. Its chance of naming the cause is a coin’s, not approximately but exactly.
The errors in the replicated column have a structure worth noticing. Calling a bend a bad run happens at exactly the test’s 5% at every size, because a bend leaves central. Calling a bad run a bend falls with the size of the discrepancy, because a larger bad observation is ever more obviously apart from its siblings. The design turns an unanswerable question into a test with a known error rate.
What the six moved runs cost
Replication looks like a luxury — six of twelve runs spent on only two settings — and the natural worry is that it buys the diagnosis with the estimate. It does the opposite.
The slope is estimated more precisely. Its standard error falls from 0.0836 to 0.0709, because the slope’s precision is set by the spread of the x values and runs at the ends contribute most to it. No run carries as much leverage: the largest falls from 0.295 to 0.235, since three runs share the top where one held it alone — which is the lesson of the essay on residuals turned into a design rule, and it also keeps any single bad run from drawing the line itself. A smooth bend is easier to detect: against a quadratic departure of the power rises from 44.7% to 66.5%, because a curvature term is estimated from the ends against the middle, and the ends are better known.
The cost is coverage. The replicated design observes eight distinct values of where the spread design observes twelve, and a departure confined to the settings it skips — or — is invisible to it, not weakly but entirely. That is the trade, stated exactly: replication makes the ends answerable and leaves gaps between them unasked. A study that has reason to expect a local feature in the interior — a threshold, a step, a narrow resonance — needs runs where the feature might be, and the two designs here are two answers to two different expectations.
What a residual plot recovers, and what it cannot
The diagnostics this line of essays has measured all work on the residuals, and it is worth being exact about which of them the design changes.
Twenty residual plots calibrated the eye against a correct model; the band the eye was standing in for priced its multiplicity; the essay on residuals showed that a raw residual understates exactly the influential points. Each of those is about seeing a discrepancy. None is about explaining one, and the design fact here sits beneath all of them.
On the spread design, a four- discrepancy at the top produces a large studentised residual at the top run and a large Cook’s distance there; a careful reader sees it 92% of the time. What the reader then does — delete the run as an outlier, or add a curvature term to the model — is a choice with no support in the data at all. Deleting it is right under one explanation and hides a real feature under the other; adding the term is right under the second and fits noise under the first. The choice is usually made by what the analyst expected, and the residual plot, which looks like evidence, has contributed nothing to it.
On the replicated design the same residual plot has something to show. The bad run is a single point far from two neighbours at the same ; the bend is three points together off the line. A reader looking at twelve points can see the difference, and the pure-error test makes the seeing into a number.
When the extreme is where the question is
The practical case for replicating the ends is strongest exactly where the ends matter most. A dose–response study’s highest dose is where toxicity or saturation would appear; a calibration’s extreme concentrations are where the instrument is least linear; an extrapolation leans entirely on the last point observed, as the prediction interval makes plain. In each case a surprising value at the extreme is either the most important finding in the study or the least interesting kind of error, and a single run there cannot say which.
The replicated design also answers a question the spread design never asks: is the spread of the errors the same at the ends as in the middle? Pure error at and at is an estimate of at each end with no model involved, and a heteroscedastic response shows up there directly. On the spread design the same question can only be asked of residuals, whose spreads already differ by the design with the model exactly right.
None of this requires a large study. Twelve runs were enough, and the rearrangement cost nothing in slope precision. What it requires is deciding before the runs are made that the ends are where a surprise would need an explanation — which is the same decision, in design form, as naming an analysis in advance: the evidence that can settle a question has to be collected before the question is asked.
The same fact in a factorial
Response-surface designs met this problem long ago and solved it in the same way, which is some evidence that the solution is not a quirk of straight lines. A two-level factorial puts every run at a corner, where a squared term equals one, so it cannot see curvature at all; the standard repair is a few runs at the centre. Those centre runs do two jobs at once, and they are the two jobs the replicates at the top did here. Their mean against the corners’ average is a test of curvature — the analogue of the top mean against the line — and their spread among themselves is pure error, an estimate of that no model assumption touches — the analogue of .
The factorial’s version makes one point clearer than the straight line does. The centre runs are replicates because the design has nowhere else to put an estimate of pure error that is free of the model, and the corners are unreplicated because each corner is already doing work in the effect estimates. A design that replicated only its corners would separate a bad corner from a curved surface in the way three runs at separate a bad run from a bend, at the cost of runs a factorial would rather spend elsewhere. Which of those is worth more is the same question as the straight-line one, asked in more dimensions, and the answer in both is that the design has to be chosen with the diagnostic in mind rather than the diagnostic chosen afterwards.
What to do with a spread design already run
Most data were not collected on a replicated design, and a reader with a surprising point at the end of a spread design still has to write something. The exact result above says what that something cannot be: a conclusion about the cause, drawn from the data.
What it can be is both answers. Fit the line with the point and without it; fit the line with a term that allows a bend at the top and without one. Report the slope and the prediction at the extreme under each, and say plainly that the data cannot choose between them. That is not a failure of analysis, and it is more useful to a reader than a confident deletion: it shows how much of the conclusion rests on a question the study did not ask. If the four readings agree, the point did not matter; if they disagree, the next study knows exactly where it needs three runs instead of one.
What replication at the end separates, and what the spread design keeps
With one run at the top of the range, a bad observation of size there and a bend of size confined there produce identical data from identical errors, so every diagnostic has the same distribution under both and none can name the cause better than chance.
With three runs at the top, the replicates’ spread about their own mean is central chi-square under a bend and noncentral under a bad run, and a rule reading it first names a four- bad run correctly 84.10% of the time and a four- bend 91.56%.
The replicated design estimates the slope more precisely — a standard error of 0.0709 against 0.0836 — carries a smaller largest leverage, 0.235 against 0.295, and detects a quadratic bend of with power 66.5% against 44.7%. It observes eight settings instead of twelve and is blind to departures confined to the four it skips.
Every rate is exact, with the spread known: normal tail areas for the top mean, and a noncentral chi-square computed as a Poisson mixture of central ones for the replicates’ spread. The hero figure uses one seeded set of twelve errors in all four panels, and the spread design’s two panels hold the same numbers, point for point.
Not claimed: that the rule used is optimal, or that the spread is known in practice. With it estimated, from the interior residuals and the replicates at both ends, the pure-error test becomes an F test on few degrees of freedom and loses power, though the identity that separates the explanations — a bend cannot move the replicates’ spread — holds exactly with the spread unknown. And the bend here is confined to the top setting; a curvature that spreads across several settings leaves traces in the interior that a spread design can read, which is a different and weaker version of the question.
Still open: how many replicates, and where
Three runs at each end is one choice. Two would give the pure-error test one degree of freedom at each end and very little power; four would give more power at the top and fewer interior settings. The trade between the number of replicates at the extremes and the number of distinct interior settings has a best point for any stated pair of worries — a bad run at an end and a local feature somewhere inside — and it has not been found here.
There is a second version of the question that matters more for large studies. Replicating every setting is what makes pure error available everywhere, and it halves the number of settings for a given budget. Replicating only the settings with the highest leverage — the ends in a straight-line study, the corners in a factorial — may buy almost all the diagnostic value for a fraction of the cost, and whether it does is a calculation over designs that this essay has only started.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A quantity that loses to a heuristic — both name experimental design, leverage, model diagnostics
- A set of pairs, not a vector — both name experimental design, leverage, model diagnostics
- Counting it exactly does not help — both name experimental design, leverage, model diagnostics
- Four datasets, one summary — both name leverage, model diagnostics, residual plot
- The summary that was meant to work — both name leverage, model diagnostics, residual plot
- Two numbers for the fit's geometry — both name leverage, model diagnostics, outlier
Named objects
A flat tag is an object no other essay names yet.
Experimental designInfluenceLack of fitLeverageModel diagnosticsOutlierPure-errorResidual plot