Twenty residual plots
Worth reading first: What normal actually looks like · Four datasets, one summary.
The previous essay ended with the standard advice: plot the residuals. The advice is correct and it is incomplete, because it asks a question nobody is equipped to answer. Does this plot look wrong? Compared with what?
The reference nobody has
An analyst inspecting a residual plot is looking for curvature, funnelling and outliers. All three are present in the figure above, and the model that produced every panel is exactly correct: a straight-line relationship with normal errors of constant variance.
Panels appear to bend. Panels appear to widen. The largest single residual across the twenty reaches well over one standard deviation of the error, and in isolation would be circled.
None of it means anything, because none of it was put there. It is what twenty-four points of noise look like, twenty times.
The problem is not that analysts are careless. It is that the comparison requires a reference distribution of pictures, and the only way to acquire one is to have looked at many plots from data known to be correct — which almost nobody has done, because the null case is never displayed.
What the panel establishes
The device turns an impression into a calibrated statement, and it is worth being precise about the strength of the claim it supports.
Put the real residual plot among nineteen generated from data where the model holds, at the same sample size, drawn identically. Then:
If the real plot cannot be picked out, the departure is not evidence. Whatever structure it shows is within what correct data produces.
If it is clearly the most extreme of the twenty, that is a one-in-twenty observation under the null. Not a proof, and a calibrated statement rather than an impression — and the one-in-twenty is the conventional threshold, arrived at by counting rather than by looking up a critical value.
The formal version is the line-up test: place the real plot at a randomly chosen position among nineteen nulls, ask an observer who has not seen the data to identify the odd one, and the probability of a correct identification under the null is exactly 1/20. That is a genuine p-value, obtained from a visual judgement, and it is valid because the null distribution was constructed rather than assumed.
Most of the value arrives before the formality. Simply having a reference — any reference — removes the failure mode that matters, which is reading noise as structure.
Nineteen decoys is exactly five per cent
The number of decoys is not a matter of taste and it is worth saying what fixes it.
Under the null the real plot is exchangeable with the nineteen generated ones, so an analyst who has not been told which is which picks it by chance with probability exactly 1/20. The lineup is a 5% test, exactly, with no asymptotics, no distributional approximation and no reference to the shape of anything.
With m decoys the level is . Nine decoys gives 10%. Ninety-nine gives 1%.
And it composes. Two analysts who have not conferred, both picking the real panel, is a 1-in-400 event — a 0.25% test built out of two people looking at twenty pictures each. That is a cheaper route to a stringent level than showing one person a hundred panels, and it is available whenever there is a second reader.
What twenty-four honest points contain
The other half of the calibration is what a single null panel should be expected to hold, and the arithmetic explains why the largest residual in the figure is unremarkable.
The expected maximum of 24 standard normals is about 2.04 standard deviations, and the chance that a panel of 24 contains a residual beyond 3 is 6.3% — one panel in sixteen.
Across all twenty panels there are 480 residuals, whose expected maximum is about 3.1.
So a perfectly correct model, shown twenty times at this sample size, is expected to produce one residual at three standard deviations and twenty at two. Circling the largest point on the page is therefore not a diagnostic at all; it is a description of what the page had to contain.
The sample size has to match
The single most common way to break the device, and the one that makes it worse than useless.
The amount a residual plot wanders depends on how many points are in it, and it shrinks as the sample grows. Twenty panels at n = 500 show almost no visible structure; twenty at n = 12 look wild. A panel generated at the wrong sample size supplies a false reference and makes the comparison feel rigorous while being wrong in a known direction.
Dragging the sample size on the figure above makes the point directly: the same correct model, the same drawing rule, and the apparent pathology largely disappears between twenty-four points and a hundred.
So the reference has to be generated for the analysis in hand. That is a small cost — a loop of twenty draws — and it is not something that can be supplied once in a textbook.
Why the eye fails in one direction
The errors here are not symmetric, and the asymmetry explains why the failure is systematic rather than random.
Visual pattern-finding is extremely sensitive and has no calibration for how much pattern noise produces. It is very good at detecting a curve that is present and very bad at rejecting one that is not, so the mistakes overwhelmingly run toward seeing structure.
That direction is the damaging one. Missing a real curvature leads to a model that is slightly wrong and usually still useful. Seeing a curvature that is not there leads to adding a quadratic term, which raises R² for free, appears to improve the model, and cannot be disconfirmed by the same eye that requested it.
So the diagnostic without a reference does not merely fail to help. It feeds the model-elaboration loop that the rest of this field is trying to interrupt.
The same problem in every diagnostic read by eye
The device generalises exactly as far as the problem does, and the problem is everywhere.
A QQ plot has its own essay and the identical structure. So does a scale–location plot, an autocorrelation function, a scree plot, a heatmap of a correlation matrix, and any map of a rate across regions. In every case a practitioner is asked whether a pattern is real, and in every case the judgement needs a reference for how much pattern arises from nothing.
Maps are the most consequential of those and the least often calibrated. A map of disease rates by district always shows clusters, because rates computed from small denominators vary enormously, and the districts with the most extreme rates are reliably the ones with the fewest people — which is regression to the mean drawn geographically. Twenty maps generated from a constant true rate would settle it in seconds and are essentially never produced.
The remedy is the same wherever the problem appears. Draw the null. Generate the same display from data with the property assumed, at the same size, several times, and put the real one among them.
The discipline that makes it honest
One rule, and without it the whole device inverts.
The reference panel must be generated before the real plot is examined. A reference produced afterwards is a reference chosen with the answer already in view, and the temptation is to keep regenerating until the real plot looks unusual — or usual, depending on what is wanted.
That is the seed-picking problem in a new setting, and it has the same character: the individual panels are honest, the selection is what does the damage, and the result is reproducible and meaningless.
The protection this site uses on its own figures is the same one available to anyone: the seeds are contiguous and fixed in advance, so the panel is whatever the first twenty seeds produced rather than whatever looked best.
What a formal test adds, and what it removes
The alternative to looking is a numerical test — for curvature, add a squared term and test its coefficient; for heteroscedasticity, one of several standard tests. Both routes have real advantages and it is worth being clear about the trade.
A test scales and a plot does not. Twenty predictors mean twenty residual plots and nobody looks at twenty; a test runs on all of them and reports numbers. For models of any size the test is the only option.
A test has a stated error rate, where a visual judgement has one only if the line-up formality is observed.
And a test is specific, which is both its strength and its limitation. A test for curvature detects curvature and is blind to a single influential point, a discrete clump, a missing interaction, or the fourth Anscombe dataset. The plot is unfocused and therefore catches things nobody thought to test for.
The honest summary is that they fail differently and neither dominates. The test catches what it was aimed at with a known error rate; the plot catches what nobody anticipated with no error rate at all. A practice that uses only tests will miss the unanticipated failures, and a practice that uses only plots will miss the subtle ones and invent some that are not there.
Which is an argument for the line-up specifically, because it is the one device that gives an unfocused visual check a stated error rate — the thing each of the two alternatives lacks.
What this costs
Almost nothing, which is the reason to insist on it.
Twenty panels of a correct model is a loop over twenty seeds, fitting the same model to simulated data with the same design. On the sample sizes where the calibration problem is worst — a few dozen observations — it takes milliseconds. There is no data cost, no assumption beyond the model being checked, and no expertise required to read the result.
Set against that, the standard practice is to look at one plot with no reference and form an impression that is known to be biased in a specific direction.
The recommendation is therefore not “look at the residuals more carefully”. Careful looking is what produces the false positives. It is: generate the null, put the real one among it, and decide whether it can be picked out — which is less work than the careful looking and produces an answer that means something.
What the null has to hold fixed
The device requires generating data “where the model is right”, and that phrase hides a choice which determines what the check is capable of detecting.
The null panels here are generated from the fitted model: the same design, the same estimated slope and intercept, normal errors with the estimated variance. So they are a reference for the question is this residual structure consistent with my fitted model being correct?
Change what is held fixed and the question changes.
Generating with the same x values means the reference shares the design, so the check cannot detect anything about the design — the fourth Anscombe dataset’s eleven points at one x will look identical in every null panel, because every null panel has the same eleven points at one x. The device is blind to it by construction.
Generating with normal errors means the check is testing shape and variance jointly with normality. A departure could be any of the three.
And generating from the fitted estimates rather than from a hypothesised truth means the reference is tuned to the data, which is conservative: the real plot is being compared against nulls that were fitted to it, so it looks more typical than it should.
None of those is a flaw as long as it is known. The rule is that the null panel defines the hypothesis being tested, and a reader should be able to say what was held fixed. Where that is unstated, the check has no determinate meaning, whatever it shows.
The practical version: generate the nulls the way the data was supposed to have arisen, and if there is uncertainty about that, generate more than one kind of null.
Where the device came from
Worth a paragraph, because knowing it is a developed method rather than an improvisation changes how readily it can be defended in a review.
The line-up test was formalised in the visual-inference literature in the late 2000s, with the protocol described above: the real plot placed at random among nineteen nulls, an observer who has not seen the data asked to pick the odd one, and the resulting p-value of 1/20 under the null. Experiments established that observers do detect real departures at rates far above chance, so the method has power as well as validity, and that observers also “detect” departures in all-null line-ups at close to the chance rate — which is the calibration this essay is arguing nobody has by default.
Two refinements from that work are worth carrying.
Several independent observers sharpen it. If three people who have not seen the data all pick the same panel, the probability under the null is 1/8000 rather than 1/20, and the test becomes powerful rather than merely valid.
And the observer must be naive. An analyst who already knows which plot is the real one cannot serve as the observer, for the same reason a seed chosen after seeing the output is not a seed. This is the practical obstacle to the formal version, and it is why most of the benefit in solo work comes from the informal version — generate the nulls, look at them, and see whether the real plot stands out.
A note on what “the model is exactly right” means here
The figure’s claim is strong and the site’s gate has to make it checkable, so it is worth saying what is asserted.
Every panel is generated from a linear relationship with normal errors of constant variance, and the fit applied to it is a linear model — so the model is correct in the strongest available sense, not merely adequate. The check the gate runs is that at least one panel shows a residual far from zero, which establishes that the reference genuinely spans the range where a real analyst would suspect a problem. A reference panel where nothing ever looked wrong would fail to serve its purpose, and the assertion is there to catch a generator that produced twenty bland pictures.
That is the shape of every check on this site: state what the figure has to demonstrate, and require it, rather than trusting that the drawing came out right.
The conclusion the field reaches
Four essays, and they compose into one recommendation.
A regression summary is a claim about a shape, and four datasets can share a summary and differ completely. The check for shape is a residual plot. A residual plot cannot be read without knowing what a correct one looks like, and this essay supplies that. And two specific failures — one point owning the slope and a statistic that rewards free parameters — have numbers attached rather than pictures, and both numbers are computable from the standard fit.
So the field’s output is four things to compute beside every regression: the maximum leverage, the largest Cook’s distance, the adjusted R² against k/(n − 1), and a residual plot with its null reference. All four are cheap, none is in the standard output, and between them they distinguish every case this field has raised.
The reason none of them is standard is the one the whole site keeps arriving at. Each answers a question the method never claimed to answer, and a method is judged by whether it does what it says.
Which is fair to the method and no comfort at all to a reader trying to work out whether a published slope means anything. The four numbers are the difference, and they cost one function call each.
One sentence
Plotting the residuals is necessary and not sufficient; what makes it sufficient is generating the same plot from data the model would have produced, and looking at both. Everything else in this essay is an elaboration of that, and the elaboration is shorter than the effort currently spent squinting at a single plot with nothing to compare it to. The reference is twenty lines of code and it converts an impression into a measurement.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A boundary for giving up — both name error rate, sample size
- A degrees of freedom that is not a count — both name error rate, sample size
- Choosing n after looking — both name error rate, sample size
- Dropping the losers — both name error rate, sample size
Named objects
A flat tag is an object no other essay names yet.
Error rateThe line-up testModel diagnosticsResidual plotSample sizeVisual inference