Residuals of a model the data chose
Worth reading first: What normal actually looks like.
Residuals that look too normal put the errors of forty rows through least-squares fits of more and more columns and read the residuals on a quantile plot. Each residual is a mix of all the errors, the mix moves towards a bell as the fit takes more columns, and a skewed error that the straightness test flags 82.9% of the time on the errors themselves was flagged 51.2% of the time on the residuals of a ten-column fit and 23.0% on a twenty-column one. Residuals made independent then tried every rotation of those residuals that has an exact null and found none that saw the shape better.
Every fit in those essays had its columns fixed before the data arrived. Most published regressions do not. Their columns were chosen — by a stepwise search, by a lasso, by keeping the interactions that came out significant — and the choice was made with the same errors the residual plot is then asked about. The earlier essay ended on the obvious suspicion: a column that lines up with the three largest errors is the column a search keeps, and the residuals it leaves have had their right tail removed on purpose.
A pool of candidates that explain nothing
The rows are forty, the outcome is pure error, and there is a pool of twenty candidate columns, each an independent normal covariate, fixed once and unrelated to anything. A fit of columns is either the first of the pool — columns fixed in advance, like the earlier essays’ designs — or the that a forward search picks, adding at each step the candidate that most reduces the residual sum of squares given the ones already in.
Nothing in the pool is related to the outcome, so every column the search picks is fitting noise, and every column it picks is fitting the particular noise in this sample. That is the whole difference between the two fits. The fixed fit’s residuals are the errors times a matrix chosen before the errors were seen; the chosen fit’s residuals are the errors times a matrix chosen by the errors.
The errors are drawn from the departures the earlier essays priced: a skewed error, lognormal with ; a light-tailed one, uniform; a heavy-tailed one, Student’s t on five degrees of freedom. Each is read by the shape test that sees it best — the straightness of the plot for the skewed error, its kurtosis for the light one, its worst point for the heavy one — against the 5% point for forty independent normal observations, the band a residual plot is ordinarily read against.
A chosen column takes the shape with it
With the intercept alone the straightness test flags the skewed error 82.5% of the time. One fixed column takes that to 79.9%, two to 76.1%, five to 65.4% and ten to 47.0% — the slide the earlier essay measured. One chosen column takes it to 66.6%: the single best of twenty noise columns costs nearly as much as five fixed ones. Two chosen columns leave 55.5%, five leave 36.8%, and ten leave 25.8%.
A light-tailed error suffers more. Fixed fits see it 63.9%, 55.4%, 34.5% and 15.8% of the time at one, two, five and ten columns; chosen fits 32.6%, 19.4%, 8.5% and 5.0%. Five columns of noise picked by a search make a uniform error almost invisible to the test built to see it: 8.5% against a 5% level. A heavy-tailed error is the least affected, 26.4% fixed and 17.7% chosen at five columns, because its signature is a single wild point, and one column can fit one point only so far.
Why a search removes the shape
The sample drawn above is ordinary. The fixed fit’s residuals keep the skewed errors’ shape: their straightness statistic is 0.1090, twice the band’s 5% point of 0.0552, and their skewness 1.283. The chosen fit’s residuals, from the same errors, give 0.0441 and 0.592 — inside the band, with half the skewness. A reader shown the second plot sees errors that look normal.
The mechanism is the one the essay before named, and it can be said precisely. A skewed sample’s shape is carried by its few largest values. A forward search ranks candidates by how much residual sum of squares they remove, and the sum of squares of a skewed sample is dominated by those same few values, so the candidate that wins is the one most aligned with them. Fitting it pulls those values in. Each step of the search does the same to whatever large residuals are left. The search is not looking for the shape; it is looking for the largest squared residuals, and in a skewed sample those are the shape.
A light-tailed sample’s shape is the opposite: too few large values, a flat middle. Its sum of squares is spread evenly, so there are no few values for the search to chase. What the chosen columns do instead is mix, and they mix harder than fixed ones: each was picked because it is aligned with this sample, so each removes more of the sample’s own variation, and what is left is a combination of the errors with less of any single error in it. Residuals that look too normal found light tails vanishing faster than skew under any mixing — 73.0% on the errors to 4.7% on twenty fixed columns — and a search reaches that floor with five.
The band was never the problem
The earlier essay’s suggestion was a band simulated by repeating the selection on simulated normal errors, so that the plot’s null would know about the search. That band can be built, and it answers a question nobody needed answered.
On normal errors the chosen fit’s residuals are flagged 6.2% of the time against the ordinary band, 5.9% against the fixed design’s band and 5.1% against the band simulated with the search. The search barely disturbs the level, because residuals of normal errors look normal however the columns were chosen — fitting the largest values of a normal sample leaves a normal-looking sample. Simulating the search moves the level by about a point.
And it moves the power by nothing useful. The skewed errors are flagged 36.8%, 36.2% and 34.0% against the three bands; the light-tailed ones 8.5%, 7.5% and 7.0%. The selection-aware band is slightly wider, so it sees slightly less. The plot of a chosen fit is not miscalibrated. It is blind, and calibrating a blind instrument gives an honest 5% test with no power.
That separates two questions the band the eye was standing in for treated as one. A band answers how often a plot of normal errors would look like this; it cannot answer how often a plot of non-normal errors would look normal, and after a search the second number is the one that has collapsed.
How many columns a search is worth
The cost of the search can be stated as a number of columns. A chosen fit of one column has the residual-plot power of a fixed fit of 4.6 columns against the skewed error; two chosen columns, of 7.8; five, of 13.4; ten, of 17.6. Against the light-tailed error the equivalents are 5.5, 8.8, 14.4 and 18.2; against the heavy-tailed one, 3.6, 6.5, 12.5 and 17.6.
The ratio is largest for the smallest searches. One chosen column is worth four or five fixed ones, because the first step of a search over twenty is a choice of the best of twenty: the column it keeps is the noise column most aligned with this sample’s largest residuals. By ten columns the search has used half the pool, and there are fewer peculiar alignments left to pick: ten chosen columns are worth 17.6 fixed ones, not forty. The pool bounds the cost — a search can never be worth more columns than it searched.
What sets the cost is the size of the pool, not the size of the model. Keep five columns and vary how many were searched: five chosen from five is no choice at all and flags the skewed error 65.4% of the time, exactly the fixed fit; five chosen from ten, 50.8%; from twenty, 36.8%; from forty, 25.9%. The light-tailed error goes 34.5%, 17.8%, 8.5% and 5.9% — at forty candidates the test is at its own level. Every doubling of the pool takes between a fifth and a third of what is left of the plot’s power against skew, at a fixed number of columns kept. Two models that both report five predictors can have residual plots of very different worth, and only the pool says which.
So a model reported as “five predictors, chosen by forward selection from twenty” has a residual plot with the power of a model with thirteen to fourteen fixed predictors. The reader who is told the final column count and reads the plot accordingly credits it with the 65.4% of five fixed columns when it has 36.8% — nearly twice the power it has against skew, and four times against light tails. is not a fit found twenty useless predictors on thirty points manufacturing an of 0.69; a search over twenty manufactures a normal-looking plot from a fraction of them.
Splitting the rows does not rescue it
The standard remedy for a selection effect is to choose on one part of the data and judge on another. Here that means running the search on the first twenty rows, fitting the columns it chose to the other twenty, and reading the quantile plot of those twenty residuals against the band for twenty normals. None of the plotted residuals’ columns was fitted to the errors being plotted, so the plot is honest.
It is also weaker. With five columns, the split plot flags the skewed error 24.8% of the time, against the chosen fit’s 36.8% on all forty rows — the same draws, so the comparison is exact. With two columns, 37.8% against 55.5%. On the light-tailed error it gains a little, 10.5% against 8.5% at five columns, and on the heavy-tailed one it loses, 11.8% against 17.7%. The split plot sees almost exactly what a fixed five-column fit on twenty rows sees — 23.7% on the skewed error — which is what it is: honest columns on half the rows.
So twenty honest rows lose to forty contaminated ones for every departure but the light-tailed, where both are nearly blind. The search’s cost was about eight columns of power; halving the rows costs more. The split is the right repair for an estimate or an interval, where contamination biases the answer, and the wrong one for a diagnostic, where the contamination only weakens a check that halving weakens further.
The plot that keeps its power is the one drawn before the search: the residuals of the intercept alone, or of whatever columns the analysis was committed to in advance, on all forty rows. With the intercept alone the skewed error is flagged 82.5% of the time — more than twice the chosen fit’s rate and more than three times the split’s.
What this does to the inference the plot was guarding
A residual plot is drawn to check an assumption the rest of the analysis leans on. If the errors are skewed, intervals for small-sample predictions are wrong in a direction a normal plot would have shown, and a tail the sample never saw priced how wrong a sample-based correction goes when the tail is the part missing. After a search, the plot looks normal on two skewed samples in three and on more than nine light-tailed samples in ten, and the analysis proceeds on an assumption the data would have refuted had the columns not been chosen from them.
It is the same structure as every selection effect in this collection. The interval after the choice found a confidence interval made too narrow by a choice its construction did not know about; the winner’s curse found an estimate made too large by being the one selected. A residual plot after selection is the diagnostic version: the check made too lenient by a choice that used the very feature it checks for.
What a regression’s residual plot should come with
Say how the columns were chosen. A plot of residuals from a searched model should be read as a plot from a model of more columns than it has, and the number of candidates searched is what fixes how many more. Twenty candidates and five kept is worth about thirteen; the report should carry both numbers.
Check the shape before searching. The errors’ shape can be read off the residuals of a model fixed in advance — the intercept alone, or the columns theory requires — before any search is run. That plot has the fixed design’s power, and its verdict is not contaminated by columns chosen to fit the same errors.
Do not reach for a selection-aware band, or a sample split, to fix power. The first fixes a level that was never broken; the second trades a contamination worth eight columns for a halving that costs more. If the question is whether the errors are normal, the residuals of a chosen model are the wrong sample to ask, and the residuals of the model fixed in advance, on every row, are the right one.
Expect light tails to vanish first. The departure the search hides most completely is the one that makes intervals too wide rather than too narrow, which is also the one least often looked for. A clean plot after a search says little about tails in either direction.
Exact, and counted
Exact: with the outcome pure error and the pool unrelated to it, every column the search picks fits noise; the forward search’s criterion is the squared projection of the current residuals on each candidate, so it is a ranking by alignment with the largest residuals.
Counted, over three thousand samples of forty at each column count and departure: the fixed fits’ power of 79.9%, 76.1%, 65.4% and 47.0% against the skewed error at one, two, five and ten columns, and the chosen fits’ 66.6%, 55.5%, 36.8% and 25.8%; the level of 6.2%, 5.9% and 5.1% on normal errors at five chosen columns against the three bands; the equivalent fixed column counts, 4.6, 7.8, 13.4 and 17.6, interpolated on a fixed-power curve counted at every column count from none to twenty.
Counted the same way: five columns chosen from pools of five, ten, twenty and forty, flagging the skewed error 65.4%, 50.8%, 36.8% and 25.9% of the time; and the split plot, the search on twenty rows and the plot on the other twenty, flagging it 24.8% at five columns and 37.8% at two, on the same error draws as the chosen fits it is compared with.
Not claimed: that every search hides as much. A lasso shrinks rather than fully fitting each chosen column, and should take less of the shape per column; a search over a pool that contains real predictors spends some of its steps on them. Both move these numbers, and neither changes the direction.
Still open: a search that shrinks instead of choosing
A forward search fits each chosen column completely. A lasso fits many columns partly, shrinking each coefficient towards zero by an amount set by a penalty, and it is how most columns are chosen now. A shrunken column removes less of the errors it is aligned with, so per column it should take less of the shape; but it keeps more columns, and each was still chosen by its alignment with this sample.
How often a lasso’s residual plot sees a skewed or light-tailed error at the penalty its cross-validation picks, how many fixed columns that plot is worth, and whether the answer depends more on the number of columns kept or on the strength of the shrinkage, are measurable on the same pool and the same errors, and have not been measured here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A number for the shape — both name normality, q–q plot, statistical power, visual inference
- A criterion is a prediction of the hold-out — both name model selection, residual, specification search
- Residuals are not the errors — both name model diagnostics, normality, q–q plot
- The plot is about the wrong quantity — both name model diagnostics, normality, q–q plot
- When the benchmark is a candidate — both name model selection, residual, specification search
- A break that was looked for — both name model selection, specification search
Named objects
A flat tag is an object no other essay names yet.
Model diagnosticsModel selectionNormalityQ–Q plotResidualSpecification searchStatistical powerVisual inference