What a summary of a scatter is a property of

A point ordinary on every axis

In a multiple regression each coefficient rests on its own number of points, n divided by the kurtosis of that covariate's residual after the others, exactly. Correlation between normal covariates inflates a coefficient's variance fivefold and leaves it resting on the same third of the points. What correlation does instead is hide the point that decides the coefficient: at a correlation of 0.9 it stands out on no single covariate's plot in 77.3% of designs of a hundred, and the largest leverage a normal design reaches depends on the number of points and covariates and on nothing else.

Worth reading first: The line that one point drew.

The points a slope rests on found that a straight line’s slope rests on the number of points divided by the covariate’s kurtosis, exactly, and that the number can be stated before a single outcome is measured: about 35 of a hundred for a normal covariate, 7.6 for a lognormal one. It ended on the case that essay could not reach. Most regressions have several covariates, a point can be ordinary on each of them and extreme in their combination, and the effective count, it said, would generalise “but not by a single kurtosis”.

It generalises by a single kurtosis after all — one per coefficient, taken of a different quantity. And the two things a reader would expect correlated covariates to do to it, they do not do. They do something else, which is visible in one picture of a hundred points.

A point that decides a coefficient, on two covariates correlated at 0.9A hundred points on two normal covariates with correlation 0.9. Dashed lines mark where the most extreme five per cent of each covariate begin. The point carrying the largest share of the first coefficient's information, 14.2%, is inside the lines on both axes, so neither covariate's plot on its own would show it. The first coefficient rests on 26.1 effective points, with a variance inflation of 6.27.-202-202first covariatesecond covariate14.2% of the first coefficientone seeded design of 100, correlation 0.9the deciding point shows on neither margin
Fig. 1 A hundred points on two normal covariates correlated at 0.9. The dashed lines mark where each covariate’s most extreme five per cent begin; the marked point carries the largest share of the first coefficient’s information and sits inside both pairs of lines. The slider sets the correlation.

Each coefficient reads one direction of the design

A fitted coefficient in a multiple regression is, like a simple slope, a weighted sum of the outcomes. The Frisch–Waugh theorem says which weights. Regress the first covariate on an intercept and all the others, and call what is left over its partial residual, rir_i for point ii. The first coefficient is then

β^1=∑iriyi∑iri2,\hat\beta_1 = \frac{\sum_i r_i y_i}{\sum_i r_i^2},

with variance σ2/∑ri2\sigma^2/\sum r_i^2, and nothing else about the design enters it. Each point contributes a share ri2/∑rj2r_i^2/\sum r_j^2 of the coefficient’s information. So the argument of the one-covariate essay goes through word for word with rr in place of the centred covariate, and the effective number of points behind the first coefficient is

neff,1=(∑ri2)2∑ri4=nkurtosis of r,n_{\text{eff},1} = \frac{\left(\sum r_i^2\right)^2}{\sum r_i^4} = \frac{n}{\text{kurtosis of } r},

exactly. The partial residual is the part of the first covariate that the others cannot predict: the direction in the design along which the first coefficient is learnt. Its kurtosis says how many points that direction is shared among, and each coefficient has its own.

Correlation costs precision, not points

The obvious expectation is that correlated covariates concentrate a coefficient’s information. They certainly cost it information: the sum ∑ri2\sum r_i^2 shrinks as the other covariates predict more of the first, and the ratio of the covariate’s own sum of squares to its partial residual’s is the variance inflation factor, the number every regression textbook uses to measure collinearity.

How many of a hundred points the first coefficient rests on, by covariate count, correlation 0.9. Median effective number of points behind the first coefficient over a thousand designs of a hundred at correlation 0.9: normal with 1, 34.9 (variance inflation 1.00); normal with 2, 35.2 (variance inflation 5.32); normal with 5, 35.1 (variance inflation 8.34); normal with 10, 34.8 (variance inflation 9.93); lognormal with 1, 7.7 (variance inflation 1.00); lognormal with 2, 9.1 (variance inflation 3.87); lognormal with 5, 11.5 (variance inflation 6.97); lognormal with 10, 15.4 (variance inflation 10.20).
Fig. 2 The effective number of points behind the first coefficient in designs of a hundred, at a correlation of 0.9 between every pair of covariates, with the variance inflation beside each bar.

For normal covariates, the count does not move. At no correlation the first coefficient of a hundred-point design rests on a median of 34.9 points with one covariate, 35.0 with two, 35.4 with five and 35.0 with ten. At a correlation of 0.9 between every pair, the variance inflation is 5.32 with two covariates, 8.34 with five and 9.93 with ten, and the effective counts are 35.2, 35.1 and 34.8 — the same third of the points, carrying a fifth or a tenth of the information they carried before.

The reason is a property of the normal law rather than of least squares. The part of one normal variable that a set of jointly normal variables cannot predict is itself normal, with a smaller variance and the same shape, so its kurtosis is three whatever the correlation, and three is what the count divides by. Correlation shrinks the direction the coefficient is learnt along. It does not change how evenly the points are spread along it.

Skewed covariates behave differently in both respects, and in the opposite direction from what one might guess. With lognormal covariates and no correlation, the first coefficient rests on 7.7 points of a hundred with one covariate, within a tenth of the 7.6 the one-covariate essay found over four thousand designs, and on 9.1 with ten. At a correlation of 0.9 it rests on 9.1 with two covariates, 11.5 with five and 15.4 with ten. Correlation makes the skewed design less concentrated. When several lognormal covariates share their large values, part of each one’s far tail is predictable from the others and is removed with them, and what is left over is less dominated by a few extreme units. The variance inflation at the same setting is 3.87, 6.97 and 10.20, so the coefficient is still learnt far less precisely than without the correlation. It is learnt from somewhat more of the points.

So the two numbers answer different questions, as the spread and the kurtosis did for a single covariate in R2R^2 is a property of the design. The variance inflation says how much information the other covariates took away. The effective count says how many points the remainder is shared among. A design report that gives only the first is silent about the second, and for a normal design the second never moves, which is why collinearity diagnostics have been able to ignore it.

A threshold that needs two numbers and no more

Two numbers for the fit’s geometry paired the largest leverage with the largest deleted residual, and its successor asked whether a threshold on the largest leverage could be calibrated at all. Over normal designs with one covariate it could: at a hundred points 95% of designs have a largest leverage below 0.123. The question for several covariates is what the reference family should be, and in particular how correlated its covariates should be taken to be.

The answer is that it does not matter, and the reason is exact. A point’s leverage in a fit with an intercept is

hi=1n+di2n−1,h_i = \frac{1}{n} + \frac{d_i^2}{n-1},

where did_i is the point’s Mahalanobis distance from the centre of the design, computed with the design’s own covariance. That distance is unchanged by any invertible linear map of the covariates, and a correlated normal design is exactly such a map of an independent one. So every leverage in a correlated normal design equals the leverage the corresponding independent design would give, and the distribution of the largest leverage over normal designs depends on the number of points and the number of covariates and on nothing else.

The largest leverage a normal design reaches, by the number of covariates and points. The 95% level of the largest leverage over normal designs. At a hundred points: 1 covariate, 0.123; 2 covariates, 0.153; 3 covariates, 0.177; 5 covariates, 0.214; 7 covariates, 0.248; 10 covariates, 0.294. At fifty points the ten-covariate level is 0.511; at two hundred, 0.162. Designs with correlation 0.9 give the same levels exactly. The conventional rule 2(p + 1)/n sits at 0.04 to 0.22 for a hundred points.
Fig. 3 The level that 5% of normal designs’ largest leverage exceeds, against the number of covariates, for fifty, a hundred and two hundred points. Open circles recompute the hundred-point designs at a correlation of 0.9 and land on the line exactly. The dashed line is the usual rule of thumb at a hundred points.

At a hundred points the 95% level is 0.123 with one covariate, 0.153 with two, 0.177 with three, 0.214 with five, 0.248 with seven and 0.294 with ten. At fifty points and ten covariates it is 0.511 — one normal design in twenty has a point with leverage above one half, with nothing unusual about it — and at two hundred points and ten covariates, 0.162. The circles drawn at a correlation of 0.9 are not close to the line; they are on it, to every digit, because they are the same numbers.

A table of that shape is all a design-stage threshold for multiple regression needs. It is a two-way table in nn and pp that could sit in any methods appendix, and a design whose largest leverage exceeds its entry is concentrated beyond what a normal design of its size and dimension produces one time in twenty.

The rule of thumb and what it is a rule about

The threshold usually quoted for leverage is 2(p+1)/n2(p+1)/n, twice the average leverage, and the dashed line in the figure is that rule at a hundred points. It sits below the curve at every dimension, and far below it where the covariates are few. Over the same normal designs of a hundred, the rule flags at least one point in every design with one covariate, with 8.39 points over the line on average; with ten covariates it still flags 77.1% of designs.

The rule is not wrong so much as answering a different question. It is a statement about a single point — whether this point’s leverage is large compared with the average — and applied to every point of a design it will find a few in every honest one, the same arithmetic that makes twenty significance tests find one. The level in the figure is a statement about the design: whether its largest leverage is unusual for a design of its size. The ratio between the two is not constant, which is why no multiple of the average could replace the table: at a hundred points the design-level threshold is 6.2 times the average leverage with one covariate and 2.7 times with ten.

Skewed covariates, in several dimensions at once

The one-covariate essay found that 15.7% of lognormal designs of a hundred have a point with leverage above one half. Adding covariates adds chances. Each lognormal covariate brings its own far tail, and the leverage of a point is its distance from the centre in all of them at once.

How often a lognormal design of a hundred has a point with leverage above one half, correlation 0. Share of a thousand lognormal designs of a hundred whose largest leverage exceeds one half, at correlation 0: 1 covariate, 16.4%; 2 covariates, 31.0%; 3 covariates, 43.9%; 5 covariates, 64.2%; 7 covariates, 78.2%; 10 covariates, 90.8%. No normal design of a hundred reaches one half at any of these dimensions.
Fig. 4 The share of lognormal designs of a hundred whose largest leverage exceeds one half, by the number of covariates. No normal design of a hundred reaches one half at any of these dimensions.

With independent lognormal covariates, the share of designs with a point above one half is 16.4% at one covariate, 31.0% at two, 43.9% at three, 64.2% at five, 78.2% at seven and 90.8% at ten. Correlating the covariates at 0.9 makes it worse, not better — 40.4% at two, 88.1% at five, 99.8% at ten — because the far units are now far on several covariates together and their Mahalanobis distance adds up in the direction the covariates share.

A regression of an outcome on ten positive, multiplicative quantities — incomes, doses, sizes, concentrations — is therefore, nine times in ten, a regression with at least one point whose own value accounts for more than half of its fitted value. Residuals are not the errors showed that such a point’s residual has less than half the variance of an ordinary one’s, so a residual plot cannot show whether the model fits it. The design says so before the outcome exists, and it says so more loudly the more covariates the model is given.

Correlation hides the point that decides

The one thing correlation does to a normal design’s coefficient, then, is not concentration. It is concealment.

The partial residual of the first covariate is its disagreement with the others. At a correlation of 0.9 between two normal covariates, that disagreement has a standard deviation of 1−0.92=0.44\sqrt{1-0.9^2} = 0.44 in the covariates’ own units. A point three standard deviations out in the direction of disagreement — as far out as any point in a design of a hundred will be — is about 1.3 units from where the other covariate predicts it, which is a perfectly ordinary value on each covariate taken alone. The hero figure is such a design: the marked point carries 14.2% of the first coefficient’s information, 3.7 times the share each point would carry if the information were spread evenly over its 26.1 effective points, and neither of its coordinates is among the five per cent most extreme values of its covariate.

How often the point that decides a coefficient is invisible on every one-covariate plot. Over a thousand normal designs of a hundred, the share in which the point carrying most of the first coefficient's information is not among the five per cent most extreme values of any single covariate. With 2 covariates: 0.1% at 0, 2.6% at 0.25, 22.9% at 0.5, 53.2% at 0.75, 77.3% at 0.9. With 5 covariates: 1.3% at 0, 11.0% at 0.25, 35.9% at 0.5, 63.3% at 0.75, 79.1% at 0.9. With 10 covariates: 2.9% at 0, 13.7% at 0.25, 39.6% at 0.5, 64.0% at 0.75, 80.7% at 0.9.
Fig. 5 The share of normal designs of a hundred in which the point carrying most of the first coefficient’s information is not among the five per cent most extreme values of any single covariate, against the correlation between every pair of covariates.

With two covariates and no correlation, the deciding point stands out on some margin in all but 0.1% of designs: it is simply the point farthest out on the first covariate. At a correlation of 0.25 it is hidden in 2.6%; at 0.5, 22.9%; at 0.75, 53.2%; at 0.9, 77.3%. With more covariates, hiding starts earlier — 11.0% at a correlation of 0.25 with five covariates and 13.7% with ten — and reaches 79.1% and 80.7% at 0.9.

The point with the largest leverage is hidden less often, because leverage measures distance in every direction, including the shared one along which extreme points are extreme on every margin. With two covariates at 0.9 it is invisible on both plots in 26.3% of designs; with ten, in 53.0%. The coefficient’s deciding point is the harder one to see, because it is defined by the single direction the covariates’ correlation makes narrow.

Skewed designs do not hide their points in the same way. In lognormal designs the far units are far on their own covariates, and the deciding point is hidden in 3.9% of designs with two covariates at 0.9 and 28.9% with ten. A skewed design concentrates its coefficients on a few points that anyone can see; a correlated normal design spreads them over a third of its points and then puts the most important of those where nobody looks.

Why the scatterplot matrix does not help

The standard first look at a multivariable design is a scatterplot matrix, or a histogram of each covariate, and twenty residual plots found that even a trained eye needs a reference to know what an ordinary plot looks like. The measurement above says that for the question of which point a coefficient rests on, the one-covariate plots are looking in the wrong direction by construction. In the hero figure the pair’s own scatterplot does show the point — it lies visibly off the diagonal — but with five or ten covariates the direction of disagreement need not lie in any pair’s plane, and the share hidden with ten covariates at a correlation of 0.25 is already 13.7%.

The plot that does look in the right direction is the one whose horizontal axis is the partial residual itself, the added-variable plot, which draws each point at its share of the coefficient’s information. Its axis is the quantity the effective count is computed from, and its most extreme point is the deciding point by definition. It is a standard plot, and it is almost always drawn after the outcome is measured, as a check on the fit. The measurement here is an argument for drawing its horizontal axis before, as a statement about the design.

Two far points, several directions

Two points that hide each other found that two far observations together can reverse a slope while each one’s deletion diagnostic says nothing, because each holds the line where the other would have left it. In several dimensions the masking has more room. Two points that disagree with the correlation in the same direction share the coefficient’s information between them, and the deletion of either leaves the other to carry it; two points out in different directions of disagreement each decide a different coefficient and may each be hidden on every margin. The line that one point drew needed a deliberately placed point to make its case. A correlated design produces the multivariable version by chance, most of the time.

A robust loss and a far x found that Huber’s loss does nothing about leverage, and in several dimensions that holds for the same reason: a point far out in a direction of disagreement pulls each fit towards itself and leaves no large residual to be down-weighted. The design-stage count is again the more useful statement. A coefficient resting on 35 effective points with a variance inflation of ten is learnt from a third of the design along a direction a tenth as wide as the covariate, and no choice of loss function changes either number.

What a report of a multivariable design can state

For each coefficient, the effective number of points, nn over the kurtosis of its covariate’s partial residual. It is exact, it needs only the covariates, and it is different for each coefficient. For normal covariates it is about a third of nn whatever the correlation; for skewed ones it is small and it moves.

For each coefficient, the variance inflation beside it. The two are not substitutes. A design can inflate a coefficient’s variance tenfold and leave its count at a third, and a skewed design can concentrate a coefficient on eight points with no inflation at all.

The largest leverage, against the table in nn and pp. For a normal reference family the table has no third dimension: correlation among the covariates changes no leverage at all. A design above its entry is unusual for its size, whatever the correlation.

The partial residual of each covariate the study cares about, plotted. It is the only one-dimensional plot that shows the point a coefficient rests on, and at the correlations observational studies routinely have, the ordinary plots miss that point more often than they show it. Balancing more than one number found that a randomised trial can hold several covariates in balance at once through their Mahalanobis distance; an observational design has no such control, and the partial residual is where its imbalance becomes a coefficient’s concentration. And adjusting for everything found that adding every measured covariate often leaves a larger bias than adding none. Each covariate added is also a direction taken away from every other coefficient’s partial residual, which is one more reason to count the points a coefficient still rests on after the adjustment.

What is exact and what was drawn

Exact: the first coefficient of a least-squares fit is its covariate’s partial residual weighted against the outcome; its effective number of points is nn over that residual’s kurtosis; and the leverages of a normal design with any correlation are those of the independent design it is a linear map of, so the reference level for the largest leverage depends on the number of points and covariates alone.

Drawn, over a thousand designs of a hundred for each setting and four thousand for the reference levels at a hundred points: the counts of 34.6 to 35.4 points for normal designs at every dimension and correlation measured; the lognormal counts of 7.7 to 15.4; the 95% levels of 0.123 to 0.294 for one to ten covariates; the share of lognormal designs above one half, 90.8% with ten independent covariates; and the 77.3% of two-covariate normal designs at a correlation of 0.9 whose deciding point stands out on neither covariate’s plot.

Not claimed: that “among the five per cent most extreme” is the threshold every reader applies to a plot. A stricter reader hides more points and a looser one fewer; the rise from almost none to more than three quarters as the correlation grows does not depend on the choice. Not claimed either that equicorrelation is how real covariates are arranged. It is the simplest arrangement that has a single number for the correlation, and the invariance of the leverage holds for every arrangement; the hiding depends on the arrangement, and a design with one pair of strongly correlated covariates among independent ones will hide the points deciding that pair’s coefficients and no others.

Still open: a term built from the others

Every covariate here is a variable in its own right. Many regressions also carry terms built from their covariates — a squared term to allow a curve, a product of two covariates to allow an interaction — and a built term is never normal, even when every variable it is built from is. The square of a standard normal covariate, centred, has kurtosis fifteen, so the coefficient on a quadratic term in an otherwise normal design should rest on something like a fifteenth of the points; the product of two independent normals has kurtosis nine.

That is a prediction from the identity above, and it has not been measured: how many points an interaction coefficient in a typical trial analysis rests on, whether the effective count behind a squared term depends on where the covariate was centred, and how often the point that decides an interaction is a point nobody would flag on any plot of the covariates it was built from. Interactions are the coefficients that subgroup claims are made on, which is reason enough to count them.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Effective sample sizeExperimental designHat matrixKurtosisLeast squaresLeverageMahalanobis distanceModel diagnosticsVariance inflation