The points a slope rests on
Worth reading first: The line that one point drew.
Two numbers for the fit’s geometry found that no summary of a scatter’s dependence can flag Anscombe’s third and fourth datasets, and that a summary of the fit’s own geometry can: the largest deleted residual, which reads the data, and the largest leverage, which reads only the design. The second is the odd one out. It can be computed before a single outcome is measured, at the stage when a sample size is chosen, and it says whether the slope the study will report is going to rest on one observation.
The essay ended by asking whether a design-stage number of that kind could be stated routinely, and whether a threshold on it could be calibrated at all, given that once a design is fixed its leverage has no sampling distribution. Both have answers, and the more useful of the two turns out to be a number the essay did not name.
Where a slope’s information sits
A least-squares slope is a weighted sum of the outcomes, and the weight on observation is proportional to its covariate’s distance from the mean, . The variance of the slope is , and each point contributes a share of the sum in the denominator — its share of the design’s information about the slope. That share is the point’s leverage less , which is why the largest leverage in a design is the largest share of the slope’s information any one point holds.
Shares can be spread evenly or concentrated, and the standard way to count how many units a set of shares is really made of is the inverse of the sum of their squares. For a design it is
exactly. The effective number of points a slope rests on is the number of points divided by the covariate’s sample kurtosis — its fourth moment over its squared second — and nothing else about the design enters.
That identity puts a familiar quantity in an unfamiliar place. Kurtosis is usually read as a property of a distribution’s tails. Read through the slope, it is the factor by which a design’s point count overstates the number of points its slope actually rests on.
What the familiar covariates give
A design of two groups, half the points at each end, has kurtosis one and rests on all of its points; every point carries the same share. A uniform covariate has kurtosis 1.8, and a slope from a hundred uniform points rests on 55.1 of them. A normal covariate has kurtosis three, and a slope from a hundred points rests on 34.9 — about a third, for the covariate every textbook assumes. The points near the middle of a normal design are nearly useless to the slope, because they sit near the mean; the handful in the tails do the work.
Skewed and heavy-tailed covariates concentrate it further. An exponential covariate gives a hundred-point slope 16.0 effective points; Student’s t on three degrees of freedom, 14.8; a lognormal covariate with σ = 1 on the log scale — the shape of incomes, doses, concentrations, firm sizes and most quantities that are positive and multiply — gives 7.6. And the lognormal’s share does not stabilise as the design grows: at a thousand points the median is 30 effective points, 3.0% of the design, because the sample kurtosis of a lognormal keeps growing with n towards a population value of 114.
The light-tailed families have settled by then — normal at about a third, uniform at a little over a half — and the heavy-tailed ones have not. For a covariate with a finite fourth moment the share converges to one over the population kurtosis; for Student’s t on three degrees of freedom the fourth moment is infinite and the share falls without limit, and for a lognormal the convergence is so slow that no study size in practice reaches it. A larger study of a skewed covariate adds points faster than it adds effective points, and reports a sample size that overstates what its slope rests on by a factor that grows with the study.
In the single design drawn above, the largest point carries 24.6% of the slope’s information by itself, and the design rests on 10.8 effective points. A study on this design reports “n = 100” and a slope whose precision is roughly what eleven evenly spread points would give — with a quarter of it decided by one observation the study cannot check against any other.
The largest leverage follows the covariate’s tail
The largest share is the largest leverage less , and its distribution over designs is set by the same tail.
For a normal covariate the median largest leverage at a hundred points is 0.083, and no design in four thousand reaches one half. For an exponential covariate the median is 0.178; for Student’s t on three degrees of freedom, 0.190; for a lognormal, 0.300, and 15.7% of lognormal designs have a point with leverage above one half — a point whose own value accounts for most of its fitted value, and whose residual therefore says almost nothing about whether the line fits it. At thirty points the lognormal’s share above one half is 40.9%.
The line that one point drew built a design with one far point to show a slope reversed by it. The curves say such designs do not have to be built. They arrive by themselves, one time in six at a hundred points, whenever the covariate is the kind of quantity most regressions in the social and life sciences are run on.
A threshold set over designs
The obstacle to a rule on the largest leverage was that it has no sampling distribution for a fixed design: a design’s leverages are what they are, and there is no null hypothesis for them to be tested against. But a threshold does not need a sampling distribution within a design. It can be calibrated over designs, the way a reference distribution is calibrated over data — by stating a reference family and asking how unusual a design would be within it.
Take the normal covariate as the reference, since it is the family the textbook arithmetic for a slope’s precision implicitly assumes. At a hundred points, 95% of normal designs have a largest leverage below 0.123. A design that exceeds it is concentrated beyond what a normal covariate would produce one time in twenty.
Measured against that level, 84.1% of exponential designs of a hundred are flagged, 83.4% of Student’s t designs and 98.6% of lognormal ones. The rule does what a design-stage alarm should: it is silent on the designs the usual precision formulas describe well and it fires on nearly every design whose covariate is skewed. It says nothing about the outcome, and it does not need to — it is a statement about what the outcome will be able to say.
What the concentration does on a correct model
None of this depends on the model being wrong. On a straight line with normal errors, the fitted slope is unbiased and its reported standard error is correct, whatever the design. What the design changes is how much any single observation matters.
Deleting the highest-leverage point from a normal design of a hundred moves the slope by more than one standard error in 0.1% of datasets. From an exponential design, 4.9%; from Student’s t, 8.3%; from a lognormal, 16.3%. At thirty points the four rates are 4.4%, 18.1%, 18.0% and 28.4%. The rates fall as the design grows, but for the heavy-tailed families slowly: at a thousand points Student’s t still gives 2.4% and the lognormal 3.8%, where the light-tailed families are at zero. In a lognormal design of thirty, more than a quarter of honest datasets have a slope that one point moves by its own standard error — so a sensitivity analysis dropping that point would report a “fragile” result on a model with nothing wrong with it, and a result that is not fragile rests on the one point nobody can check.
That is the practical meaning of an effective sample of 7.6. The formula for the standard error knows about the concentration — it uses , which the far points dominate — so the reported precision is right. What it does not report is that the precision was bought from very few observations, and that any problem with those few, a recording error or a genuine curvature in the region only they occupy, reaches the slope at full strength.
Two far points are not twice as safe
A natural response to one point carrying a quarter of the slope is to hope for two. The effective number of points counts them correctly — two points each carrying a fifth of the information contribute about two effective points, not forty — but the diagnostics that read the data count them worse than one. Two points that hide each other placed a second far observation beside a first and found their Cook’s distances fell from 24.1 for one to 0.966 and 0.772 for two, neither crossing the conventional line, while together they reversed the slope. Each point’s deletion leaves the other to hold the line where it was.
So a skewed design with a small cluster of far points is the worst case for data-stage checking and the case the design-stage numbers describe without difficulty: the effective number of points is small, the largest leverage is moderate, and the fact that the slope rests on a handful of units is visible before the outcome is measured and invisible after.
A robust line does not create points
Robust regression is often reached for when a design looks like the lognormal one, and it answers a different question. A robust loss and a far x found that Huber’s loss, the standard robust line, downweights observations with large residuals and does nothing about observations with large leverage — a far point pulls the line towards itself and so never has a large residual to be downweighted for. Estimators that do resist leverage, like least trimmed squares, resist it by discarding the far points, and the start an efficient robust line inherits found them choosing between competing halves of the data when a far cluster is present.
Neither route creates information the design does not hold. A robust fit on a design with 7.6 effective points either follows those points or discards them, and in the second case its slope rests on the remaining ninety-odd points, which are nearly useless to it because they sit close together. The honest descriptions of the two outcomes are a slope decided by a few points and a slope decided by many points with little spread, and the design-stage numbers say which of the two a study is buying before it chooses an estimator.
Why the kurtosis, and not the spread
is a property of the design found that the spread of the covariate decides while the relationship stays fixed: five studies differing only in how far apart they placed their x values reported from 0.021 to 0.849. The spread is the second moment of the design. The effective number of points is the fourth moment over the square of the second, and the two answer different questions. The spread says how much the design can tell about the slope; the kurtosis says how many of the points the telling is shared among.
A design can have a large spread and a small effective sample — a lognormal covariate spreads widely because of a few far points — or a modest spread and a large one, as a tight two-group design does. A study choosing where to take its measurements controls both, and the usual advice to spread the x values as widely as possible addresses only the first. Spread achieved by including a few extreme units buys precision by concentrating it.
What a design can state before the data
The effective number of points, n over the design’s kurtosis. It needs only the covariate values, it is exact, and it translates a design into a count a reader already understands: a study of a hundred incomes whose slope rests on eight of them. It belongs beside the sample size in a protocol for the same reason the sample size itself belongs there — it is the number the precision is really computed from. A randomised trial meets the same asymmetry from another direction: balancing a skewed covariate found that which functions of a covariate a balancing rule should hold changes once the covariate is not symmetric, and a skewed covariate is exactly the one whose far tail concentrates a slope’s information.
The largest leverage, beside a threshold calibrated over a reference family. At a hundred points the normal family’s 95% level is 0.123. A design above it is concentrated in a way the textbook precision arithmetic does not picture, and a design above one half has a point whose fit cannot be checked by its residual.
A transformation, if the covariate’s scale is a choice. Taking the logarithm of a lognormal covariate makes it normal and gives the slope back its thirty-five effective points out of a hundred. Whether the relationship is linear on the log scale is a separate question, but it is one a study should ask before accepting a design in which eight points decide the answer.
The two diagnostics that read the data, afterwards. Two numbers for the fit’s geometry paired the largest leverage with the largest deleted residual. The design-stage numbers do not replace the second; they say in advance how much the second will be able to see, because a residual at a point with leverage near one is near zero whatever the point’s truth.
What is measured here and what is exact
The effective number of points a slope rests on is n divided by the design’s sample kurtosis, as an identity.
At a hundred points, its median is 34.9 for a normal covariate, 16.0 for an exponential and 7.6 for a lognormal with σ = 1, over four thousand designs each; the lognormal’s is 30 of a thousand at a thousand points.
The largest leverage exceeds one half in 15.7% of lognormal designs of a hundred and in none of four thousand normal ones. A threshold at the normal family’s 95% level, 0.123, flags 98.6% of lognormal designs.
On a correct line with unit errors, deleting the highest-leverage point moves the slope by more than its standard error in 16.3% of lognormal datasets of a hundred and 0.1% of normal ones.
Not claimed: that the normal family is the right reference for every field; a design field whose covariates are always skewed may prefer to calibrate against its own. Not claimed either that concentration is always a defect. A deliberately extreme design — two groups at the ends of a dose range — concentrates nothing and is efficient, and a design that concentrates its information in a few far points can be right when those points are measured with special care. The number says where the information is, not whether it should be there.
Still open: many covariates at once
Everything here is one covariate. A multiple regression’s leverages are the diagonal of the full hat matrix, and the information about each coefficient is concentrated or spread according to the design’s joint shape — a point can be ordinary on every covariate separately and extreme in their combination, as a young patient with a high dose is. The effective number of points generalises, but not by a single kurtosis: it depends on which coefficient is asked about, and on how correlated the covariates are.
How concentrated typical multivariable designs are, whether a threshold on the largest leverage can be calibrated over a reference family of correlated covariates the way it was over a single normal one, and how often the point that decides a coefficient is invisible on every one-dimensional plot of the design, are the measurements a design-stage report for multiple regression would need, and none of them has been made here.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A quantity that loses to a heuristic — both name experimental design, hat matrix, leverage, model diagnostics
- Counting it exactly does not help — both name experimental design, hat matrix, leverage, model diagnostics
- Residuals are not the errors — both name experimental design, influence, leverage, model diagnostics
- Three runs at the end of the line — both name experimental design, influence, leverage, model diagnostics
- What a chosen probe finds — both name effective sample size, experimental design, leverage, model diagnostics
- A penalty is a trace — both name effective sample size, hat matrix, least squares
Named objects
A flat tag is an object no other essay names yet.
Effective sample sizeExperimental designHat matrixInfluenceKurtosisLeast squaresLeverageModel diagnostics