A term built from the others
Worth reading first: The line that one point drew.
A point ordinary on every axis found that each coefficient of a multiple regression rests on its own number of points: the number of points divided by the kurtosis of that covariate’s partial residual, the part of it the other covariates cannot predict. For normal covariates that is about a third of the points whatever their correlation, and the essay closed on the coefficients it had not counted. Many regressions carry terms built from their covariates — a square to let a relation bend, a product to let one covariate’s effect depend on another — and a built term is never normal, even when everything it is built from is.
The identity does not care how a column was made. So the prediction was that a squared normal covariate’s coefficient rests on about a fifteenth of the points and the product of two independent normals on about a ninth. Both are right as statements about the population. Neither is what a design of a hundred points delivers, and the ways they fail are worth more than the prediction.
A built column and its partial residual
The coefficient on any column of a regression is its partial residual’s weighted sum of the outcomes, , with the column’s residual after an intercept and every other column. Point carries a share of the coefficient’s information and the effective number of points is over the kurtosis of . A built column is just a column, so all of this holds for it, and what has to be worked out is what its partial residual is.
For the terms a trial or an observational analysis actually fits, the partial residual is simple, because a built column is usually uncorrelated with the columns it was built from:
- the product of two normal covariates , correlated , is uncorrelated with each of them — every term of is an odd moment — so its residual after both is , whose kurtosis works out to : nine for uncorrelated covariates, 12.84 at a correlation of one half, fifteen in the limit;
- the square of a normal covariate, adjusted for the covariate, leaves at any mean , a centred chi-square with one degree of freedom, kurtosis fifteen;
- a treatment’s interaction with a normal covariate, with treatment coded in two equal arms, is uncorrelated with both, and its residual has the kurtosis of itself — three, exactly the covariate’s own;
- a treatment’s interaction with a two-level subgroup, coded as well, is and has kurtosis one: every point carries the same share.
What each kind of coefficient rests on
The ranking the population predicts is the ranking a design of a hundred shows. A treatment’s interaction with a two-level subgroup rests on 94.6 of a hundred points, nearly all of them; with a normal covariate, on 34.9, the same as the covariate’s own coefficient at 34.5. The product of two uncorrelated normals rests on 16.6, and the square of a normal covariate on 12.2. A skewed covariate is where the counts collapse: a treatment’s interaction with a lognormal covariate rests on 10.6 points, and the square of a lognormal covariate on 6.9.
That last pair is where the population and the design part company most. The population share for the treatment-by-lognormal interaction is one over the lognormal’s kurtosis, 113.9 — less than one point in a hundred — and for the square of a lognormal covariate the residual’s kurtosis is about 28.7 million. A design of a hundred cannot show numbers like those, because the kurtosis of a lognormal is set by values so rare that a hundred draws almost never include one. What the design shows instead is its own sample kurtosis, which is always smaller, and a count that is better than the population’s and still bad.
The most informative single point says the same thing in a form a reader can check on a real dataset: in a median design of a hundred it carries 7.5% of a covariate’s own coefficient, 14.9% of the product of two covariates, 19.8% of a square, 21.6% of a treatment’s interaction with a lognormal covariate and 36.1% of a lognormal covariate’s square. A subgroup claim built on a skewed biomarker’s interaction with treatment is, in a typical trial of a hundred, a claim about which way one or two patients went.
Why a design of a hundred rests on more than the population says
The effective count is a ratio of the design’s second moment squared to its fourth, and the fourth moment of a built term is carried by rare points: the square of a normal covariate is large only when the covariate is two or three standard deviations out, and its fourth power only when it is further out still. A design of twenty-five usually has no such point, so its built column looks less peaked than the population’s, and its coefficient rests on 24.1% of its points rather than the population’s 6.7%. At a hundred points the share is 12.2%, at four hundred 8.9%, at sixteen hundred 7.4%. The product converges the same way from 26.9% to 12.0%, and a covariate’s own coefficient, whose kurtosis is three and set by ordinary points, is within a few per cent of its third from the start.
So the prediction was right about the direction and wrong about the size at ordinary sample sizes, and the error is in the reassuring direction. It is also the error that makes the count fragile. A design whose built column happens to contain one of the rare points drops at once towards the population’s count; the median design of a hundred rests on twelve points behind its square, and the one design drawn for the centring figure below rests on 4.88. The spread across designs is the thing to carry: the points a slope rests on found the same for a skewed covariate’s own coefficient, and a built term makes every covariate skewed in the column that counts.
A subgroup split and its continuous version
The two ends of the ranking are the same question asked two ways. A trial that wants to know whether its treatment works differently for patients with a high baseline value can split the covariate at its median and fit the treatment’s interaction with the split, or keep the covariate continuous and fit the interaction with the covariate itself. The split’s interaction rests on nearly every patient, 94.6 of a hundred for a two-level characteristic, because a column has no extreme points; the continuous interaction rests on 34.9.
The count favours the split, and the count is not the whole story. The split throws away where each patient sits within their half, which a baseline cut in two priced at a factor of of the information a covariate carries, and it answers a coarser question — whether the average effect differs between halves, not whether it changes along the covariate. The continuous interaction keeps the information and pays for it by concentrating it: the patients far from the covariate’s mean carry the slope of the treatment effect, exactly as the points far from a covariate’s mean carry a single regression’s slope, and the treatment column changes none of that. Neither is free, and the effective count is what makes the second cost visible. And the choice of where to cut is its own hazard: five places to cut one variable counted how often a chosen cut manufactures a reversal, which is the price of letting the split be chosen after the data.
What a larger design buys a built term
The share of points behind a built term falls as a design grows, and that looks like the wrong direction. The count itself does not fall: behind a normal covariate’s square it is about 6 points at twenty-five, 12.2 at a hundred, 35.6 at four hundred and 118 at sixteen hundred. What falls is the share, because a larger design draws more of the rare values that carry a built column’s fourth moment, and each one it draws takes a larger slice of the coefficient than an ordinary point does.
So a large design is not a design in which every point contributes equally to its squared term; it is one that has found the points that dominate it. That is the regime in which one point can draw a line and two can hide each other, moved from the design’s covariates to the columns the analysis built from them — and a regression diagnostic that looks for high-leverage points on the covariates will not look there.
What centring does and does not change
The standard advice for a model with a squared term is to centre the covariate first, on the grounds that and are highly correlated when is far from zero and the correlation inflates the variances. The advice is right about one coefficient and says nothing about the other.
The columns and span the same space for any , since , and the coefficient on the square is the same in both: the same estimate, the same variance, and the same partial residual, so the same effective points to the last digit. In the drawn design it rests on 4.88 points at every mean — a design that happened to include one of the rare values — and centring does not touch that.
What centring changes is the linear coefficient, because it changes what the linear coefficient is. Uncentred, it is the slope of the fitted curve at ; centred, at the covariate’s mean. When the mean is two standard deviations from zero, the uncentred slope is read where there are almost no data, its variance is inflated by , and it rests on 6.5 points instead of 25.8. Nothing about the fit is better or worse; one parametrisation asks a question the design can answer and the other asks one it cannot. The variance inflation that centring removes was never a property of the data — it was the price of asking for the slope somewhere the data are not.
The point that decides an interaction
The essay this one continues found the point that decides a covariate’s own coefficient hidden from both covariates’ plots in most designs once the covariates are strongly correlated. For an interaction it is the other way round.
A covariate’s own coefficient is decided by the point with the largest partial residual, the direction the other covariate cannot predict, and when the covariates are correlated that direction is diagonal to both axes: a point can be far out along it and ordinary on each axis. At a correlation of 0.9 the deciding point is hidden from both plots in 80.1% of designs of a hundred; at no correlation, never.
The product’s deciding point is the one with the largest . At no correlation the largest products come from points that are moderately large on both covariates at once — two standard deviations on each gives a product of four, as large as a point four standard deviations out on one and at one on the other — and in 14.2% of designs the deciding point is inside both covariates’ central 95%. As the correlation rises the product becomes more like a square, dominated by points extreme on both together, which stand out on both plots: the hidden share falls to 8.5% at a quarter, 2.8% at a half and 0.1% at 0.9. The hero figure’s slider shows the change on one design.
This matters for the claims interactions are fitted to support. One characteristic that really matters hunted for a real modifier among eight two-level characteristics, and a two-level modifier’s interaction rests on nearly every patient, as the counts above show; the difficulty there was multiplicity. An interaction between two continuous covariates is decided, one design in seven, by a patient nobody would pick out from either covariate’s distribution, and the usual check — plot each covariate, look for outliers — passes.
What the counts say about fitting interactions
A treatment’s interaction with a normal covariate is not short of points. It rests on the same third of them as the covariate’s own coefficient, and its difficulty is variance, which is the familiar fourfold cost of estimating a difference of slopes rather than a slope. Nothing about which points decide it is worse.
An interaction of two continuous covariates, or a squared term, rests on half as many points or fewer, 16.6 and 12.2 of a hundred, and on about one in nine or one in fifteen in the population a large design converges to. Its most informative point carries a seventh to a fifth of the coefficient.
A skewed covariate makes any built term’s count unreliable. A treatment’s interaction with a lognormal biomarker rests on 10.6 points in a typical design of a hundred and on less than one in the population; its largest point carries a fifth of the coefficient. The count in any one design depends on whether that design happened to include one of the values that make the population’s kurtosis 114, and a reader cannot tell from the design which case they are in except by computing the count.
The count predicts what one point can do to the fitted coefficient. On a correct model with no interaction at all, deleting the single point that carries the most of a coefficient’s information moves that coefficient by more than its own standard error in 0.1% of designs of a hundred for a covariate’s own coefficient and 0.25% for a treatment’s interaction with a normal covariate — and in 3.65% for the product of two normal covariates, 9.05% for the square of a normal covariate, 13.6% for a treatment’s interaction with a lognormal covariate and 41.9% for the square of a lognormal one. An interaction reported as significant in a design like the last two has a good chance of being a statement about one patient, and the check that finds out — refit without the point that carries the most of it — needs only the design to say which point that is.
And the count is computable before any outcome exists. Every number here is a property of the design matrix alone, so an analysis plan that intends to fit an interaction can state, from the covariates it will have, how many points the interaction will rest on and which patient will carry most of it — the design-stage check a design’s two numbers proposed for leverage, extended to the terms an analysis builds.
The population kurtoses are exact moment calculations, and each for a normal covariate was checked against a design of two hundred thousand rows; the square’s effective count was checked to be the same to seven decimal places centred and uncentred; and the fitted interaction coefficient was checked to equal its partial residual’s weighted sum of the outcomes. The medians are over a thousand seeded designs of each kind. Not measured: interactions of more than two covariates, where the built column is a product of three and its kurtosis larger again; splines and other bases, where a covariate is spread across several built columns whose coefficients are read together; and outcomes that are not linear in the coefficients, where the information a point carries depends on the fitted values as well as the design.
Still open: a curve read through several columns
A squared term is the simplest way to let a relation bend and a poor one; analyses that care about the shape fit a spline, which spreads one covariate across several built columns, each supported on a stretch of the covariate’s range. No single coefficient of a spline is interpreted, so the effective points behind any one of them is the wrong count. The right one is the count behind the fitted curve at a chosen point — the value of the relation at a covariate value of interest — which is a linear combination of several coefficients and has its own weights on the outcomes.
Those weights are computable from the design alone, as every count here was, and they are local: a spline’s value at a point is decided mostly by the data nearby, so the count behind it should depend on how many points fall near the chosen value and how the knots were placed. Whether a spline’s fitted value in the tail of a skewed covariate rests on fewer points than the squared term’s coefficient does, and whether moving a knot can change that count more than adding data does, has not been measured here, and it would say how much of a bent relation’s shape the ends of a design can support.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a chosen probe finds — both name effective sample size, experimental design, leverage, model diagnostics
- A quantity that loses to a heuristic — both name experimental design, leverage, model diagnostics
- A set of pairs, not a vector — both name experimental design, leverage, model diagnostics
- Counting it exactly does not help — both name experimental design, leverage, model diagnostics
- Residuals are not the errors — both name experimental design, leverage, model diagnostics
- The bread and the filling — both name kurtosis, least squares, leverage
Named objects
A flat tag is an object no other essay names yet.
Effective sample sizeExperimental designInteractionKurtosisLeast squaresLeverageModel diagnosticsSubgroup analysisVariance inflation