What a summary of a scatter is a property of

A term built from the others

A regression's squared terms and interactions are columns built from its covariates, and each coefficient on one rests on the number of points divided by the kurtosis of what the built column adds — fifteen for the square of a normal covariate, nine for the product of two uncorrelated ones, three for a treatment's interaction with a normal covariate, which is the same as the covariate's own coefficient. A design of a hundred rests on more than the population says, 12.2 points behind a square against 6.7, because a hundred draws rarely hold the tail that sets the kurtosis. Centring changes nothing about the square's coefficient, to the last digit. And correlation between two covariates hides the point that decides each one's own coefficient while it reveals the point that decides their interaction.

Worth reading first: The line that one point drew.

A point ordinary on every axis found that each coefficient of a multiple regression rests on its own number of points: the number of points divided by the kurtosis of that covariate’s partial residual, the part of it the other covariates cannot predict. For normal covariates that is about a third of the points whatever their correlation, and the essay closed on the coefficients it had not counted. Many regressions carry terms built from their covariates — a square to let a relation bend, a product to let one covariate’s effect depend on another — and a built term is never normal, even when everything it is built from is.

The identity does not care how a column was made. So the prediction was that a squared normal covariate’s coefficient rests on about a fifteenth of the points and the product of two independent normals on about a ninth. Both are right as statements about the population. Neither is what a design of a hundred points delivers, and the ways they fail are worth more than the prediction.

A point that decides an interaction, on two covariates correlated at 0A hundred points on two normal covariates with correlation 0. Dashed lines mark where the most extreme five per cent of each covariate begin. The point carrying the largest share of the interaction's information, 10.2%, is inside the lines on both axes, so neither covariate's plot on its own would show it. The coefficient rests on 19.0 effective points.-202-202first covariatesecond covariate10.2% of the interactionone seeded design of 100, correlation 0the deciding point shows on neither margin
Fig. 1 A hundred points on two normal covariates. The marked point carries the largest share of the information behind the coefficient on their product in a fit that also has each covariate; the dashed lines mark where each covariate’s most extreme five per cent begin. At no correlation it sits inside both pairs of lines. The slider sets the correlation between the covariates.

A built column and its partial residual

The coefficient on any column of a regression is its partial residual’s weighted sum of the outcomes, β^j=∑riyi/∑ri2\hat\beta_j = \sum r_i y_i/\sum r_i^2, with rr the column’s residual after an intercept and every other column. Point ii carries a share ri2/∑r2r_i^2/\sum r^2 of the coefficient’s information and the effective number of points is nn over the kurtosis of rr. A built column is just a column, so all of this holds for it, and what has to be worked out is what its partial residual is.

For the terms a trial or an observational analysis actually fits, the partial residual is simple, because a built column is usually uncorrelated with the columns it was built from:

  • the product of two normal covariates x1x2x_1x_2, correlated ρ\rho, is uncorrelated with each of them — every term of E[x12x2]E[x_1^2 x_2] is an odd moment — so its residual after both is x1x2−ρx_1x_2 - \rho, whose kurtosis works out to (9+42ρ2+9ρ4)/(1+ρ2)2(9 + 42\rho^2 + 9\rho^4)/(1 + \rho^2)^2: nine for uncorrelated covariates, 12.84 at a correlation of one half, fifteen in the limit;
  • the square of a normal covariate, adjusted for the covariate, leaves (x−μ)2−1(x - \mu)^2 - 1 at any mean μ\mu, a centred chi-square with one degree of freedom, kurtosis fifteen;
  • a treatment’s interaction with a normal covariate, with treatment coded ±1\pm1 in two equal arms, is uncorrelated with both, and its residual t⋅xt\cdot x has the kurtosis of xx itself — three, exactly the covariate’s own;
  • a treatment’s interaction with a two-level subgroup, coded ±1\pm1 as well, is ±1\pm1 and has kurtosis one: every point carries the same share.

What each kind of coefficient rests on

How many of a hundred points each kind of coefficient rests on, for terms built from the covariates. Median effective number of points behind each coefficient over a thousand designs of a hundred, with the population share in brackets: treatment × a two-level subgroup, 94.6 (100.0); a normal covariate's own coefficient, 34.5 (33.3); treatment × a normal covariate, 34.9 (33.3); product of two normals, uncorrelated, 16.6 (11.1); product of two normals, correlated 0.9, 12.2 (6.7); square of a normal covariate, 12.2 (6.7); treatment × a lognormal covariate, 10.6 (0.88); square of a lognormal covariate, 6.9 (0.00).
Fig. 2 How many of a hundred points each kind of coefficient rests on: the median over a thousand designs as a bar, and the population share, one over the built column’s residual kurtosis, as a tick. Grey bars are the two references, a two-level subgroup and a covariate’s own coefficient.

The ranking the population predicts is the ranking a design of a hundred shows. A treatment’s interaction with a two-level subgroup rests on 94.6 of a hundred points, nearly all of them; with a normal covariate, on 34.9, the same as the covariate’s own coefficient at 34.5. The product of two uncorrelated normals rests on 16.6, and the square of a normal covariate on 12.2. A skewed covariate is where the counts collapse: a treatment’s interaction with a lognormal covariate rests on 10.6 points, and the square of a lognormal covariate on 6.9.

That last pair is where the population and the design part company most. The population share for the treatment-by-lognormal interaction is one over the lognormal’s kurtosis, 113.9 — less than one point in a hundred — and for the square of a lognormal covariate the residual’s kurtosis is about 28.7 million. A design of a hundred cannot show numbers like those, because the kurtosis of a lognormal is set by values so rare that a hundred draws almost never include one. What the design shows instead is its own sample kurtosis, which is always smaller, and a count that is better than the population’s and still bad.

The most informative single point says the same thing in a form a reader can check on a real dataset: in a median design of a hundred it carries 7.5% of a covariate’s own coefficient, 14.9% of the product of two covariates, 19.8% of a square, 21.6% of a treatment’s interaction with a lognormal covariate and 36.1% of a lognormal covariate’s square. A subgroup claim built on a skewed biomarker’s interaction with treatment is, in a typical trial of a hundred, a claim about which way one or two patients went.

Why a design of a hundred rests on more than the population says

The share of a design's points a built term's coefficient rests on, against the size of the design. Median effective share: a covariate's own coefficient 0.381, 0.357, 0.345, 0.342, 0.337, 0.334; the product of two uncorrelated normals 0.269, 0.203, 0.166, 0.143, 0.131, 0.120; the square of a normal 0.241, 0.162, 0.122, 0.100, 0.089, 0.074, at 25, 50, 100, 200, 400, 1600 points. The population shares are one third, one ninth and one fifteenth.
Fig. 3 The median share of a design’s points behind a covariate’s own coefficient, behind the product of two uncorrelated normal covariates, and behind the square of a normal covariate, against the number of points in the design, with each population share dashed.

The effective count is a ratio of the design’s second moment squared to its fourth, and the fourth moment of a built term is carried by rare points: the square of a normal covariate is large only when the covariate is two or three standard deviations out, and its fourth power only when it is further out still. A design of twenty-five usually has no such point, so its built column looks less peaked than the population’s, and its coefficient rests on 24.1% of its points rather than the population’s 6.7%. At a hundred points the share is 12.2%, at four hundred 8.9%, at sixteen hundred 7.4%. The product converges the same way from 26.9% to 12.0%, and a covariate’s own coefficient, whose kurtosis is three and set by ordinary points, is within a few per cent of its third from the start.

So the prediction was right about the direction and wrong about the size at ordinary sample sizes, and the error is in the reassuring direction. It is also the error that makes the count fragile. A design whose built column happens to contain one of the rare points drops at once towards the population’s count; the median design of a hundred rests on twelve points behind its square, and the one design drawn for the centring figure below rests on 4.88. The spread across designs is the thing to carry: the points a slope rests on found the same for a skewed covariate’s own coefficient, and a built term makes every covariate skewed in the column that counts.

A subgroup split and its continuous version

The two ends of the ranking are the same question asked two ways. A trial that wants to know whether its treatment works differently for patients with a high baseline value can split the covariate at its median and fit the treatment’s interaction with the split, or keep the covariate continuous and fit the interaction with the covariate itself. The split’s interaction rests on nearly every patient, 94.6 of a hundred for a two-level characteristic, because a ±1\pm1 column has no extreme points; the continuous interaction rests on 34.9.

The count favours the split, and the count is not the whole story. The split throws away where each patient sits within their half, which a baseline cut in two priced at a factor of 2/π2/\pi of the information a covariate carries, and it answers a coarser question — whether the average effect differs between halves, not whether it changes along the covariate. The continuous interaction keeps the information and pays for it by concentrating it: the patients far from the covariate’s mean carry the slope of the treatment effect, exactly as the points far from a covariate’s mean carry a single regression’s slope, and the treatment column changes none of that. Neither is free, and the effective count is what makes the second cost visible. And the choice of where to cut is its own hazard: five places to cut one variable counted how often a chosen cut manufactures a reversal, which is the price of letting the split be chosen after the data.

What a larger design buys a built term

The share of points behind a built term falls as a design grows, and that looks like the wrong direction. The count itself does not fall: behind a normal covariate’s square it is about 6 points at twenty-five, 12.2 at a hundred, 35.6 at four hundred and 118 at sixteen hundred. What falls is the share, because a larger design draws more of the rare values that carry a built column’s fourth moment, and each one it draws takes a larger slice of the coefficient than an ordinary point does.

So a large design is not a design in which every point contributes equally to its squared term; it is one that has found the points that dominate it. That is the regime in which one point can draw a line and two can hide each other, moved from the design’s covariates to the columns the analysis built from them — and a regression diagnostic that looks for high-leverage points on the covariates will not look there.

What centring does and does not change

The standard advice for a model with a squared term is to centre the covariate first, on the grounds that xx and x2x^2 are highly correlated when xx is far from zero and the correlation inflates the variances. The advice is right about one coefficient and says nothing about the other.

What centring a covariate changes in a model with its square: effective points behind each coefficient, against the covariate's mean. One design of a hundred normal points shifted to mean μ. The square's coefficient rests on 4.88 points at every μ, centred or not. The linear coefficient rests on 25.8 centred at every μ, and uncentred on 23.9, 9.7, 7.5, 6.8, 6.5, 6.2, 6.0 at μ = 0, 0.5, 1, 1.5, 2, 2.5, 3, where its variance inflation, one plus twice the square of μ, is 1, 1.5, 3, 5.5, 9, 13.5, 19.
Fig. 4 One design of a hundred normal points shifted to have mean μ, fitted with the covariate and its square: the effective points behind the linear coefficient with the covariate centred and uncentred, and behind the square’s coefficient, which is identical either way.

The columns [1,x,x2][1, x, x^2] and [1,x−c,(x−c)2][1, x - c, (x - c)^2] span the same space for any cc, since (x−c)2=x2−2cx+c2(x-c)^2 = x^2 - 2cx + c^2, and the coefficient on the square is the same in both: the same estimate, the same variance, and the same partial residual, so the same effective points to the last digit. In the drawn design it rests on 4.88 points at every mean — a design that happened to include one of the rare values — and centring does not touch that.

What centring changes is the linear coefficient, because it changes what the linear coefficient is. Uncentred, it is the slope of the fitted curve at x=0x = 0; centred, at the covariate’s mean. When the mean is two standard deviations from zero, the uncentred slope is read where there are almost no data, its variance is inflated by 1+2μ2=91 + 2\mu^2 = 9, and it rests on 6.5 points instead of 25.8. Nothing about the fit is better or worse; one parametrisation asks a question the design can answer and the other asks one it cannot. The variance inflation that centring removes was never a property of the data — it was the price of asking for the slope somewhere the data are not.

The point that decides an interaction

The essay this one continues found the point that decides a covariate’s own coefficient hidden from both covariates’ plots in most designs once the covariates are strongly correlated. For an interaction it is the other way round.

How often the point that decides a coefficient stands out on neither covariate's plot, for a main effect and for an interaction. Over a thousand designs of a hundred on two normal covariates, the point carrying the most of a covariate's own coefficient is inside both covariates' central 95% in 0.0%, 2.8%, 21.6%, 55.8%, 80.1% of designs at correlations 0, 0.25, 0.5, 0.75, 0.9; the point carrying the most of their product's coefficient, in 14.2%, 8.5%, 2.8%, 0.5%, 0.1%.
Fig. 5 How often the point carrying the most of a coefficient’s information stands out on neither covariate’s plot — lies inside both covariates’ central 95% — against the correlation between the two covariates, for a covariate’s own coefficient and for their product’s.

A covariate’s own coefficient is decided by the point with the largest partial residual, the direction the other covariate cannot predict, and when the covariates are correlated that direction is diagonal to both axes: a point can be far out along it and ordinary on each axis. At a correlation of 0.9 the deciding point is hidden from both plots in 80.1% of designs of a hundred; at no correlation, never.

The product’s deciding point is the one with the largest ∣x1x2−ρ∣|x_1x_2 - \rho|. At no correlation the largest products come from points that are moderately large on both covariates at once — two standard deviations on each gives a product of four, as large as a point four standard deviations out on one and at one on the other — and in 14.2% of designs the deciding point is inside both covariates’ central 95%. As the correlation rises the product becomes more like a square, dominated by points extreme on both together, which stand out on both plots: the hidden share falls to 8.5% at a quarter, 2.8% at a half and 0.1% at 0.9. The hero figure’s slider shows the change on one design.

This matters for the claims interactions are fitted to support. One characteristic that really matters hunted for a real modifier among eight two-level characteristics, and a two-level modifier’s interaction rests on nearly every patient, as the counts above show; the difficulty there was multiplicity. An interaction between two continuous covariates is decided, one design in seven, by a patient nobody would pick out from either covariate’s distribution, and the usual check — plot each covariate, look for outliers — passes.

What the counts say about fitting interactions

A treatment’s interaction with a normal covariate is not short of points. It rests on the same third of them as the covariate’s own coefficient, and its difficulty is variance, which is the familiar fourfold cost of estimating a difference of slopes rather than a slope. Nothing about which points decide it is worse.

An interaction of two continuous covariates, or a squared term, rests on half as many points or fewer, 16.6 and 12.2 of a hundred, and on about one in nine or one in fifteen in the population a large design converges to. Its most informative point carries a seventh to a fifth of the coefficient.

A skewed covariate makes any built term’s count unreliable. A treatment’s interaction with a lognormal biomarker rests on 10.6 points in a typical design of a hundred and on less than one in the population; its largest point carries a fifth of the coefficient. The count in any one design depends on whether that design happened to include one of the values that make the population’s kurtosis 114, and a reader cannot tell from the design which case they are in except by computing the count.

The count predicts what one point can do to the fitted coefficient. On a correct model with no interaction at all, deleting the single point that carries the most of a coefficient’s information moves that coefficient by more than its own standard error in 0.1% of designs of a hundred for a covariate’s own coefficient and 0.25% for a treatment’s interaction with a normal covariate — and in 3.65% for the product of two normal covariates, 9.05% for the square of a normal covariate, 13.6% for a treatment’s interaction with a lognormal covariate and 41.9% for the square of a lognormal one. An interaction reported as significant in a design like the last two has a good chance of being a statement about one patient, and the check that finds out — refit without the point that carries the most of it — needs only the design to say which point that is.

And the count is computable before any outcome exists. Every number here is a property of the design matrix alone, so an analysis plan that intends to fit an interaction can state, from the covariates it will have, how many points the interaction will rest on and which patient will carry most of it — the design-stage check a design’s two numbers proposed for leverage, extended to the terms an analysis builds.

The population kurtoses are exact moment calculations, and each for a normal covariate was checked against a design of two hundred thousand rows; the square’s effective count was checked to be the same to seven decimal places centred and uncentred; and the fitted interaction coefficient was checked to equal its partial residual’s weighted sum of the outcomes. The medians are over a thousand seeded designs of each kind. Not measured: interactions of more than two covariates, where the built column is a product of three and its kurtosis larger again; splines and other bases, where a covariate is spread across several built columns whose coefficients are read together; and outcomes that are not linear in the coefficients, where the information a point carries depends on the fitted values as well as the design.

Still open: a curve read through several columns

A squared term is the simplest way to let a relation bend and a poor one; analyses that care about the shape fit a spline, which spreads one covariate across several built columns, each supported on a stretch of the covariate’s range. No single coefficient of a spline is interpreted, so the effective points behind any one of them is the wrong count. The right one is the count behind the fitted curve at a chosen point — the value of the relation at a covariate value of interest — which is a linear combination of several coefficients and has its own weights on the outcomes.

Those weights are computable from the design alone, as every count here was, and they are local: a spline’s value at a point is decided mostly by the data nearby, so the count behind it should depend on how many points fall near the chosen value and how the knots were placed. Whether a spline’s fitted value in the tail of a skewed covariate rests on fewer points than the squared term’s coefficient does, and whether moving a knot can change that count more than adding data does, has not been measured here, and it would say how much of a bent relation’s shape the ends of a design can support.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Effective sample sizeExperimental designInteractionKurtosisLeast squaresLeverageModel diagnosticsSubgroup analysisVariance inflation