What a summary of a scatter is a property of

A value read off a curve

No coefficient of a spline is read; what is read is the fitted curve's value at a covariate value of interest, and that value is a weighted sum of the outcomes with weights the design fixes before any outcome is seen. In the tail of a lognormal covariate a natural spline's value rests on 4.8 of a hundred points, fewer than the squared term's coefficient's 6.9, and the squared term's own value there rests on 10.2 — yet it is three times further from the truth, because its extra points are borrowed from where the curve has a different shape and buy no precision at all. At four hundred points, moving the last knot past the value raises its count threefold and halves its error; doubling the data does neither.

Worth reading first: The line that one point drew.

A term built from the others counted the points behind the coefficient on a squared term or an interaction: the number of points divided by the kurtosis of what the built column adds, fifteen for the square of a normal covariate and 6.9 of a hundred for the square of a lognormal one. It ended on the analysis that people who care about a curve’s shape actually run. A spline spreads one covariate across several built columns, each supported on a stretch of its range, and no single coefficient of it means anything. What is read is the curve: the fitted value of the relation at a covariate value of interest, a dose, an age, a concentration.

That value has its own count. It is a linear combination of the coefficients, so it is a linear combination of the outcomes, with weights that the design alone determines. The questions here are how many points those weights rest on, where along the covariate’s range, under a spline and under the squared term it usually replaces — and what decides the count in the tail, where a skewed covariate has few points and the relation is often most interesting.

The weight each of a hundred outcomes carries in the fitted value at the 97.5th percentile of a lognormal covariateOne design of a hundred lognormal covariate values. The fitted value at a covariate value of 7.10 is a weighted sum of the outcomes, with weights fixed by the design. Under a line and its square it rests on 10.7 effective points, the largest carrying 19.1% of its variance. Under a natural spline with knots at 0.17, 0.57, 1.18, 3.95 it rests on 3.8, the largest carrying 41.1%.00.1000.200the point's covariate value (log scale)weight on the point's outcome0.20.512510x₀ = 7.10a line and its square: 10.7 pointsa natural spline: 3.8 pointsone design of 100 lognormal valuesweights from the design alone, before any outcome
Fig. 1 One design of a hundred lognormal covariate values: the weight each point’s outcome carries in the fitted value at the 97.5th percentile of the covariate, under a line and its square and under a natural spline with four knots, against the point’s covariate value on a logarithmic scale. The slider moves the value to the median and the 90th percentile.

A fitted value is a weighted sum

Any least-squares fit with design matrix XX predicts the relation at x0x_0 as f^(x0)=b(x0)⊤β^\hat f(x_0) = b(x_0)^\top \hat\beta, where b(x0)b(x_0) is the row the design would have at x0x_0. Since β^=(X⊤X)−1X⊤y\hat\beta = (X^\top X)^{-1} X^\top y,

f^(x0)=∑iwi yi,w=X(X⊤X)−1b(x0),\hat f(x_0) = \sum_i w_i\, y_i, \qquad w = X (X^\top X)^{-1} b(x_0),

and the weights wiw_i depend on the covariate values and the model and on nothing else. Point ii contributes σ2wi2\sigma^2 w_i^2 to the value’s variance. On the rule the points a slope rests on introduced, it carries a share wi2/∑w2w_i^2/\sum w^2 of the value’s information, and the effective number of points the value rests on is

neff=(∑wi2)2∑wi4,n_{\text{eff}} = \frac{\left(\sum w_i^2\right)^2}{\sum w_i^4},

the number of points that would carry it if each carried an equal share. A second number sits beside it: 1/∑wi21/\sum w_i^2, the number of points whose plain average would be as precise. The first measures how much one point can move the value, the second how noisy it is, and in the tail of a skewed design they come apart.

None of this needs an outcome. Like the leverages that two numbers for the fit’s geometry read off a design before any measurement, the weights are a property of where the covariate values fell and of which model was chosen, so the count behind a value can be known — and printed — at the moment the design is fixed.

Three models are compared on the same designs: a straight line; a line and its square, which is what a “check for curvature” usually adds; and a natural cubic spline with knots at the 5th, 35th, 65th and 95th percentiles of the design’s own covariate values, cubic between its knots and straight beyond the outer two. Every model contains the straight line, so in every one the weights sum to one and reproduce x0x_0 exactly; what differs is how they are spread.

The hero figure shows one design. At the median of the covariate every model spreads its weight widely: the square’s value rests on 95.3 points and the spline’s on 47.5. At the 97.5th percentile, x0=7.10x_0 = 7.10 on a covariate whose median is one, the spline’s weights collapse onto the handful of points out there — 3.8 effective points, the largest carrying 41.1% of the variance — while the square’s stay spread across the design at 10.7, with points far below x0x_0 carrying weight of both signs.

Where along the range a value rests on its points

One design is an anecdote, so the count is taken over four hundred designs of a hundred, at six places along the covariate’s range.

How many of a hundred points a fitted value rests on, by where it is read, lognormal covariate. Median over 400 designs of a hundred lognormal covariate values. A straight line: 88.2, 85.1, 13.7, 8.7, 7.5, 7.3. A line and its square: 86.5, 44.3, 27.5, 20.8, 10.2, 3.3. A natural spline, four knots: 51.3, 36.1, 28.4, 17.7, 4.8, 3.4 — at the 50th, 75th, 90th, 95th, 97.5th, 99th percentiles. At the 97.5th the spline's value rests on 4.8 points with the largest carrying 35.2% of its variance; the square's on 10.2.
Fig. 2 The median effective number of points behind each model’s fitted value, over four hundred designs of a hundred lognormal covariate values, against where the value is read: the 50th to the 99th percentile of the covariate.

On a lognormal covariate, the spline’s value at the median rests on 51.3 points, at the 75th percentile on 36.1, at the 90th on 28.4, at the 95th on 17.7, at the 97.5th on 4.8 and at the 99th on 3.4. The fall is steep between the 95th and the 97.5th percentiles because that is where the last knot sits: inside it, the spline is a cubic fitted to the points nearby; beyond it, a straight line whose slope is decided by the few points at the edge of the data. At the 97.5th percentile the largest single point carries 35.2% of the value’s variance.

So the answer to the first question is yes. The spline’s value in the tail of a skewed covariate rests on fewer points than the squared term’s coefficient did — 4.8 against 6.9 — and the coefficient was already the weakest count the earlier essay found for any built term. A curve’s value in the tail of a skewed design is, in a typical study of a hundred, a statement about four or five people.

The line and its square spread differently. The square’s value rests on 86.5 points at the median, 27.5 at the 90th percentile, 20.8 at the 95th and 10.2 at the 97.5th, more than twice the spline’s count there. Only at the 99th percentile, beyond nearly every observation, do the two meet, at 3.3 and 3.4: there every model is extrapolating and every model’s value is decided by whichever points lie furthest out. The straight line is the odd one: it rests on 88.2 points at the median and 7.5 at the 97.5th percentile, because a line’s value far out is governed by its slope, and a slope on a lognormal design rests on its few largest points — the situation the line that one point drew built by hand, here arising from nothing more unusual than a skewed covariate.

On a normal covariate the shapes are milder and the ranking is the same. The spline’s value at the 97.5th percentile rests on 7.8 points, the square’s on 10.8 and the line’s on 35.6. A normal design has its tail points spread symmetrically and fewer of them extreme, so every count is higher; the spline is still the one that concentrates its weight where the value is read.

What the borrowed points are worth

A higher count looks like a better value, and the earlier essays in this series treated it as one, because for a coefficient with nothing else to trade against, more points behind it is simply more evidence. A fitted value under a wrong model has something to trade against: its bias.

What each fitted value's points are worth: its error about log x, lognormal covariate. Root mean squared error of the fitted value about the curve log x, with noise of standard deviation 0.5, over 400 designs of a hundred — exact given each design, as squared bias plus variance. A straight line: 0.278, 0.531, 0.501, 0.477, 0.819, 1.789. A line and its square: 0.285, 0.242, 0.274, 0.497, 0.807, 2.344. A natural spline, four knots: 0.104, 0.088, 0.127, 0.176, 0.257, 0.468. At the 97.5th percentile the square's value rests on 10.2 points and errs by 0.807; the spline's rests on 4.8 and errs by 0.257.
Fig. 3 Root mean squared error of each model’s fitted value about the curve log⁡x\log x, which none of the three models contains, with noise of standard deviation 0.5, over four hundred designs of a hundred lognormal covariate values. The error is exact given each design: the squared bias of the weighted sum plus its variance. Logarithmic scale.

Take the relation to be log⁡x\log x — concave, flattening in the tail, the shape a dose-response or an income effect often has — with noise of standard deviation one half. At the 97.5th percentile the square’s value rests on 10.2 points and errs by 0.807 on average; the spline’s rests on 4.8 points and errs by 0.257. Across the range the spline’s error is 0.104 at the median, 0.088 at the 75th percentile, 0.127 at the 90th and 0.176 at the 95th; the square’s is 0.285, 0.242, 0.274 and 0.497. At the 99th percentile the square misses by 2.344, nearly five times the noise’s standard deviation, and the spline by 0.468.

This is four datasets with one summary in a quieter form. There, four very different relations produced the same line, and the line said nothing about which one it had been fitted to. Here, a parabola fitted to a concave relation fits the middle of the data well enough that its residual plot gives little away, and its value in the tail is still wrong by more than the noise.

The square’s extra points are the reason for its error rather than a protection against it. A parabola fitted to a lognormal design is decided mostly by the bulk of the data below 3, and its value at 7 is the parabola’s continuation, not anything the points near 7 said. The points far below x0x_0 carrying weight of both signs in the hero figure are that continuation, visible: the value at x0x_0 leans on the curvature measured in the middle of the range, and when the curve’s shape in the tail differs from its shape in the middle, every one of those borrowed points pulls the value the wrong way.

And they do not even buy precision. At the 97.5th percentile the square’s value is as precise as the average of 8.9 points and the spline’s as the average of 9.6. The spread-out weights that make the square’s count twice the spline’s include large weights of opposite sign that cancel in the mean and add in the variance. The count says the square’s value depends on more of the data; the precision says it is no less noisy for that; the error says the dependence was on the wrong data. Reading the count alone, as a measure of how much evidence stands behind a value, would have got the comparison backwards.

On the normal covariate with sin⁡x\sin x as the relation the same ordering holds at the 97.5th percentile: the spline errs by 0.218, the straight line by 0.323 and the square by 0.416. Wherever the relation’s shape in the tail is not the shape in the middle, a fitted value in the tail should rest on the tail’s points, and a spline is the model that makes it do so.

What a knot moves

The spline’s count fell off a cliff between the 95th and 97.5th percentiles because the last knot was at the 95th. Beyond it, the natural spline is a straight line, and a straight line’s value far out rests on the points with the largest leverage, which on a lognormal design are very few. That suggests the count at a tail value is set as much by where the last knot is as by how many points there are, and it can be measured directly.

What moving the last knot does to the points behind a tail value, against what adding points does. Median effective points behind the natural spline's fitted value at the 97.5th percentile of a lognormal covariate, 7.10, over 400 designs at each size, with the first three knots at the 5th, 35th and 65th percentiles and the last at the 80th, 90th, 95th, 99th. Last knot at the 80th: 5.1, 7.9, 12.8, 20.0. Last knot at the 90th: 5.0, 7.7, 12.7, 21.4. Last knot at the 95th: 5.0, 7.9, 13.7, 26.3. Last knot at the 99th: 7.3, 15.1, 40.9, 103.2 — at 100, 200, 400, 800 points.
Fig. 4 The median effective number of points behind the natural spline’s value at the 97.5th percentile of a lognormal covariate, against the number of points in the design, with the last knot at the 80th, 90th, 95th and 99th percentile of the design’s covariate values and the other three fixed. Both axes logarithmic.

With the last knot at the 95th percentile the count at the 97.5th percentile is 5.0 at a hundred points, 7.9 at two hundred, 13.7 at four hundred and 26.3 at eight hundred: doubling the design a little less than doubles the count, as it should, since the value is extrapolated from a stretch of the range whose share of the points is fixed. With the last knot moved to the 99th percentile — past the value, so that the value is read inside the cubic rather than off the straight tail — the count is 7.3, 15.1, 40.9 and 103.2.

At four hundred points, moving the knot from the 95th to the 99th percentile raises the count from 13.7 to 40.9, threefold. Doubling the design to eight hundred with the knot where it was raises it to 26.3. The knot moves the count more than the data does. And it moves the error more: at four hundred points the value’s error about log⁡x\log x falls from 0.187 to 0.103 when the knot moves, and is 0.198 at eight hundred points with the knot left at the 95th percentile — no better at all, because the error there is the straight tail’s bias, which no amount of data reduces.

At a hundred points the knot can do little. The 99th percentile of a hundred values is the second largest of them, there is nothing between the 95th and the 99th to fit a cubic to, and moving the knot raises the count only from 5.0 to 7.3 and lowers the error from 0.260 to 0.231. A knot is worth moving only where there are points to move it among, which means a design large enough to have a tail.

Moving the last knot inwards does the opposite of what intuition might expect. At the 80th or 90th percentile the value at the 97.5th is deeper into the straight tail, the count barely changes — 12.8 and 12.7 at four hundred points — and the error rises to 0.256 and 0.232, because the line extrapolated from further in misses more of the curve’s bend. A knot placed inside the range where the value is read is a decision that the relation is a straight line out there, and the value inherits that decision’s error however many points are behind it.

The same arithmetic explains a familiar complaint about spline fits, that their ends are wild. The end of a natural spline is a straight line fitted to the last few points, and a straight line fitted to few points with large leverage swings with each of them. The wildness is the count: a value resting on four or five points moves when one of them does, and a robust loss does not help when the points carrying the value are the ones whose covariate values are far out rather than whose outcomes are.

Why the count of a coefficient was the wrong count

The counts in the essays before this one were all counts behind a coefficient: a slope, a square, an interaction. A summary of the whole fit that turned out to belong to the design made the same distinction for a summary; here it is made for one number read off it. A coefficient is a global quantity — every point’s outcome enters it with a weight set by that point’s partial residual — and its count is nn over a kurtosis because the partial residual’s fourth moment is what concentrates its information. For a curve’s value at a point, the relevant weights are not the partial residual of any column. They are the row of the fit’s projection that belongs to x0x_0, and they are concentrated near x0x_0 to the degree the model allows its shape to vary along the range.

That is why the two counts disagree in direction as well as size. The squared term’s coefficient rests on 6.9 points of a lognormal hundred because the square’s partial residual is dominated by a few large values; the square’s fitted value at the 97.5th percentile rests on 10.2 because it borrows the middle of the range to fix the curvature there. Neither is the count behind “what is the relation at this covariate value”. The spline’s value at the same place rests on 4.8, and that is the honest number: the relation in the tail of a skewed design is known from the tail’s points and from nothing else, unless something outside the data says the shape continues.

What a reader can do with the weights

Report the count behind a curve’s value at the covariate values that matter, not the counts behind its coefficients. The weights need only the design and the model, so the count can be printed before any outcome is measured, and a value read in a tail resting on four or five points should be presented as one.

Do not take a higher count as a better value. Under a squared term a tail value rests on twice the points a spline’s does and errs three times as much about a concave curve, with no gain in precision. More points behind a value is more evidence only if the model is right about the shape between those points and the value.

Put the last knot past the values that will be read, when the design has points out there to support it. At four hundred lognormal points that triples the count behind the 97.5th-percentile value and halves its error, which doubling the data does not do. At a hundred points it barely matters, because there is no tail to place a knot in.

Every count and error here is exact given each design — the weights are computed from the design by solving the model’s normal equations at x0x_0, and the error is the weighted sum’s squared bias about the stated curve plus its variance — and medians and averages are over four hundred designs at each size. Every model’s weights are checked to sum to one and to reproduce a straight line exactly, and the spline’s weights are checked to draw less of the value’s variance from points far from it than the square’s do. The reading that a value resting on more points is the better value is refused: in the lognormal tail the square’s value rests on more than twice the spline’s points and has more than twice its error.

Still open: a knot chosen from the data

Every knot here was placed at a stated percentile of the covariate, which uses the design and nothing else, and so leaves every count and every error exact given the design. In practice the number of knots and their places are often chosen by a criterion computed from the outcomes — a cross-validated error, an information criterion, a penalty whose size is tuned on the fit. The figure on knots says that choice moves a tail value’s count by a factor of three and its error by a factor of two, which is exactly the kind of choice a cut that fitted best found flatters the precision it reports.

A penalised spline, which keeps many knots and shrinks the curvature between them by an amount chosen from the data, makes the same choice continuously. Its fitted value is still a weighted sum of the outcomes, but the weights now depend on the outcomes through the tuning, and the count behind a tail value is a random quantity rather than a property of the design. How far the count reported at the chosen penalty overstates the count the value really rests on, across the designs of this essay, has not been measured.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Bias-varianceCurvatureEffective sample sizeExtrapolationKurtosisLeast squaresLeverageModel diagnosticsSpline