A standard error for a model that is wrong

What a wrong model estimates

A straight line fitted to a curved truth converges on the tangent at its own design's mean. Two honest studies of one world, fitting the same wrong model, report 2.600000 and 1.600000, and neither is in error.

Worth reading first: Four datasets, one summary.

Two groups measure the same world. The response really is 1+0.6x+0.5x21 + 0.6x + 0.5x^2; both groups fit a straight line, because a straight line is what the software offers and because nobody has told either of them about the quadratic term. One group’s covariate is spread evenly over the interval from nought to two. The other’s is spread exponentially, with exactly the same mean. The first group’s slope converges on 1.600000 and the second’s on 2.600000, and no amount of data moves either of them towards the other.

Neither group has made a mistake. Neither has a smaller sample, a worse instrument or a sloppier analyst. They have fitted the same wrong model to the same truth, and the number a wrong model converges to turns out to be a property of the covariate distribution as much as of the world. That number — the population projection — is what every standard error in this field is an interval for, and it is not the truth’s own parameter.

That is worth stating flatly, because the usual account of misspecification is about bias: the estimate is off by an amount, and with enough care the amount can be corrected. Here there is no amount. The two slopes are a whole unit apart, they are both consistent for something, and the two somethings are different. A correction would have to know which design’s answer is wanted, and nothing in either study says.

The number a straight line converges to

Least squares against a population, rather than against a sample, minimises E[(yb0b1x)2]E[(y - b_0 - b_1 x)^2], and its solution is the familiar ratio of a covariance to a variance. With the truth a0+a1x+curvx2a_0 + a_1 x + \mathrm{curv}\cdot x^2 and the covariate’s mean mm, variance vv and third central moment μ3\mu_3, the covariance carries the quadratic term through Cov(x,x2)=2mv+μ3\mathrm{Cov}(x, x^2) = 2mv + \mu_3, and the slope the fit converges to is

β1=a1+curv(2m+μ3v).\beta_1^* = a_1 + \mathrm{curv}\left(2m + \frac{\mu_3}{v}\right).

Read that against a symmetric covariate, where μ3\mu_3 vanishes, and it says something a reader can check by differentiating: the limit is a1+2curvma_1 + 2\,\mathrm{curv}\cdot m, which is the truth’s own tangent slope at x=mx = m. A straight line fitted to a curve converges on the tangent at the design’s own mean, and on nothing about the design except its mean.

So the arithmetic already contains the finding, and the finding is about which moments appear. The spread does not appear at all; the mean appears doubled; the skewness appears divided by the variance. Every one of those is a property of where the covariate was measured rather than of what is true, and the truth enters only through a1a_1 and curv\mathrm{curv}, which both studies share.

One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 1.6000 and the same spread moved to [1, 3] gives 2.6000, while widening it to [0, 4] gives 2.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 2.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 1.0000 apart, and neither is making an error.
Fig. 1 The slope a straight line converges to under four covariate distributions, in closed form and counted over 2,500 fits of two hundred rows apiece. The even spread over the unit interval doubled gives 1.6000 and the exponential spread with the same mean gives 2.6000.

Four designs, and the pair that is a whole unit apart

Four covariate distributions, one truth, and two independent routes to each number. The closed form above reads the design’s central moments and touches no data. The count fits 2,500 samples of two hundred rows apiece and averages the slopes. The two agree on the three symmetric designs at every digit the sweep can resolve — 1.6017 against 1.6000 with a standard error of 0.0025, 2.6017 against 2.6000 at the same precision, 2.6006 against 2.6000 at 0.0015 — which is a t of 0.70, 0.70 and 0.40.

The pair that matters is the first and the last. An even spread over the doubled unit interval and an exponential spread both have mean 1.0000, and their targets are 1.600000 and 2.600000. The exponential design’s variance is 1.0000 against the even one’s 0.3333 and its standardised skewness is 2.000 against 0.000, and it is the skewness that does the work: μ3/v\mu_3/v is 2.0000 for the exponential and nought for every symmetric design, so the exponential design’s line is chasing the curve out into a tail the even design never visits.

A reader who has met the four datasets that share every summary statistic will recognise the shape of the complaint and should notice that this one is worse. There the summaries were identical and the data differed; here the data-generating truth is identical and the summaries differ, because the summaries are about a model that is not it. The same reversal is the subject of the essay that treats Simpson’s reversal as a region rather than a table: correct arithmetic, honestly applied, arriving at two answers that cannot both be quoted.

Widening a design changes nothing, and that is the proof

The strongest evidence that the target belongs to the design rather than to the truth is a number that does not move.

Take the even spread over the unit interval doubled, and shift it up by one: the target goes from 1.600000 to 2.600000, moving by twice the shift, exactly as a1+2curvma_1 + 2\,\mathrm{curv}\cdot m says it must. Now take the same spread and widen it, from the interval ending at two to the interval ending at four. The covariate’s variance goes from 0.3333 to 1.3333 — four times as much spread, twice the range, a design a practitioner would describe as much better — and the target is 2.600000, the same number as the shifted design’s, to every digit. The sweep reports the difference between them as 0 exactly rather than as a small number.

That is the whole argument in one comparison. If the fitted slope were converging on some feature of the truth, a design that covers four times as much of the truth would converge on something different from one that covers a quarter of it. It does not. If it were converging on an average of the truth’s local slopes, the wider design’s average would differ from the narrower one’s. It does not. The one thing the two designs share is their mean, and the one thing the targets share is the tangent slope at it.

What a straight line through a curve converges to. A quadratic truth over an even spread over [0, 2], and the straight line least squares converges to on it. The line is not a fit to these points: it is the population projection, computable before any data exists, with slope 1.6000 and intercept 0.6667. It is above the truth in the middle of the design and below it at both ends, by 0.3234 at the low end and 0.3234 at the high, and its average against the covariate is exactly nought — which is what a projection is. The misfit is a systematic error the fit cannot remove, and it goes into the sandwich's filling as though it were noise.
Fig. 2 The quadratic truth over the even design, and the straight line least squares converges to on it. The line is above the truth in the middle and below it at both ends, by 0.3234 at each end and by 0.1667 at the worst point in the middle.

What the line leaves behind

The gap between the truth and the line it projects onto is not noise, and it is not removable. On the even design it is +0.3234 at the low end of the range and +0.3234 at the high end, and −0.1667 at its most negative in the middle — a smooth arch, entirely determined before a single observation is drawn.

Two properties of that arch matter for everything that follows. The first is that its average against the covariate is exactly nought, which is what makes the line a projection: the fit’s own first-order conditions are satisfied, the residuals have mean nought and are uncorrelated with the covariate, and every diagnostic that reads those two quantities reports nothing wrong. That is the same fact the essay measuring what a residual plot can actually show prices from the other direction, and it is why a misfit of this size can survive a competent analysis.

The second is that the misfit is a systematic quantity sitting inside what the fit calls the error. It goes into every variance estimate as though it were noise, which is the connection to the field’s arithmetic of the bread and the filling: the filling of a robust variance estimate is the squared residual, and a squared residual here contains the arch as well as the error. Misspecification does not merely displace the estimand; it inflates and re-shapes the thing a standard error is computed from.

What a straight line through a curve converges to. A quadratic truth over an exponential spread with mean 1, and the straight line least squares converges to on it. The line is not a fit to these points: it is the population projection, computable before any data exists, with slope 2.6000 and intercept 0.0000. It is above the truth in the middle of the design and below it at both ends, by 0.9900 at the low end and 4.4394 at the high, and its average against the covariate is exactly nought — which is what a projection is. The misfit is a systematic error the fit cannot remove, and it goes into the sandwich's filling as though it were noise.
Fig. 3 The same truth and the same fitted line over the exponential design, which reaches to 5.2983 rather than to 1.9900. The misfit runs from 0.9900 at the low end to 4.4394 at the high, against 0.3234 at both ends of the even design.

The exponential design’s version of the arch is the reason its slope is a unit steeper. The design reads out to 5.2983 rather than to 1.9900, so the curve it is chasing has climbed much further, and the misfit runs from 0.9900 at the low end to 4.4394 at the high end rather than sitting symmetrically at 0.3234 at both. A design with a long right tail puts a small number of observations a long way out, and a straight line fitted through them tilts to reach them. There is nothing pathological about the design; an exponential covariate is what a duration, an income or a concentration usually looks like.

The curvature scales the disagreement; the design creates it

An obvious objection is that 0.5 is a large quadratic term, and that the whole effect is a matter of having chosen a truth curved enough to make a point. The sweep answers it by doubling the curvature and reading the same four numbers.

At curv=1\mathrm{curv} = 1 the four targets are 2.6000, 4.6000, 4.6000 and 4.6000, and the gap between the two same-mean designs is 2.0000 rather than 1.0000 — exactly twice. That is what the closed form says: the curvature multiplies 2m+μ3/v2m + \mu_3/v, so it is a scale factor on the disagreement and not the source of it. Set the curvature to nought and all four designs converge on 0.6000, which is the truth’s own slope, because the model is then correct.

So the curvature decides how large the disagreement is and the design decides whether there is one at all. A reader who wants the effect to be small can have it small; what the reader cannot have is a curvature at which two designs with the same mean agree. The only way to that is μ3=0\mu_3 = 0 on both, which is a statement about where the covariate was measured.

One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 2.6000 and the same spread moved to [1, 3] gives 4.6000, while widening it to [0, 4] gives 4.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 4.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 2.0000 apart, and neither is making an error.
Fig. 4 The same four designs with the omitted quadratic term doubled. Every target moves by exactly the doubling, and the two same-mean designs are 2.0000 apart rather than 1.0000.

The projection is a limit, and a small sample is not sitting on it

One number on the first table does not agree with its closed form, and it is the interesting one. On the exponential design the counted mean slope at two hundred rows is 2.5463 against a projection of 2.6000, a standard error of 0.0067 and a t of −8.01. The three symmetric designs agree; this one does not.

That was a failing assertion before it was a finding. The natural reading is an error in the closed form, and it is the wrong reading: sweeping the sample size shows the gap closing. The mean fitted slope on the exponential design is short of its own projection by 0.1662 at fifty rows, 0.0952 at a hundred, 0.0489 at two hundred and 0.0259 at four hundred, and the successive ratios are 1.75, 1.94 and 1.89. A quantity that halves as the sample doubles is a quantity going to nought like one over the sample size, which is what a finite-sample bias looks like and not what a wrong constant looks like.

The mechanism is that a fitted slope is a ratio of two sample sums, and the expectation of a ratio is not the ratio of the expectations. On a design whose covariate is redrawn every time, the denominator (xixˉ)2\sum(x_i - \bar{x})^2 is itself random, and on a skewed design it is very skewed: most samples miss the far tail, a few contain it, and the ones that contain it have a much larger denominator. Averaging over that is not the same as substituting the population moments. On the symmetric designs the same sweep reads −0.0002, 0.0021, 0.0030 and 0.0007 against standard errors between 0.0056 and 0.0020 — no shortfall at any size.

The projection is a limit, not an estimand. How far the mean fitted slope sits from the projection its own design defines, at four sample sizes, over 2500 draws apiece with the covariate redrawn each time. On a symmetric design the two agree at every size — the readings are -0.0036, 0.0021, 0.0025, -0.0004. On an exponential design they do not: the fitted slope is short by -0.1736 at 50 rows and -0.0269 at 400, and the shortfall falls by factors of 1.73, 1.97, 1.89 as the sample doubles. A slope is a ratio of two sample sums, so the number it converges to need not be the number it is centred on, and quoting the limit as the estimand of a small sample is the same substitution this field is about one level up.
Fig. 5 How far the mean fitted slope sits from the projection its own design defines, at four sample sizes on two designs. The exponential design is short by 0.1662 at fifty rows and 0.0259 at four hundred; the even design shows no shortfall at any size.

This matters more than a footnote about a skewed design, because it is the same substitution one level down. The argument of this essay is that the estimand of a wrong model is the projection rather than the truth’s parameter. The sweep then finds that the projection is a limit rather than the number a small sample is centred on. So a reader who corrects the first substitution and quotes the projection has not finished: at fifty rows on this design the projection is six per cent away from where the fitted slope actually sits, and calling it the estimand is the same class of error as calling the truth’s slope the estimand. The honest statement is that a wrong model has a target which depends on the design, and an estimator that reaches that target only asymptotically.

What could have produced these numbers without the claim being true

Four candidates, and the sweep rules out three of them.

Simulation noise. Every counted number carries a standard error and the sweep reports it. The three symmetric designs sit at t of 0.70, 0.70 and 0.40 against their closed forms, and the one design that disagrees does so at −8.01 — a separation of two orders of magnitude in the same units. A reader who suspects that 2,500 draws is thin is entitled to; the answer is that the quantity being separated, 1.0000 between the two same-mean designs, is a hundred and fifty times the largest standard error on the table.

A wrong closed form. The closed form is an expectation of a polynomial in the centred covariate against that design’s own central moments, and it reads no data at all; the count fits samples and reads no moments. The two share no arithmetic. They agree on three designs of four, and the fourth’s disagreement moves like one over the sample size rather than staying put, which is the signature of the count being short rather than of the form being wrong.

Non-constant error variance. The errors in this sweep have a constant variance. That is deliberate: it means the disagreement between the designs cannot be blamed on the shape of the noise, and everything a robust standard error is usually reached for is switched off. What is left is functional form alone, which is the point — the essay that separates a bias landing in the intercept from one landing in the slope makes the same separation for a different defect, and it is worth keeping the two apart.

A fixed design in one case and a redrawn one in the other. The covariate is redrawn on every draw of the sweep, so the counted targets are averages over designs as well as over errors. Holding the covariate fixed removes the finite-sample shortfall on the skewed design and changes none of the closed forms, because the closed forms are population quantities either way. That is the one candidate the sweep does not rule out so much as name: the shortfall in the previous section is the redrawing, and it is reported rather than suppressed.

Where the claim runs out. The result is about a model that is wrong in its functional form, and it says nothing at all about a model that is right. Set the quadratic term to nought and every design gives 0.6000; there is no design-dependence to find, and the whole of this field’s machinery collapses onto the textbook answer. That is not a limitation to apologise for — it is the statement that the effect is caused by the misspecification rather than by least squares.

It is also stated for least squares, where the projection has a closed form and the target solves a normal equation. For a general M-estimator the target solves an estimating equation instead, and it moves with the loss function as well as with the design: a least-absolute-deviations line and a Huber line fitted to the same curved truth on the same covariate converge on three different slopes. That is a stronger version of this essay’s claim and it is not measured here, because a single closed-form engine covering four designs was worth more than a second estimator with only a simulated route. It is named rather than assumed.

And there is a design at which the disagreement is arbitrarily large. Nothing in β1=a1+curv(2m+μ3/v)\beta_1^* = a_1 + \mathrm{curv}(2m + \mu_3/v) is bounded, so a covariate with a long enough tail can put the projection anywhere. What stops that being a reductio is that μ3/v\mu_3/v is a measured property of a design a reader can look at: the essay on what one observation can do to a slope is the extreme case of the same arithmetic with one point rather than a tail, and the leverage that essay computes is exactly what this field’s account of the four corrections reads.

What this leaves for a standard error to do

The estimand under a wrong model is the design’s projection. A standard error is therefore an interval for that number, and the question of whether it is a good interval is a separate question from whether the projection is the number anybody wanted.

There is a second reason to state the estimand before the standard error, and it is arithmetic rather than philosophical. The variance a slope actually has, on a design and under an error variance, is a ratio of two population quantities — the design’s second moment inverted twice, with the errors’ contribution in the middle — and the design appears in all three factors. So a design that decides the estimand also decides the spread of the estimator around it, and the two effects need not point the same way. The exponential design is the extreme case here: its kurtosis is 9.0000 against 1.8000 for every even spread, and the same fourth moment that put its projection a unit away is what sets how far a robust standard error there can sit from a model-based one.

The filling is the design's fourth moment. n times the variance of the slope under the sandwich and under the model, in closed form, on four covariate distributions at a variance leaning towards the edges of the design (γ = 0.8). The bread A = E[xx′] is the same object in both; only the filling B = E[ε²xx′] differs, and its slope element is E[u²σ²(u)] rather than E[σ²]E[u²]. The ratio is exactly 1 + γ(κ − 1) with κ the design's own kurtosis, so the even spreads read ×1.6400 on a kurtosis of 1.80 and the exponential reads ×7.4000 on a kurtosis of 9.00. The size of the correction is a property of the covariate's tail, and no part of it is a property of how much noise there is.
Fig. 6 The slope’s variance under the sandwich and under the model, in closed form on the four designs, with the error variance leaning towards the edges of each. The ratio is 1.6400 on every even spread and 7.4000 on the exponential one, on kurtoses of 1.8000 and 9.0000.

Those two questions are worth keeping apart, because the literature runs them together under the word robust. An interval that covers the projection 95% of the time is doing its job perfectly while telling a reader nothing about the truth’s slope. And an interval that misses the projection is failing at a job it could have done, which is what the field’s coverage measurements count. The first defect is not repairable by a variance estimate; the second is exactly what a variance estimate is for.

What is left unmeasured here is the one a practitioner would ask for: given two studies reporting 1.6 and 2.6, is there an estimator of the truth’s tangent at a stated point that both could compute? There is — fit the quadratic — and the reason it is not the answer is that the whole field assumes the quadratic is the term nobody knew to include. A model that is wrong in a way somebody has named is a model that can be fixed. The interesting case, and the one every standard error in the arithmetic that follows is computed under, is the model that is wrong in a way nobody has named yet, where the projection is not a thing to be corrected but the only thing there is.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Asymptotic biasBiasClosed formConsistencyCurvatureEstimandFunctional formLeast squaresModel misspecificationMonte CarloPopulation projectionProjectionSampling variationSkewnessStandard error