The distribution itself

Where the derivative is zero

The delta method reads a standard error off a tangent line, and at a flat point the tangent says the spread is zero. The interval built on it for a squared mean covers 99.991% there and 85.978% one and a half standard errors away, with nearly every miss on the same side — and the law it should have used is a χ², not a normal.

Worth reading first: Sums of almost anything · What the 95% refers to.

The central limit theorem makes a mean approximately normal. The delta method extends the promise to anything smooth computed from a mean. Expand g(x̄) around μ as far as its tangent line,

g(xˉ)g(μ)+g(μ)(xˉμ),g(\bar x) \approx g(\mu) + g'(\mu)\,(\bar x - \mu),

and the right-hand side is a constant plus a multiple of an approximately normal quantity, so g(x̄) is approximately normal with standard error |g′(μ)|·σ/√n. It is how a standard error gets attached to a log odds ratio, a coefficient of variation, a correlation, an elasticity, a return level and a forecast many steps ahead. Most software that prints a standard error beside a derived quantity is printing this one.

The approximation has an obvious hole, and the hole is not out at the edges of anything. Where the derivative is zero the tangent line is flat, the multiple of the normal quantity is zero, and the delta method reports that g(x̄) has no spread at all. That is wrong in kind rather than in degree, and it is worth seeing exactly how wrong, because points where a derivative vanishes are ordinary: the square of a mean at zero, a variance p(1 − p) at one half, a product of two effects when both are absent, a squared correlation where there is no correlation.

The squared mean, which can be computed exactly

Take the plainest case. Draw n observations from a normal distribution with mean μ and standard deviation σ, and estimate μ² by x̄². The derivative of x² is 2x, so the delta method’s standard error is 2|μ|σ/√n, which is zero at μ = 0.

Everything about this estimate depends on one number: how far the mean is from the flat point, measured in standard errors of the mean,

δ=nμ/σ.\delta = \sqrt{n}\,\mu/\sigma.

Write Z = √n·x̄/σ, which is exactly normal with mean δ and variance one. The estimate is Z² in units of σ²/n, the target is δ², and Z² has a non-central χ² distribution on one degree of freedom. Its mean is δ² + 1, its variance 2 + 4δ² and its skewness 22(1+3δ2)/(1+2δ2)3/22\sqrt2\,(1+3\delta^2)/(1+2\delta^2)^{3/2}. The delta method’s version is a normal with mean δ², variance 4δ² and no skewness at all.

The squared estimate 1 standard errors from the flat point, exact and linearised. At δ = √n·μ/σ = 1 the exact law of the squared estimate has mean 2.00, variance 6.00 and skewness 2.177; the delta method's normal has mean 1.00, variance 4.00, no skewness, and 30.85% of its mass below zero, where a square cannot go. The Kolmogorov distance between them is 0.3085.
Fig. 1 The squared estimate one standard error from the flat point. The solid curve is its exact law, a non-central χ² on one degree of freedom; the dashed curve is the normal the delta method puts in its place, with the part of it below zero, where a square cannot go, shaded.

The figure draws both at δ = 1, a mean one standard error from zero. The exact law has mean 2.00 against the tangent line’s 1.00, variance 6.00 against 4.00, and skewness 2.177 against none. The delta method’s normal puts 30.85% of its mass below zero, where a square cannot go, and the largest vertical gap between the two distribution functions — the Kolmogorov distance — is 0.3085. Those are the same number: at δ = 1 the worst disagreement is exactly the approximating normal’s mass on impossible values, Φ(−½).

At the flat point itself the approximation is not approximately wrong. It is a point mass.

The squared estimate at the flat point, exact and linearised. At μ = 0 the tangent line is flat, so the delta method gives the square of the mean no spread at all: all of its mass sits at zero. The exact law of n·x̄²/σ² is a χ² on one degree of freedom, with mean 1 and variance 2, and the Kolmogorov distance between the two is 1.0000.
Fig. 2 The same comparison at μ = 0. The tangent line is horizontal, so the delta method puts every draw of the squared mean at zero; the exact law of n·x̄²/σ² is a χ² on one degree of freedom, unbounded at zero with mean one. The distance between the two is 1.

A different limit, not a worse approximation

The tangent line is the first term of a Taylor series, and where it vanishes the second term is all that is left:

g(xˉ)g(μ)12g(μ)(xˉμ)2.g(\bar x) - g(\mu) \approx \tfrac12\, g''(\mu)\,(\bar x - \mu)^2 .

The square of an approximately normal quantity with variance σ²/n is σ²/n times a χ² on one degree of freedom, so

n{g(xˉ)g(μ)}    12g(μ)σ2χ12.n\,\{g(\bar x) - g(\mu)\} \;\longrightarrow\; \tfrac12\, g''(\mu)\,\sigma^2\,\chi^2_1 .

Three things change at once, and none of them is a matter of accuracy.

The scale is n rather than √n. The estimate settles on its target faster than the central limit theorem’s rate, which sounds like good news and is the reason a standard error computed at the usual rate describes nothing.

The shape is a χ², one-sided and unbounded at zero, where the delta method expected a symmetric bell.

The side is fixed by the sign of g″. Every draw of g(x̄) lies on the same side of g(μ): above it where the flat point is a minimum, below it where it is a maximum. A symmetric interval around a quantity that can only err in one direction is already telling a reader something false before any number is put in it.

For normal data and g(x) = x² the second-order statement is exact at every n, which is what makes the squared mean worth studying first. Every number below about it is a closed form in Φ, and each is counted again on 40,000 draws with a stated seed, because a number with one route to it has nothing to disagree with.

What a plug-in interval does with a slope it has to estimate

Nobody evaluates the delta method’s standard error at the true μ, because μ is the thing being estimated. The interval actually reported plugs in the estimate,

xˉ2±z2xˉσ/n,or in standard unitsZ2±2zZ.\bar x^2 \pm z \cdot 2|\bar x|\,\sigma/\sqrt n, \qquad\text{or in standard units}\qquad Z^2 \pm 2z\,|Z| .

That interval is never of zero width, since |Z| is never exactly zero, so the question of how often it covers has a real answer, and the answer is exact. The interval contains δ² when |Z² − δ²| ≤ 2z|Z|, and solving the two quadratic inequalities that splits into gives

z2+δ2z    Z    z2+δ2+z.\sqrt{z^2+\delta^2} - z \;\le\; |Z| \;\le\; \sqrt{z^2+\delta^2} + z .

Below that band the whole interval lies under the truth; above it, the whole interval lies over. Both are normal probabilities, so the coverage and both kinds of miss are closed forms at every δ.

At the flat point the band runs from 0 to 2z, and the interval covers whenever |Z| is under 3.92. It covers 99.991% of the time. An interval built on a standard error that is zero over-covers at the one place where that zero is correct, because the plug-in slope is never zero and a squared estimate can never be negative, so almost every interval stretches down past the target of nought.

One and a half standard errors away, it covers 85.978%

Coverage of two 95% intervals for a squared mean, against the distance from the flat point. The plug-in delta interval covers 99.991% at the flat point, falls to 85.978% at δ = 1.45, and is still 94.36% at δ = 8. At its worst 13.86% of intervals sit wholly below the truth and 0.16% above it. The exact interval for μ, squared, never covers under 95%. Dots are counts of 40,000 draws each.
Fig. 3 Above, the coverage of two nominal 95% intervals for μ² against δ; below, the share of plug-in delta intervals lying wholly below and wholly above the truth. Lines are the closed forms and dots are counts of 40,000 draws at each δ.

Moving off the flat point the coverage falls: 98.759% at δ = 0.25, 95.557% at δ = 0.5, 88.289% at δ = 1. The worst is at δ = 1.45, where the interval covers 85.978% — counted on its own draws, 86.098% ± 0.173% — and from there it climbs back slowly: 87.628% at δ = 2, 91.018% at δ = 3, 92.603% at δ = 4 and 94.356% at δ = 8. A mean eight standard errors from zero is a mean nobody would describe as near zero, and the interval there is still more than half a point short of its promise.

The distance is what decides it, not the sample. A study of 10,000 observations of a quantity whose mean is 0.0145 standard deviations from zero sits at δ = 1.45, exactly where a study of 100 sits when the mean is 0.145 standard deviations out. More data moves a given μ away from the flat point in these units, and it also narrows the band of μ inside which the damage is done — but the depth of the trough does not change with n at all. There is always a region of parameter values where the interval covers 85.978%, and the most a larger study does is make that region smaller in the units of the data.

That is a precise sense in which a 95% interval is a statement about a procedure and not about the data in hand: this procedure’s statement is false on a set of parameter values that no amount of data removes.

Why nearly every miss falls on one side

The lower panel is the part a single coverage number conceals. At δ = 1.45 the intervals that miss lie wholly below the truth 13.86% of the time and wholly above it 0.16% of the time. A two-sided 95% interval promises 2.5% on each side, and this one breaks the promise on the low side by more than a factor of five while keeping it on the high side with a factor of fifteen to spare.

The mechanism is the plug-in slope. A draw with |Z| small produces an estimate near zero and a half-width near zero together, so the interval is short and sits under δ². A draw with |Z| large produces a large estimate and a proportionately large half-width, generous enough to reach back down over the target. The estimate is on average too high — its mean is δ² + 1 — and yet the misses are too low, because the interval’s width moves with the estimate and moves the wrong way.

This is why the side of the misses is the side of the stationary value. The squared mean has a minimum at its flat point, the short intervals collect near that minimum, and the minimum is below every target. A function whose flat point is a maximum will collect its short intervals above, and a binomial variance below shows it doing exactly that.

The imbalance fades at the same slow rate as everything else here. At δ = 20 the interval covers 94.895% and its misses are still 3.115% below against 1.990% above.

It is worth separating this from the return-level interval whose misses also sit below the truth. There the quantity is a steep function of a shape parameter reached through an exponent, and the likelihood is skewed; here the function is flat and the interval’s own width is the cause. The two failures share only the habit of placing a symmetric interval around an asymmetric law.

Squaring the interval instead of linearising the square

The repair needs no approximation at all. An exact interval for μ is x̄ ± zσ/√n. The values μ² takes across that interval form its image under squaring: from the smaller of the two squared endpoints to the larger, or from zero when the interval straddles zero. That image contains μ² whenever the interval for μ contains μ, and also whenever it happens to contain −μ instead, so it cannot cover less than 95% at any δ.

The upper panel of the coverage figure draws it as the second line. It is exactly 95.000% at the flat point, where μ and −μ coincide and the second chance adds nothing; it rises to 97.500% at δ = 1.45, where an interval for μ often reaches across zero to −μ; and it is back to 95.003% by δ = 3 and 95.000% by δ = 4. It never falls under its promise anywhere on the sweep, which is the difference between approximating a function and transforming an interval.

The rule is general and cheap: transform the endpoints of an exact interval, not the standard error of an estimate. For a monotone function the image of an interval is an interval with exactly the same coverage; for a function that folds, like x², the image covers at least as often. What the rule needs is an exact interval for the quantity that was actually measured, and what it costs is that the answer may be lopsided around the estimate. The delta method is what gets used when that seems like too much trouble, and at δ = 1.45 the trouble saved is one draw in seven landing on the wrong side of the answer.

The same rule is why the narrowest of several intervals is so often the one that misses: the delta interval here is narrower than the squared interval at every δ where it fails, and narrower for exactly the reason it fails.

The error is set by √n·μ/σ, and not by n

How fast the delta method's normal approaches the exact law of a squared mean. Kolmogorov distance between the exact law of Z² and N(δ², 4δ²), against δ on logarithmic axes. δ times the distance is 0.309, 0.317, 0.183, 0.162, 0.154, 0.150, 0.148 at δ = 1, 2, 4, 8, 16, 32, 64, approaching φ(√2) = 0.1468. A distance of 0.01 needs δ = 15.40, which is n = 237 observations when μ is one standard deviation from the flat point and a hundred times that when it is a tenth of one.
Fig. 4 Kolmogorov distance between the exact law of the squared estimate and the delta method’s normal, against δ, on logarithmic axes. The number above each point is δ times the distance; the dashed line is the limit those numbers approach.

The coverage figure says the approximation is poor near the flat point and adequate far from it. The rate says how far is far. Measured as the Kolmogorov distance between the exact law of Z² and the delta method’s N(δ², 4δ²), the error is 0.3085 at δ = 1 and 0.0023 at δ = 64, and δ times the distance reads 0.309, 0.317, 0.183, 0.162, 0.154, 0.150 and 0.148 at δ = 1, 2, 4, 8, 16, 32 and 64. It settles, so the error falls like 1/δ.

The constant it settles on can be derived rather than read off. On the delta method’s own scale the squared estimate is

Z2δ22δ  =  ε+ε22δ,εN(0,1),\frac{Z^2 - \delta^2}{2\delta} \;=\; \varepsilon + \frac{\varepsilon^2}{2\delta}, \qquad \varepsilon \sim N(0,1),

so its distribution function at u is Φ(u − u²/2δ) to first order in 1/δ, which differs from Φ(u) by φ(u)·u²/2δ. That gap is largest at u = ±√2, where it is φ(√2)/δ, and φ(√2) = e⁻¹/√(2π) = 0.1468. The last measured product, 0.148, sits a little over 1% above it and is still falling towards it.

The obvious first attempt at the constant gets it wrong, and the way it goes wrong is informative. Reading the error off the skewness alone — the leading term of an Edgeworth correction, (γ/6)·φ(0), with γ·δ tending to 3 — gives 1/(2√(2π)), about 0.1995, and the measured products head somewhere else. The squared estimate’s mean is δ² + 1 and not δ², and that displacement is of the same order in 1/δ as the skewness. Leaving it out does not give a slightly worse constant; it gives one that is wrong by more than a third, and the two terms partly cancel rather than add.

The consequence is the practical content of the whole essay. A distance of 0.01 needs δ = 15.40, which is n = 237 observations when μ is one standard deviation from the flat point, and a hundred times as many when μ is a tenth of a standard deviation from it. There is no sample size that is large enough in general: for every n there are values of μ close enough to the flat point that the delta method is as wrong as it was at n = 10. The tail of a sum converges last because of where a distribution is read; this converges last because of where the parameter is, and a rule of thumb stated in n cannot see either.

The same limit, for data that are not normal

For normal data the χ² law is exact at the flat point. For anything with a finite variance it is only a limit, and the essay on sums of a skewed, one-sided source supplies the obvious source to watch it arrive on.

The squared mean of a skewed source, against the second-order limit. Kolmogorov distance of n·x̄², for means of n centred exponential draws, from χ² on one degree of freedom: 0.0802, 0.0176, 0.0083, 0.0048 at n = 2, 8, 32, 128, on 20,000 means each. The same count on a normal source reads 0.0050, 0.0047, 0.0070, 0.0060, which is sampling noise. The first-order law puts every draw at zero and is at distance 1 at every n.
Fig. 5 The squared mean of centred exponential draws, n·x̄², against a χ² on one degree of freedom, on 20,000 means at each n. The normal source beside it is the control, and the shaded band is three times the distance a perfectly matched sample of that size shows by chance.

Take X = E − 1 with E exponential: mean zero, variance one, skewness two. The second-order delta method says n·x̄² tends to a χ² on one degree of freedom, and the Kolmogorov distance from it, on 20,000 means at each n, is 0.0802 at n = 2, 0.0176 at n = 8, 0.0083 at n = 32 and 0.0048 at n = 128. On a normal source, where the χ² is exact, the same count reads 0.0050, 0.0047, 0.0070 and 0.0060 — the size of disagreement 20,000 draws show against a law that fits perfectly, whose median is about 0.0059. By n = 128 the skewed source cannot be told from its limit at this resolution.

The skewness of the source is visible only at the smallest n, where a sum of two exponentials is still lopsided enough that its square leans. What it never does is look like the first-order answer. The tangent line puts every draw at zero and is at distance 1 at every n, so there is no sample size at which it becomes the right description of a quantity at a flat point. The second term is the description, and it is a different distribution from the one the first term promised, not a correction to it.

A binomial variance at one half

The variance of a single yes-or-no outcome is p(1 − p), estimated by p̂(1 − p̂), and its derivative 1 − 2p is zero at p = ½. The delta interval is p̂(1 − p̂) ± z·|1 − 2p̂|·√(p̂(1 − p̂)/n), and because p̂ takes only n + 1 values its coverage can be summed over every possible sample rather than simulated.

Exact coverage of the delta interval for a binomial variance, n = 100. Every count from 0 to 100 summed at each p. At p = ½, where p(1 − p) is flat, the interval covers 99.98%; its worst on this grid is 84.38% at p = 0.57, a flat-point distance of 1.41. Its misses fall above the truth, because the flat point of p(1 − p) is a maximum, and the squared mean's fall below, because its flat point is a minimum.
Fig. 6 Exact coverage of the delta interval for p(1 − p) at a hundred trials, every count summed, at proportions from 0.30 to 0.70, with the misses split by side beneath. The dashed verticals mark where √n·|p − ½|/√(p(1 − p)) is 1.45, the squared mean’s worst distance.

At a hundred trials the interval covers 99.98% at p = ½. One step away, at p = 0.49, it covers 92.17%, and its worst on the grid is 84.38% at p = 0.57 and its mirror image, where the flat-point distance √n·|p − ½|/√(p(1 − p)) is 1.41 — within a few hundredths of the squared mean’s 1.45. At twenty-five trials the worst is 80.06% at p = 0.6, and at four hundred it is 86.17% at p = 0.53. The ragged edge between neighbouring proportions is discreteness laid over the flat-point failure, the same raggedness an interval for a proportion shows as its sample grows.

The misses fall the other way from the squared mean’s. At p = 0.57, 15.35% of intervals lie wholly above the truth and 0.28% below it, and at p = ½ an interval cannot miss above at all, because p̂(1 − p̂) never exceeds ¼. The squared mean’s flat point is a minimum and its misses fall below; the variance’s flat point is a maximum and its misses fall above. The short intervals collect at the stationary value, on whichever side of the target that happens to be, and the prediction from the previous section holds with the sign reversed.

This is not a curiosity about coin flips. An estimated Bernoulli variance sits inside every standard error for a proportion, every heterozygosity and every design effect built from one, and the region where its own interval misbehaves is the region near a half where such quantities are most often computed.

A product at the origin, and a test that almost never rejects

The size of two tests that a product of estimates is zero, with the first path at zero. The Sobel statistic a·b/√(a²·SE(b)² + b²·SE(a)²) is exactly normal with variance ¼ when both paths are zero, so a 5% test rejects 0.0089% of the time. With the second path 6 standard errors out it rejects 3.71%. The test that requires both paths to be significant rejects 0.25% at the origin and 5.00% at the far end. Dots are counts: 40 rejections in 400,000 at the origin.
Fig. 7 The rejection rate of two 5% tests that a product of two independent estimates is zero, with the first estimate’s true value at zero and the second’s moving out, on a logarithmic scale. The Sobel line is a quadrature; the dots are counts.

The indirect effect in a mediation analysis is a product a·b of two separately estimated paths: the treatment’s effect on the mediator, and the mediator’s effect on the outcome — the path a direct effect is separated from when a covariate the treatment caused is adjusted for. The Sobel test divides the product by its delta-method standard error, √(a²·SE(b)² + b²·SE(a)²). The gradient of a·b is (b, a), which is zero when both paths are.

At that point something unusually clean happens. With both paths standardised the statistic is za·zb/√(za² + zb²) for two independent standard normals, and that ratio is itself exactly normal with variance ¼. Comparing it with 1.96 is comparing a standard normal with 3.92, so a nominal 5% test rejects 0.0089% of the time. Counted, it rejects 40 times in 400,000, against 35 expected; a quadrature over zb reproduces the closed form to every digit printed here.

The same 0.0089% is the squared mean’s plug-in interval missing at its own flat point, and for the same reason: both events are a standard normal falling beyond 3.92. Two different procedures, one delta method, one number.

Holding the first path at zero and moving the second out, the size rises — 0.0187% at half a standard error, 0.0589% at one, 0.3622% at two, 1.1564% at three, 2.2347% at four — and is still only 3.7091% at six. The test that simply requires both paths to be significant has size 0.25% at the origin and 5.00% at the far end. Both are conservative near the origin, and the Sobel test is conservative there by a factor of more than five hundred. The origin is where the hypothesis of no mediation is most often true, and it is also where the test’s power against a small indirect effect is least, because a test that almost never rejects under the null has spent none of its error rate on finding anything.

Other flat points already met here

A squared correlation at zero correlation is the squared mean again: n·r² tends to a χ² on one degree of freedom, whose mean of one is the 1/(n − 1) that a useless predictor adds to R². The inflation measured there is not a small-sample bias waiting for a correction. It is the whole of the estimate’s law at a flat point, and the delta method’s standard error for R² at zero, which is zero, is the first-order answer to it.

The height of a fitted optimum is another. The fitted curve is flat at its maximum by construction, so the delta method gives the uncertainty in where the maximum is no weight at all in how high it is. That is exactly backwards from the difficulty of locating it, which is a ratio whose honest interval is sometimes unbounded.

Two nearby failures are not flat points and should not be mistaken for them. A population spread that estimates to exactly zero on a third of eight-group datasets sits on the boundary of its parameter space, where the limit is a mixture with a point mass rather than a χ². And a forecast carried through a biased persistence estimate fails the delta method through curvature rather than flatness: the tangent there has a slope, and the slope is too small a description of a function that bends away from it.

Which of these numbers are proved

Every coverage for the squared mean is a closed form in Φ, and every one is also counted, with the counts inside four standard errors of the closed forms at twelve distances from zero to twenty. The squared exact interval’s floor of 95% is not a measurement; it follows from the image argument and holds at every δ, counted or not. The binomial coverages are sums over every count and involve no simulation at all.

The Sobel size at the origin rests on an identity — za·zb/√(za² + zb²) is N(0, ¼) for independent standard normals — which is proved, and which is checked here by a quadrature and by a count. The rate constant φ(√2) comes from a first-order expansion; what is shown is that the exact distances approach it, not a bound on how fast they do. For non-normal data the χ² limit is a theorem, and the approach to it at finite n is only measured, on one source.

The refusals are the same standards turned on things that must fail them. The sweep standard the squared exact interval passes — within half a point of 95% at every distance counted — is applied to the plug-in delta interval, and rejects it at 85.978%. The tangent line’s point mass, offered as the limit of n·x̄² at the flat point, is rejected at a distance of 1. And the Sobel test’s nominal size, claimed to within a point of 5% at the origin, is rejected at 0.0089%. Each of those checks passes the procedure it was written for, which is the only evidence that it is capable of failing anything.

Where this goes next

The other place the delta method’s answer is wrong in kind is not a flat point but a place where the function stops being smooth at all: a ratio whose denominator can be zero. The ratio of two normal means is where widening a delta interval rescues nothing, where no interval that is always finite can hold its coverage, and where the exact answer has to be the whole line some of the time.

After that, the flat point in more than one dimension, which nothing here has touched. The squared length of a mean vector, R² with several predictors, or any smooth function at an interior minimum has a Hessian where the squared mean has a single g″, and the second-order law becomes a weighted sum of χ² variables whose weights are the Hessian’s eigenvalues and whose count is its rank. At a saddle it is a difference of χ² variables that can take either sign. That question is distinct from this one in the one respect that matters most: the side the misses fall on is no longer fixed by the sign of a single second derivative, so the rule that the short intervals collect at the stationary value has to be restated, and may not survive.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Binomial proportionCentral limit theoremChi-squareClosed formConfidence intervalConvergence rateCoverageDelta methodExact enumerationKolmogorov–SmirnovNon-central χ²Second-order delta methodSkewnessSobel testStationary point