Where the derivative is zero
Worth reading first: Sums of almost anything · What the 95% refers to.
The central limit theorem makes a mean approximately normal. The delta method extends the promise to anything smooth computed from a mean. Expand g(x̄) around μ as far as its tangent line,
and the right-hand side is a constant plus a multiple of an approximately normal quantity, so g(x̄) is approximately normal with standard error |g′(μ)|·σ/√n. It is how a standard error gets attached to a log odds ratio, a coefficient of variation, a correlation, an elasticity, a return level and a forecast many steps ahead. Most software that prints a standard error beside a derived quantity is printing this one.
The approximation has an obvious hole, and the hole is not out at the edges of anything. Where the derivative is zero the tangent line is flat, the multiple of the normal quantity is zero, and the delta method reports that g(x̄) has no spread at all. That is wrong in kind rather than in degree, and it is worth seeing exactly how wrong, because points where a derivative vanishes are ordinary: the square of a mean at zero, a variance p(1 − p) at one half, a product of two effects when both are absent, a squared correlation where there is no correlation.
The squared mean, which can be computed exactly
Take the plainest case. Draw n observations from a normal distribution with mean μ and standard deviation σ, and estimate μ² by x̄². The derivative of x² is 2x, so the delta method’s standard error is 2|μ|σ/√n, which is zero at μ = 0.
Everything about this estimate depends on one number: how far the mean is from the flat point, measured in standard errors of the mean,
Write Z = √n·x̄/σ, which is exactly normal with mean δ and variance one. The estimate is Z² in units of σ²/n, the target is δ², and Z² has a non-central χ² distribution on one degree of freedom. Its mean is δ² + 1, its variance 2 + 4δ² and its skewness . The delta method’s version is a normal with mean δ², variance 4δ² and no skewness at all.
The figure draws both at δ = 1, a mean one standard error from zero. The exact law has mean 2.00 against the tangent line’s 1.00, variance 6.00 against 4.00, and skewness 2.177 against none. The delta method’s normal puts 30.85% of its mass below zero, where a square cannot go, and the largest vertical gap between the two distribution functions — the Kolmogorov distance — is 0.3085. Those are the same number: at δ = 1 the worst disagreement is exactly the approximating normal’s mass on impossible values, Φ(−½).
At the flat point itself the approximation is not approximately wrong. It is a point mass.
A different limit, not a worse approximation
The tangent line is the first term of a Taylor series, and where it vanishes the second term is all that is left:
The square of an approximately normal quantity with variance σ²/n is σ²/n times a χ² on one degree of freedom, so
Three things change at once, and none of them is a matter of accuracy.
The scale is n rather than √n. The estimate settles on its target faster than the central limit theorem’s rate, which sounds like good news and is the reason a standard error computed at the usual rate describes nothing.
The shape is a χ², one-sided and unbounded at zero, where the delta method expected a symmetric bell.
The side is fixed by the sign of g″. Every draw of g(x̄) lies on the same side of g(μ): above it where the flat point is a minimum, below it where it is a maximum. A symmetric interval around a quantity that can only err in one direction is already telling a reader something false before any number is put in it.
For normal data and g(x) = x² the second-order statement is exact at every n, which is what makes the squared mean worth studying first. Every number below about it is a closed form in Φ, and each is counted again on 40,000 draws with a stated seed, because a number with one route to it has nothing to disagree with.
What a plug-in interval does with a slope it has to estimate
Nobody evaluates the delta method’s standard error at the true μ, because μ is the thing being estimated. The interval actually reported plugs in the estimate,
That interval is never of zero width, since |Z| is never exactly zero, so the question of how often it covers has a real answer, and the answer is exact. The interval contains δ² when |Z² − δ²| ≤ 2z|Z|, and solving the two quadratic inequalities that splits into gives
Below that band the whole interval lies under the truth; above it, the whole interval lies over. Both are normal probabilities, so the coverage and both kinds of miss are closed forms at every δ.
At the flat point the band runs from 0 to 2z, and the interval covers whenever |Z| is under 3.92. It covers 99.991% of the time. An interval built on a standard error that is zero over-covers at the one place where that zero is correct, because the plug-in slope is never zero and a squared estimate can never be negative, so almost every interval stretches down past the target of nought.
One and a half standard errors away, it covers 85.978%
Moving off the flat point the coverage falls: 98.759% at δ = 0.25, 95.557% at δ = 0.5, 88.289% at δ = 1. The worst is at δ = 1.45, where the interval covers 85.978% — counted on its own draws, 86.098% ± 0.173% — and from there it climbs back slowly: 87.628% at δ = 2, 91.018% at δ = 3, 92.603% at δ = 4 and 94.356% at δ = 8. A mean eight standard errors from zero is a mean nobody would describe as near zero, and the interval there is still more than half a point short of its promise.
The distance is what decides it, not the sample. A study of 10,000 observations of a quantity whose mean is 0.0145 standard deviations from zero sits at δ = 1.45, exactly where a study of 100 sits when the mean is 0.145 standard deviations out. More data moves a given μ away from the flat point in these units, and it also narrows the band of μ inside which the damage is done — but the depth of the trough does not change with n at all. There is always a region of parameter values where the interval covers 85.978%, and the most a larger study does is make that region smaller in the units of the data.
That is a precise sense in which a 95% interval is a statement about a procedure and not about the data in hand: this procedure’s statement is false on a set of parameter values that no amount of data removes.
Why nearly every miss falls on one side
The lower panel is the part a single coverage number conceals. At δ = 1.45 the intervals that miss lie wholly below the truth 13.86% of the time and wholly above it 0.16% of the time. A two-sided 95% interval promises 2.5% on each side, and this one breaks the promise on the low side by more than a factor of five while keeping it on the high side with a factor of fifteen to spare.
The mechanism is the plug-in slope. A draw with |Z| small produces an estimate near zero and a half-width near zero together, so the interval is short and sits under δ². A draw with |Z| large produces a large estimate and a proportionately large half-width, generous enough to reach back down over the target. The estimate is on average too high — its mean is δ² + 1 — and yet the misses are too low, because the interval’s width moves with the estimate and moves the wrong way.
This is why the side of the misses is the side of the stationary value. The squared mean has a minimum at its flat point, the short intervals collect near that minimum, and the minimum is below every target. A function whose flat point is a maximum will collect its short intervals above, and a binomial variance below shows it doing exactly that.
The imbalance fades at the same slow rate as everything else here. At δ = 20 the interval covers 94.895% and its misses are still 3.115% below against 1.990% above.
It is worth separating this from the return-level interval whose misses also sit below the truth. There the quantity is a steep function of a shape parameter reached through an exponent, and the likelihood is skewed; here the function is flat and the interval’s own width is the cause. The two failures share only the habit of placing a symmetric interval around an asymmetric law.
Squaring the interval instead of linearising the square
The repair needs no approximation at all. An exact interval for μ is x̄ ± zσ/√n. The values μ² takes across that interval form its image under squaring: from the smaller of the two squared endpoints to the larger, or from zero when the interval straddles zero. That image contains μ² whenever the interval for μ contains μ, and also whenever it happens to contain −μ instead, so it cannot cover less than 95% at any δ.
The upper panel of the coverage figure draws it as the second line. It is exactly 95.000% at the flat point, where μ and −μ coincide and the second chance adds nothing; it rises to 97.500% at δ = 1.45, where an interval for μ often reaches across zero to −μ; and it is back to 95.003% by δ = 3 and 95.000% by δ = 4. It never falls under its promise anywhere on the sweep, which is the difference between approximating a function and transforming an interval.
The rule is general and cheap: transform the endpoints of an exact interval, not the standard error of an estimate. For a monotone function the image of an interval is an interval with exactly the same coverage; for a function that folds, like x², the image covers at least as often. What the rule needs is an exact interval for the quantity that was actually measured, and what it costs is that the answer may be lopsided around the estimate. The delta method is what gets used when that seems like too much trouble, and at δ = 1.45 the trouble saved is one draw in seven landing on the wrong side of the answer.
The same rule is why the narrowest of several intervals is so often the one that misses: the delta interval here is narrower than the squared interval at every δ where it fails, and narrower for exactly the reason it fails.
The error is set by √n·μ/σ, and not by n
The coverage figure says the approximation is poor near the flat point and adequate far from it. The rate says how far is far. Measured as the Kolmogorov distance between the exact law of Z² and the delta method’s N(δ², 4δ²), the error is 0.3085 at δ = 1 and 0.0023 at δ = 64, and δ times the distance reads 0.309, 0.317, 0.183, 0.162, 0.154, 0.150 and 0.148 at δ = 1, 2, 4, 8, 16, 32 and 64. It settles, so the error falls like 1/δ.
The constant it settles on can be derived rather than read off. On the delta method’s own scale the squared estimate is
so its distribution function at u is Φ(u − u²/2δ) to first order in 1/δ, which differs from Φ(u) by φ(u)·u²/2δ. That gap is largest at u = ±√2, where it is φ(√2)/δ, and φ(√2) = e⁻¹/√(2π) = 0.1468. The last measured product, 0.148, sits a little over 1% above it and is still falling towards it.
The obvious first attempt at the constant gets it wrong, and the way it goes wrong is informative. Reading the error off the skewness alone — the leading term of an Edgeworth correction, (γ/6)·φ(0), with γ·δ tending to 3 — gives 1/(2√(2π)), about 0.1995, and the measured products head somewhere else. The squared estimate’s mean is δ² + 1 and not δ², and that displacement is of the same order in 1/δ as the skewness. Leaving it out does not give a slightly worse constant; it gives one that is wrong by more than a third, and the two terms partly cancel rather than add.
The consequence is the practical content of the whole essay. A distance of 0.01 needs δ = 15.40, which is n = 237 observations when μ is one standard deviation from the flat point, and a hundred times as many when μ is a tenth of a standard deviation from it. There is no sample size that is large enough in general: for every n there are values of μ close enough to the flat point that the delta method is as wrong as it was at n = 10. The tail of a sum converges last because of where a distribution is read; this converges last because of where the parameter is, and a rule of thumb stated in n cannot see either.
The same limit, for data that are not normal
For normal data the χ² law is exact at the flat point. For anything with a finite variance it is only a limit, and the essay on sums of a skewed, one-sided source supplies the obvious source to watch it arrive on.
Take X = E − 1 with E exponential: mean zero, variance one, skewness two. The second-order delta method says n·x̄² tends to a χ² on one degree of freedom, and the Kolmogorov distance from it, on 20,000 means at each n, is 0.0802 at n = 2, 0.0176 at n = 8, 0.0083 at n = 32 and 0.0048 at n = 128. On a normal source, where the χ² is exact, the same count reads 0.0050, 0.0047, 0.0070 and 0.0060 — the size of disagreement 20,000 draws show against a law that fits perfectly, whose median is about 0.0059. By n = 128 the skewed source cannot be told from its limit at this resolution.
The skewness of the source is visible only at the smallest n, where a sum of two exponentials is still lopsided enough that its square leans. What it never does is look like the first-order answer. The tangent line puts every draw at zero and is at distance 1 at every n, so there is no sample size at which it becomes the right description of a quantity at a flat point. The second term is the description, and it is a different distribution from the one the first term promised, not a correction to it.
A binomial variance at one half
The variance of a single yes-or-no outcome is p(1 − p), estimated by p̂(1 − p̂), and its derivative 1 − 2p is zero at p = ½. The delta interval is p̂(1 − p̂) ± z·|1 − 2p̂|·√(p̂(1 − p̂)/n), and because p̂ takes only n + 1 values its coverage can be summed over every possible sample rather than simulated.
At a hundred trials the interval covers 99.98% at p = ½. One step away, at p = 0.49, it covers 92.17%, and its worst on the grid is 84.38% at p = 0.57 and its mirror image, where the flat-point distance √n·|p − ½|/√(p(1 − p)) is 1.41 — within a few hundredths of the squared mean’s 1.45. At twenty-five trials the worst is 80.06% at p = 0.6, and at four hundred it is 86.17% at p = 0.53. The ragged edge between neighbouring proportions is discreteness laid over the flat-point failure, the same raggedness an interval for a proportion shows as its sample grows.
The misses fall the other way from the squared mean’s. At p = 0.57, 15.35% of intervals lie wholly above the truth and 0.28% below it, and at p = ½ an interval cannot miss above at all, because p̂(1 − p̂) never exceeds ¼. The squared mean’s flat point is a minimum and its misses fall below; the variance’s flat point is a maximum and its misses fall above. The short intervals collect at the stationary value, on whichever side of the target that happens to be, and the prediction from the previous section holds with the sign reversed.
This is not a curiosity about coin flips. An estimated Bernoulli variance sits inside every standard error for a proportion, every heterozygosity and every design effect built from one, and the region where its own interval misbehaves is the region near a half where such quantities are most often computed.
A product at the origin, and a test that almost never rejects
The indirect effect in a mediation analysis is a product a·b of two separately estimated paths: the treatment’s effect on the mediator, and the mediator’s effect on the outcome — the path a direct effect is separated from when a covariate the treatment caused is adjusted for. The Sobel test divides the product by its delta-method standard error, √(a²·SE(b)² + b²·SE(a)²). The gradient of a·b is (b, a), which is zero when both paths are.
At that point something unusually clean happens. With both paths standardised the statistic is za·zb/√(za² + zb²) for two independent standard normals, and that ratio is itself exactly normal with variance ¼. Comparing it with 1.96 is comparing a standard normal with 3.92, so a nominal 5% test rejects 0.0089% of the time. Counted, it rejects 40 times in 400,000, against 35 expected; a quadrature over zb reproduces the closed form to every digit printed here.
The same 0.0089% is the squared mean’s plug-in interval missing at its own flat point, and for the same reason: both events are a standard normal falling beyond 3.92. Two different procedures, one delta method, one number.
Holding the first path at zero and moving the second out, the size rises — 0.0187% at half a standard error, 0.0589% at one, 0.3622% at two, 1.1564% at three, 2.2347% at four — and is still only 3.7091% at six. The test that simply requires both paths to be significant has size 0.25% at the origin and 5.00% at the far end. Both are conservative near the origin, and the Sobel test is conservative there by a factor of more than five hundred. The origin is where the hypothesis of no mediation is most often true, and it is also where the test’s power against a small indirect effect is least, because a test that almost never rejects under the null has spent none of its error rate on finding anything.
Other flat points already met here
A squared correlation at zero correlation is the squared mean again: n·r² tends to a χ² on one degree of freedom, whose mean of one is the 1/(n − 1) that a useless predictor adds to R². The inflation measured there is not a small-sample bias waiting for a correction. It is the whole of the estimate’s law at a flat point, and the delta method’s standard error for R² at zero, which is zero, is the first-order answer to it.
The height of a fitted optimum is another. The fitted curve is flat at its maximum by construction, so the delta method gives the uncertainty in where the maximum is no weight at all in how high it is. That is exactly backwards from the difficulty of locating it, which is a ratio whose honest interval is sometimes unbounded.
Two nearby failures are not flat points and should not be mistaken for them. A population spread that estimates to exactly zero on a third of eight-group datasets sits on the boundary of its parameter space, where the limit is a mixture with a point mass rather than a χ². And a forecast carried through a biased persistence estimate fails the delta method through curvature rather than flatness: the tangent there has a slope, and the slope is too small a description of a function that bends away from it.
Which of these numbers are proved
Every coverage for the squared mean is a closed form in Φ, and every one is also counted, with the counts inside four standard errors of the closed forms at twelve distances from zero to twenty. The squared exact interval’s floor of 95% is not a measurement; it follows from the image argument and holds at every δ, counted or not. The binomial coverages are sums over every count and involve no simulation at all.
The Sobel size at the origin rests on an identity — za·zb/√(za² + zb²) is N(0, ¼) for independent standard normals — which is proved, and which is checked here by a quadrature and by a count. The rate constant φ(√2) comes from a first-order expansion; what is shown is that the exact distances approach it, not a bound on how fast they do. For non-normal data the χ² limit is a theorem, and the approach to it at finite n is only measured, on one source.
The refusals are the same standards turned on things that must fail them. The sweep standard the squared exact interval passes — within half a point of 95% at every distance counted — is applied to the plug-in delta interval, and rejects it at 85.978%. The tangent line’s point mass, offered as the limit of n·x̄² at the flat point, is rejected at a distance of 1. And the Sobel test’s nominal size, claimed to within a point of 5% at the origin, is rejected at 0.0089%. Each of those checks passes the procedure it was written for, which is the only evidence that it is capable of failing anything.
Where this goes next
The other place the delta method’s answer is wrong in kind is not a flat point but a place where the function stops being smooth at all: a ratio whose denominator can be zero. The ratio of two normal means is where widening a delta interval rescues nothing, where no interval that is always finite can hold its coverage, and where the exact answer has to be the whole line some of the time.
After that, the flat point in more than one dimension, which nothing here has touched. The squared length of a mean vector, R² with several predictors, or any smooth function at an interior minimum has a Hessian where the squared mean has a single g″, and the second-order law becomes a weighted sum of χ² variables whose weights are the Hessian’s eigenvalues and whose count is its rank. At a saddle it is a difference of χ² variables that can take either sign. That question is distinct from this one in the one respect that matters most: the side the misses fall on is no longer fixed by the sign of a single second derivative, so the rule that the short intervals collect at the stationary value has to be restated, and may not survive.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name binomial proportion, closed form, confidence interval, coverage
- A simulation that stops when it looks settled — both name binomial proportion, closed form, confidence interval, coverage
- The check worth more than the check — both name binomial proportion, closed form, coverage, exact enumeration
- The same draws for both methods — both name binomial proportion, closed form, confidence interval, coverage
- A block size that changes — both name chi-square, confidence interval, coverage
- A bound written for a coin — both name central limit theorem, convergence rate, skewness
Named objects
A flat tag is an object no other essay names yet.
Binomial proportionCentral limit theoremChi-squareClosed formConfidence intervalConvergence rateCoverageDelta methodExact enumerationKolmogorov–SmirnovNon-central χ²Second-order delta methodSkewnessSobel testStationary point