A flat point with more than one direction
Worth reading first: The optimum is a ratio, and its interval is sometimes the whole line.
The essay on the delta method at a flat point found what the delta method does at a flat point: the tangent line says the spread is zero, the interval built on it covers 99.991% there and 85.978% one and a half standard errors away, and nearly every miss is on the same side — fixed, it argued, by the sign of a single second derivative.
A function of several means has no single second derivative. It has a Hessian, and a Hessian has a sign only when its eigenvalues agree. What happens when they do not is the question, and the answer changes which part of the one-variable finding survives.
The law is a weighted sum of squared normals with the Hessian’s eigenvalues as weights, and where the weights cancel, two of its moments are exactly zero.
Three surfaces with the same gradient
All three have a stationary point at the origin and differ only in curvature:
- a bowl, , with eigenvalues 2 and 2;
- a valley, , with eigenvalues 2 and 8;
- a saddle, , with eigenvalues 2 and −2.
Expanding g about the stationary point, the first-order term vanishes and the second-order one is ½( − μ)′H( − μ). Scaled by n it becomes ½ Z′HZ, where Z is a pair of independent standard normals — and diagonalising H turns that into , a weighted sum of independent variables on one degree of freedom each.
Everything that follows is a reading of that sum. Its mean is , its variance , and its skewness — closed forms with nothing in them but the eigenvalues.
The bowl’s law is a on two degrees of freedom with skewness 2.023 against a predicted 2.000. The valley’s is a weighted sum with skewness 2.638 against 2.623. The saddle’s is a difference of two ’s with skewness −0.017 against 0.000 exactly.
The three histograms are worth reading as three answers to one question. A reader asked what happens to an estimate at a flat point would reasonably expect one answer, since the flatness is the same in all three: the gradient is zero, the tangent plane is horizontal, and the first order of the expansion contributes nothing. What differs is only the second order, and the second order is what is left.
That is why the eigenvalues carry everything. The law is a weighted sum with them as weights and no other quantity from the problem enters — not the dimension, not the function’s value, not the shape of the surface away from the point. Two functions with the same Hessian at a stationary point have the same second-order law there, however different they look elsewhere.
What the single second derivative was standing in for
The single-variable case attributed the behaviour to the sign of g″, and the multivariable case says what that sign was doing. It was fixing the sign of — which is to say, the bias.
At a stationary point the estimate ĝ = g() is systematically wrong by half the Hessian’s trace, in units of /n. For the bowl that is 2.008 measured against 2.000 predicted; for the valley 5.028 against 5.000; for the saddle −0.006 against 0.000.
At a saddle the bias is exactly zero and it is zero for an uninformative reason. The two directions’ contributions cancel: the estimate is too high along one axis and too low along the other by the same amount in expectation, and the cancellation says nothing about the estimate being good. Its spread is unchanged — the saddle’s law has standard deviation 2.017 against the bowl’s 2.014 — so the estimate is exactly as variable and no longer biased.
That is the single-variable finding’s limit stated precisely: it extends to any definite Hessian and it stops at an indefinite one, and where it stops it stops by cancellation rather than by repair.
There is a practical version of that worth stating, because the bias is the part a study can do something about. is a number: at forty observations with unit standard errors, a bowl’s estimate of g is too high by 0.05 on a quantity whose own scale is set by /n = 0.025, so the bias is twice the natural unit of the problem. At four hundred observations it is a tenth of that in absolute terms and the same multiple of the unit — because both scale like 1/n, which is the shape of a second-order bias and the reason it does not shrink relative to what it is biasing.
The estimate’s spread is also of order /n at a stationary point, which is what makes the ratio fixed. So the bias-to-noise ratio at a flat point does not improve with the sample at all, which is the opposite of the ordinary situation and is worth knowing before quoting a large-n argument.
The cancellation buys nothing
If the bias were the whole story, the saddle would be the well-behaved case. It is not.
At the stationary point itself both intervals cover 99.96% and 99.99% — the single-variable essay’s finding, arriving unchanged, and for the same reason: the estimated gradient is near zero, so the interval is nearly a point, and it happens to sit on the right side of the truth almost always.
Half a standard error away the bowl covers 93.93% and the saddle 91.52%. The saddle’s dip is the deeper one, so the case with no bias is the case with the worse interval.
The two quantities are answering different questions and the saddle separates them. Bias is about where the estimate sits on average; coverage is about whether an interval built from an estimated gradient contains the truth, and the estimated gradient is bad at a flat point whatever the curvature does. Removing a bias by cancellation does not fix a standard error that is reading a tangent plane which is not there.
Where a saddle actually turns up
A saddle sounds like a constructed case and it is the ordinary one for several quantities that have already come up.
A difference of two squared means — the contrast between two groups’ squared effects, or the excess of one variance-like quantity over another — has Hessian diag(2, −2) exactly. A correlation-like quantity written as a ratio of a covariance to a product of spreads is stationary where the covariance is zero, and its Hessian there is indefinite. And any quantity built as one thing minus another thing, each of which is itself at a minimum, inherits a saddle from the subtraction.
That last shape is the common one and it is worth spelling out: a study comparing two quantities that are each estimated well near their own minima gets a difference whose second-order behaviour is a saddle’s, and the bias in each disappears from the difference while the coverage problem in each does not. Differencing removes the bias and keeps the failure, which is exactly the wrong half to keep and is the reason this case deserves an essay.
Where the rule about the side goes
The single-variable essay’s other finding was that nearly every miss falls on the same side. That one survives in a weaker form, and the weakening is worth stating precisely because it is easy to over-claim.
Near the stationary point the misses do concentrate on one side, for both the bowl and the saddle: at half a standard error out, 14.6% of the bowl’s misses and 5.0% of the saddle’s are above the truth, so most are below in both. The concentration is a property of the interval being short where the gradient is small rather than of the Hessian’s signature.
So the honest version is that the single-variable case found two things and attributed both to one cause. The bias is the Hessian’s trace and vanishes at a saddle. The one-sidedness of the misses is about the estimated standard error collapsing and does not. They looked like one finding because in one variable g″ controls both.
The dimension, and what it does not change
One question the three surfaces do not answer is whether any of this is about having two directions rather than one, so it is worth separating what the dimension does.
The law in one variable is , a single scaled on one degree of freedom, with mean and skewness — always positive when is, always negative when it is not. In two it is , which for a definite Hessian is still one-signed and is less skewed, because a sum of two chi-squares is closer to normal than one: the bowl’s skewness is 2.023 against a single ’s 2.828, and a bowl in ten dimensions would be closer still.
So the dimension does two things and only one of them is interesting. It makes the second-order law more normal, which is a quantitative softening of the one-variable problem. And it admits indefinite Hessians, which is a qualitative change and is the whole of this essay — a case with no counterpart in one variable, since a single second derivative is either positive, negative or zero.
What a defensible reading looks like
Report the estimated gradient with its own standard error. It costs nothing and it is the one reading that separates a study in this regime from one that is not. That is the same discipline a second number beside a promise asks for everywhere here: one figure describes what was estimated and another says whether the estimate means what it appears to.
Compute the trace before trusting an estimate near a stationary point. is a bias with a closed form, so it is correctable when the Hessian is known and quantifiable when it is estimated. It is also the number that says how large a sample the bias stops mattering at, which is a design question rather than an analysis one.
Treat a quantity built by subtraction with extra care. A difference of two things each at its own minimum has a saddle’s Hessian, so it is the shape most likely to arrive without anyone choosing it — and its bias cancels while its coverage does not, which is exactly the pattern a routine check would read as reassuring. It is the same asymmetry a robust standard error shows: the reassuring number and the informative number are not the same number.
Do not read a zero bias as a well-behaved case. The saddle is the worst of the three on coverage and the best on bias, which is the clearest available demonstration that the two properties are not substitutes. This is the same complaint an interval that covers and says nothing makes in a different field: one good property is not a summary.
Report the distance from the stationary point in standard errors. Every figure here is indexed by it, and it is the quantity that decides whether any of this applies: at five standard errors out the interval covers 94.90% and behaves like an ordinary one. A study that knows it is near a flat point has a different problem from one that is not, and the boundary is about half a standard error rather than a matter of degree.
And use a reference distribution matched to the law. The second-order limit is a weighted sum of chi-squares, not a normal, and reading a statistic against a normal at a flat point is exactly the error the first-order delta method makes one term earlier. Where the eigenvalues are known, the weighted sum’s quantiles are computable and an interval built from them behaves; where they are not, it is the estimation of a Hessian that has to be priced.
What the single-variable case would have predicted
Running the one-variable account forward gives three predictions, and the measurements sort them.
It predicts a large over-coverage at the point itself, because the interval collapses to nearly a point on the right side of the truth. That holds: 99.96% and 99.99%.
It predicts a coverage dip near the point, for the same reason one step out. That holds too: 93.93% and 91.52% at half a standard error.
And it predicts a bias with a fixed sign. That is the one that does not survive, and it is the one the one-variable case could not have distinguished from the others, because a single second derivative fixes all three at once. What the extra dimension supplies is a case where the third comes apart from the first two — which is the only way a conflated account can be taken apart, and is why the saddle earns an essay rather than a paragraph. The same move separates a coverage from a width elsewhere here: find the case where two properties that usually move together do not.
The two routes, and which is which
Every number in this essay exists twice, and the pairing is worth naming because it is what makes a claim about a limiting law checkable at all.
The closed forms come from the moments of a weighted sum of chi-squares: mean , variance , skewness . They involve no data, no sample size and no simulation — only three power sums of the Hessian’s eigenvalues.
The measurements come from forty thousand draws at each surface, with the estimate computed the way a practitioner would compute it and no expansion used anywhere.
They agree to 0.055 at worst across nine comparisons. That is what licenses the essay’s central claim to be stated as an identity rather than as a fit: the law at a stationary point is the Hessian’s eigenvalues, and the simulation is not evidence for a pattern but a check on an algebraic derivation.
What a study would see
Nothing above is available to a study, and it is worth saying what is, because the gap decides whether this is a caution or a method.
A study has an estimate of g and a standard error computed from an estimated gradient. It does not know that it is near a stationary point, does not know the Hessian, and has no diagnostic that fires. What it can see is a symptom: the estimated gradient is small relative to its own sampling error, which is a comparison the data supports and which nothing in a standard output prints.
That symptom is the same one a weak instrument produces — a first-stage coefficient small against its own noise, with everything downstream inheriting a division by something near zero — and the responses available are the same two: a reference distribution built for the case, or a statement that the quantity is not identified at this precision.
What differs is that a weak instrument has a conventional diagnostic and a flat point does not. An F statistic on a first stage is printed by default; the ratio of an estimated gradient to its own standard error is not printed by anything, and a reader has no way to ask whether the interval in front of them was built on a tangent plane that exists.
What is claimed here and what is not
Two dimensions, equal variances, no correlation. The means are independent with a common standard error, so H’s eigenvalues and the law’s weights are the same numbers. With correlated means the weights are the eigenvalues of H times the covariance matrix rather than of H alone, which changes every number here and none of the structure. That generalisation is stated rather than measured.
The Hessian is known. Every closed form above uses the true eigenvalues, and a study would estimate them from the same data that produced the estimate. The estimation error is not priced here and it is not obviously small: a Hessian estimated near a stationary point is estimated where the function is flattest, which is where curvature is hardest to see.
The saddle’s zero is exact by construction. The eigenvalues are 2 and −2 because the surface was written that way, so the trace is zero identically rather than approximately. A saddle whose eigenvalues do not cancel exactly has a bias equal to half whatever they sum to, and the symmetry of its law goes with it — the measured skewness of −0.017 is a simulation reading of a quantity that is zero only because the surface is symmetric.
The surfaces are chosen, not found. A bowl, a valley and a saddle are the three cases a two-by-two Hessian admits up to scaling, so the set is complete for two dimensions — and it is not a sample of anything. What is claimed is that the law’s moments are the eigenvalues, which is algebra, and that the coverage behaves as measured at these three, which is three readings.
Forty observations throughout. The figures are at n = 40, which is where the second-order term is visible against the first. At a much larger n the region in which the delta interval misbehaves shrinks — it is the region within about a standard error of the stationary point, and a standard error shrinks — so the failure becomes narrower without becoming milder. That is the same behaviour the single-variable case reported and it is reproduced rather than re-derived here.
And the coverage figures at the stationary point are not a success. 99.96% is not a good interval; it is an interval so wrong about its own width that it contains the truth by accident. The single-variable case established that and this one reproduces it, and both numbers should be read as the failure they are.
The argument in one line
A function’s stationary point is where the first-order expansion contributes nothing, so the second order is the whole of the behaviour — and in one variable the second order is one number whose sign decides everything. In several it is a matrix, and a matrix has a sign only when its eigenvalues agree.
Everything above follows from that sentence. The bias is half the trace, so it cancels where they cancel. The skewness is a ratio of power sums, so it is zero where the law is a difference of chi-squares. The coverage failure is about the estimated gradient and is indifferent to all of it, which is why a saddle can have no bias and the worst interval of the three. And the single-variable essay’s two findings looked like one because a single second derivative controls both quantities at once, which is an accident of the dimension rather than a fact about flat points — the same kind of accident a design effect turns out to be in a crossed study, where one number summarises a design only when the design is simple enough to have one.
Still open: the Hessian as an estimated quantity
Everything above treats the curvature as known and asks what it implies. A study has neither the Hessian nor the knowledge that it is at a stationary point, and both have to be read off the same data.
The two problems interact in an unhelpful way. Deciding whether the gradient is zero is a test whose power is worst exactly at the point where the answer matters, and estimating a Hessian is estimating second derivatives, which needs more data than estimating first ones. So the procedure a practitioner would need — test for flatness, and if flat, estimate the curvature and use the weighted- reference — is a two-stage procedure whose first stage is weak and whose second is expensive.
What that costs, whether the two-stage procedure’s coverage beats simply using the first-order interval and accepting the failure, and whether the selection in the first stage biases the second, are the questions these essays have not answered. They are the same shape as the charge a searched break has to pay, arriving as a question about curvature rather than about a change point.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A block size that changes — both name chi-square, confidence interval, coverage, monte carlo
- A coverage table with its own error — both name closed form, confidence interval, coverage, monte carlo
- A simulation that stops when it looks settled — both name closed form, confidence interval, coverage, monte carlo
- A width rule on skewed outcomes — both name closed form, coverage, monte carlo, skewness
- An interval that carries its scale — both name closed form, confidence interval, coverage, monte carlo
- Intervals for the findings — both name closed form, confidence interval, coverage, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Asymptotic biasCentral limit theoremChi-squareClosed formConfidence intervalCoverageDelta methodEigenvalueMonte CarloNon-central χ²Saddle pointSecond-order delta methodSkewnessStationary point