Testing for a flat point first
Worth reading first: The optimum is a ratio, and its interval is sometimes the whole line.
Where the derivative is zero showed what goes wrong when a standard error is read off a tangent line at a point where the tangent is flat, and a flat point with more than one direction replaced the single second derivative with a Hessian and its eigenvalues. Both treated the geometry as known. A study does not know whether it is at a stationary point. It has a sample mean, a gradient evaluated there that is never exactly zero, and a decision to make about which law to read its interval from.
The procedure a careful analyst would reach for has two stages. First, test whether the gradient is zero. If the test finds a gradient, the point is not flat, the tangent line is a fair approximation, and the delta method’s interval is used. If the test does not reject, the point may be flat, and the interval is read instead off the law that holds at a stationary point — a in one dimension, a weighted sum of variables in several. Each stage on its own is a recognised method. What has not been asked is what the pair does together.
The rule, written out for a squared mean
The squared mean is the case where everything can be computed exactly. In units of the standard error, the estimate of the mean is Z, normal with mean δ and variance one, and the quantity wanted is . Four intervals are in play.
The tangent interval is , with z = 1.96. It covers when |Z| lies in a band between and , and its worst coverage is 85.98%, at δ = 1.45, as the earlier essay found.
The second-order interval treats as exactly , which it is when δ = 0. Its 95% interval for is , floored at zero: the quantiles at 97.5% and 2.5%, subtracted. It covers exactly 95% at the flat point, and nowhere else.
The two-stage rule tests flatness with |Z| against a critical value c, uses the second-order interval when |Z| ≤ c, and the tangent interval otherwise. A 5% pretest has c = 1.96.
The squared exact interval takes the exact interval for the mean, Z ± z, and squares it — reporting every value of the square the interval for the mean contains. It covers at least 95% at every δ, because it contains whenever the interval for the mean contains δ.
Every one of these is a function of |Z| alone, so each covers when |Z| falls in a union of intervals, and its coverage is a sum of normal probabilities. Nothing below about one dimension is simulated.
The two-stage rule is closer to the label than the tangent in exactly one place: at the flat point itself, where the tangent over-covers — 99.99% at δ = 0 — and the rule covers 97.49%, most of its samples being declared flat and handed the interval calibrated for exactly that point. A step away it is worse. At δ = 0.5 it covers 65.78%, at δ = 1 52.17%, and its worst is 49.72% at δ = 1.96. It does not rejoin the tangent until δ = 3.39, and from there on the two rules are identical, for an exact reason: past that distance the tangent’s own band starts above 1.96, so every sample the pretest declares flat is one the tangent would have missed anyway.
Why the worst is about one half
The location of that worst is not a coincidence, and once it is seen the size follows. The worst sits at δ = 1.96, which is the pretest’s own critical value.
At that distance, Z is normal around 1.96, so it falls below the critical value about half the time. On those samples the pretest declares the point flat and the second-order interval is reported. Its upper end is , and the truth is . A sample declared flat has , so its upper end is below the truth — the interval cannot reach from any sample that chose it. On the other half the tangent interval is reported, and it covers nearly always. The rule’s coverage is one half times nearly nothing plus one half times nearly everything.
The figure separates the two branches, and each tells its own story. The pretest’s power is poor in exactly the region that matters: at δ = 1, where the second-order law is already badly wrong, it declares flatness 83.0% of the time, and at δ = 1.45 still 69.5%. And conditioning on that verdict makes the second-order interval worse than it is on its own. At δ = 1 it covers 44.93% of all samples, and 42.5% of the samples the pretest hands it. At δ = 2 it covers 34.22% of all samples and 0.0% of the ones the pretest hands it. The selection does not merely fail to protect the second stage; it chooses the samples on which the second stage is certain to miss.
The tangent branch is biased the other way. Among the samples the pretest declares steep, the tangent interval covers 99.5% at δ = 1.45, against 85.98% on all samples. The samples where the tangent fails are the ones with |Z| small — the estimate close to the flat point, where the plug-in slope is nearly zero and the interval nearly a point — and those are precisely the ones the pretest routes away from it. So the two-stage rule takes the tangent’s failures away from the tangent and gives them to an interval that fails on them worse.
That is the general shape, and it is the same one the charge a searched break has to pay turned up in a different place: a procedure chosen by a test on the same data inherits the test’s errors at the point where the test is least able to tell. A pre-test for a robust standard error recovered some of what it was meant to; this one recovers nothing, because its two branches fail on complementary samples.
No pretest level rescues it
The natural repair is to tune the pretest. A stricter test declares flatness more often and puts the threshold further out; a laxer one declares it less often and keeps the second-order interval for samples very close to zero.
It does not help. At every pretest level from 1% to 30% the rule’s worst is within two points of one half — 49.55% at 1%, 49.72% at 5%, 50.38% at 20%, 51.82% at 30% — because the argument above does not depend on which critical value is chosen: the worst moves to wherever the critical value is, and is about one half there. Laxer pretests move the worst inwards, where the second-order interval is less wrong and is used less often: 58.84% at a 50% pretest, 80.61% at 80%. From 90% upward the pretest almost never declares flatness, and the rule is the tangent interval with its 85.98%. The two-stage rule’s worst coverage climbs towards the tangent’s as the first stage is switched off, and never passes it.
The reason no level works is that the second-order law is right at one point. The reference describes when δ is exactly zero, and its interval is calibrated to that point alone. At δ = 2 it covers 34.22% of samples, and its worst over the range to δ = 12 is 8.22%. A pretest cannot pick out the one value where it holds, because no test distinguishes δ = 0 from δ = 0.3 with any power, and at δ = 0.3 the second-order interval is already 20 points short. Its usefulness lasts exactly as far as the test’s blindness does.
What a study would report
The units hide how ordinary these distances are. A trial of a hundred patients a side, estimating the square of a standardised difference in means — an effect size reported as a proportion of variance, or the squared distance between two groups — has a standard error for the difference of about 0.14 standard deviations. A true difference of 0.14 is δ = 1. That is a small effect, the kind a study would describe as “consistent with none”, and it is where the rule fails worst.
At that distance the pretest declares the point flat in 83.0% of studies. Each of those reports an interval for the squared effect whose upper end is its own estimate, in standard units, and in 42.5% of them that upper end lies above the truth. In the other 57.5% the interval sits entirely below the true squared effect, and its report reads as a confident statement that the effect is smaller than it is. The remaining 17.0% of studies, declared steep, report a tangent interval that covers almost every time. Put together, the study’s own procedure reports the truth about half the time, and the studies that miss do not know it: their pretest said there was nothing to see, and their interval agreed.
The crossing point has a closed form that makes the dependence on the pretest explicit. The two-stage rule and the tangent coincide once the tangent’s lower edge, , passes the critical value c, which happens at . With a 5% pretest c equals z and the crossing is = 3.39. A stricter pretest pushes it further out and widens the region where the rule can fail; a laxer one pulls it in and shrinks the region, at the cost of using the second-order interval less often where it was right. Nowhere along that dial does the region in which the rule fails disappear while the second-order interval is still being used.
In two dimensions, a bowl and a saddle
The lead into this asked about the Hessian, and two dimensions are where the second-order law stops being a single . The same rule can be run there: test stationarity with the squared length of the standardised mean against , use the law when the test does not reject, and the tangent otherwise. The law is for the bowl , and twice a product of two independent normals for the saddle , whose 97.5% point is 4.364. The coverages are counted, on 40,000 draws at each distance.
On the bowl the rule fails as it does in one dimension. Its worst, 59.5% at a distance of two, compares with the tangent’s worst of 93.7% — the tangent does better on a bowl than on the squared mean, because a two-dimensional estimate is rarely close to the point in both coordinates at once, so its plug-in gradient is rarely near zero.
The saddle is the first case anywhere in this field where the second-order law is a genuine improvement. Within about one and a half standard errors of the saddle point the two-stage rule covers more than the tangent: 96.7% against 94.3% at 1.5. The reason is the saddle’s law. It is symmetric about zero, since the two eigenvalues cancel, so its interval is centred on the estimate rather than shifted below it, and a symmetric interval centred on the estimate survives a small displacement of the truth in a way the interval does not. Given the pretest’s verdict of stationarity, the second-order interval still covers 95.8% at 1.5.
But the saddle’s improvement runs out where the pretest’s power is still weak. At 2.5 the rule covers 82.2% and at 3 80.1%, against the tangent’s 90.5% and 91.3%: given a verdict of stationarity at 3, the second-order interval covers 16.1%. So the rule trades a few points near the point for ten points further out, and its worst is below the tangent’s worst on both surfaces.
The interval that needs no test
There is an interval that never needed the decision. Take the exact 95% interval for the mean — a disc in two dimensions, an interval in one — and report the set of values the function takes on it. If the mean is in the disc, its image is in the reported set, so the reported set covers at least 95% at every distance from every stationary point, for any function at all. On the squared mean it covers between 95% and 97.50%; on the saddle between 98.5% and 100%; on the bowl it is 95% at the stationary point and between 97.9% and 98.9% elsewhere.
What it costs is width, and in one dimension the cost is exact. Whenever the squared interval is , the same width as the tangent’s . Only when the estimate is within 1.96 of the flat point does it become , wider by — and that is the same set of samples on which the tangent interval collapses.
The ratio is 1.272 at the flat point, 1.130 at δ = 1, 1.025 at δ = 2 and 1.000 by four. So the honest interval costs at most 27% more width, spends it precisely where the tangent was about to miss, and spends nothing elsewhere. The two-stage rule’s extra step buys a worst coverage of 49.72% in exchange for saving that width; the shortest interval is the one that misses for the same reason here as there.
What happened to the Hessian
The question this took up had two parts: the stationarity test, and the Hessian that would have to be estimated if the test said flat. For a known function of means, which is every case above, the second part never arrives. The Hessian of is a fixed matrix; nothing about it is estimated. For a function that is not quadratic, the Hessian is evaluated at the estimate rather than the truth, and that adds the third derivative’s effect, which is of smaller order near the point.
So the expensive half of the two-stage procedure is not where its failure comes from. The whole of the damage measured above comes from the first stage — from choosing a law with a test that cannot see the difference between the one point where the law is right and the neighbourhood where it is wrong. Estimating the curvature better would leave every number unchanged.
The case where the Hessian genuinely has to be estimated is a fitted surface, where the function itself comes from the data. That is the setting of the run that confirms it, where the height a fitted optimum predicts is a maximum over a random field, and its curvature is an estimate with a sampling distribution of its own.
What is claimed here and what the counts rest on
Every number about the squared mean is exact: each interval covers when |Z| lies in a union of intervals, and each coverage, each branch probability and each conditional coverage is a sum of normal probabilities. The worst coverages are searched over δ from 0 to 12 in steps of 0.005, which is fine enough that the reported minima are within a few thousandths of a point of the true ones. The width ratio is a quadrature in Z.
The two-dimensional numbers are counted, on 40,000 draws at each of nine distances along one axis, one stated seed each; their standard errors are about a tenth of a point near 95% and two tenths near 80%. Only one direction of approach to the stationary point is measured. On the saddle the direction matters — along the diagonal the function is flat to first order, not just at the point — and the two-stage rule’s behaviour there is not measured.
What is not claimed: that every two-stage procedure fails. A pretest whose two branches failed on different samples, rather than on complementary ones, could improve on both; the saddle within one and a half standard errors is an example of a law robust enough to survive being chosen. What is claimed is narrower: for the squared mean and the bowl, the test-then-choose rule is worse than the tangent at its worst for every pretest level, and the squared exact interval is better than both at every distance.
Still open: a pretest whose branches fail apart
The saddle shows that the rule can help when the second-order law is robust to a small displacement, and the squared mean shows it cannot when the law is one-sided. That suggests a different second stage: not the reference at the flat point, but an interval for that is exact at every δ — the inversion of the noncentral law of , which is what the squared exact interval approximately is — used only when the pretest does not reject, with the tangent elsewhere.
Such a rule’s branches would no longer fail on complementary samples, because its flat branch would be right throughout the region the pretest cannot rule out. Whether it then beats the squared exact interval used alone — in width, since both would cover — whether the pretest buys anything once the flat branch is exact, and whether the same construction exists on a saddle, where the noncentral law of a difference of squares has two parameters rather than one, are the measurements this leaves. The closed forms above extend to the first of them without simulation; a ratio whose interval has to be the whole line is the reminder that an exact interval is sometimes forced to be one nobody wants to report, and sums of anything the reminder that the normal law this all departs from was never the problem.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Intervals for the findings — both name closed form, confidence interval, coverage, interval width, monte carlo, selection effect
- A block at every starting row — both name closed form, confidence interval, coverage, interval width
- A block size that changes — both name chi-square, confidence interval, coverage, monte carlo
- A coverage table with its own error — both name closed form, confidence interval, coverage, monte carlo
- A schedule that reads the mean — both name confidence interval, coverage, monte carlo, selection effect
- A simulation that stops when it looks settled — both name closed form, confidence interval, coverage, monte carlo
Named objects
A flat tag is an object no other essay names yet.
Chi-squareClosed formConditional coverageConfidence intervalCoverageDelta methodEigenvalueInterval widthMonte CarloSaddle pointSecond-order delta methodSelection effectSelective inferenceStationary point