Two intervals for one return level
Worth reading first: Three shapes, one limit.
Three hundred records at each of four lengths, one generalised extreme value fit apiece, and two 95% intervals for the hundred-block level read off each fit. The level they are about is 4.0330 and it is known in closed form, so this is coverage of a number rather than agreement between two estimates.
The delta-method interval — the symmetric one, the one a package reports first — covers 80.3% at twenty-five blocks, 86.3% at fifty, 87.7% at a hundred and 89.0% at two hundred. The profile-likelihood interval, on the same fits of the same records, covers 94.7%, 94.7%, 94.0%, 94.7%. Eight times the data buys the symmetric interval 8.7 points and leaves it six short of its promise; the other one is at its promise from twenty-five blocks and stays there.
That much is a familiar shape here: an interval that says 95% is making a checkable statement about a procedure, and the one taught first covers 87.6%. What is not familiar is which end of it goes wrong. 99.24% of the symmetric interval’s misses are the whole interval sitting below the truth. At twenty-five blocks it misses low 19.7% of the time and high 0.0% — not one record in three hundred.
Two intervals from one likelihood
Both intervals come off the same three fitted parameters, so any difference between them is a difference between two ways of turning a likelihood into an interval and cannot be a difference between two samples.
The delta-method interval propagates the fit’s covariance through the return level’s own gradient. The gradient is written out rather than differenced —
with — because the interval’s failure and a differencing error would otherwise be tangled together, and the whole point is to know which is which. The variance is , the interval is the estimate plus and minus 1.96 standard errors, and it is symmetric by construction. Not by assumption, by construction: there is no arrangement of the arithmetic in which it comes out lopsided.
The profile-likelihood interval reparameterises. The return level replaces the location as a parameter, , the other two are maximised out at each candidate , and the interval is every whose deviance from the maximum is below the one-parameter chi-square cut of 3.841458820694124. It has no reason to be symmetric and it is not.
Each of its endpoints is found by doubling outward from the estimate until the deviance crosses the cut and then bisecting, with the two-parameter maximisation warm-started from the previous step. That is what makes counting the coverage of three hundred of them at four record lengths affordable at all; a linear outward walk would cost three times the maximisations for the same answer.
The counted coverage
The symmetric interval never covers. Not at any of the four lengths, and not close: six points short at its best. The usual defence of an asymptotic interval is that it is asymptotic, and the sweep is arranged so that defence can be tested rather than asserted — eight times the record, from twenty-five blocks to two hundred, is a substantial slice of what a real record could ever be. It buys 8.7 points and would need about twice as many again to arrive, on a trend that is visibly flattening.
The profile interval is at its stated level immediately. 94.7, 94.7, 94.0, 94.7 against a nominal 95, on three hundred records apiece — a counted 94.7% at that many records carries a standard error of 1.3 points, so all four readings are one standard error from nominal and none of them is distinguishable from it. There is no small-sample penalty here to report.
This is the same finding as four intervals for one dataset where the narrowest is the one that fails its stated level, and it has the same practical shape: the appealing interval and the honest interval are not the same interval, and the appealing one is appealing because of the property that makes it wrong.
Which endpoint fails
The expectation going in was that the lower endpoint would be the problem — that a symmetric interval on a right-skewed quantity would fail by not reaching low enough. It is the other way round, and the failure is almost entirely one-sided.
At twenty-five blocks the symmetric interval sits entirely below the truth on 19.7% of records and entirely above it on 0.0% of them. At two hundred blocks it is 10.7% and 0.3%. Across the whole sweep, 99.24% of its misses are the interval sitting short. The profile interval, by contrast, misses 3.7% low and 1.7% high at twenty-five blocks and 3.3% and 2.0% at two hundred — lopsided, but lopsided the way a 95% interval is allowed to be.
A two-sided coverage figure cannot tell those two apart, and the difference is the whole of what is wrong with the first one. An interval that misses 20% of the time by being too low everywhere and one that misses 20% of the time evenly are not the same instrument. The first is a systematic understatement of a level with a safety margin attached to it; the second is an interval that is merely too narrow. Only one of them is dangerous in the direction the quantity is used.
And it is worth being clear that “misses low” means the reported upper endpoint is below the truth. A record whose interval sits short does not merely have a lower bound that is too high; it has an upper bound that is too low, so the whole of the reported range is on the safe-looking side of a level that is in fact higher. Reading it as an engineering margin gives a margin that is not there.
Why the upper endpoint cannot reach
The reason is in the other interval’s shape. At twenty-five blocks the profile interval’s upper arm is 5.686 times its lower one; at fifty it is 3.256, at a hundred 2.229 and at two hundred 1.749. The likelihood for a far-out level is skewed, and it is skewed for a reason that is arithmetic rather than incidental: the shape enters the return level through . A symmetric change in becomes a multiplicative change in , with a long arm upward, and the arm lengthens with the period being read.
So the profile interval reaches 6.2154 at twenty-five blocks where the symmetric one reaches 4.8406. Both start from an estimate of about 4.01. The symmetric interval is not too narrow in the ordinary sense — it is the wrong shape, and forced onto a skewed likelihood it can only lose on one side. Widening it symmetrically would fix the coverage by adding length at the bottom, where it is not needed, which is why the failure survives a longer record: at two hundred blocks the two intervals are almost the same width and the asymmetry has fallen to 1.749, and that is exactly where the coverage gap has narrowed to six points.
That asymmetry is the same fact the essay that read a level with no data in it found from the other side, where the fitted level’s bias at a hundred blocks is −0.0144 against an error of 0.2779. A distribution skewed to the right has a mean near its target and a median below it, and a symmetric interval is centred on the estimate rather than on either.
The wider interval is also a moved interval
Calling the profile interval “wider” is the natural summary and it is not what the endpoints say.
At twenty-five blocks the symmetric interval runs from 3.1816 to 4.8406 and the profile interval runs from 3.5993 to 6.2154. The profile interval’s lower endpoint is four tenths higher than the symmetric one’s. It gives up 0.4177 at the bottom and takes 1.3748 at the top, for a net widening of about 0.96 — so more than half of what it does is a translation and not an inflation.
At two hundred blocks the same thing happens on a smaller scale: 3.7877 to 4.2441 against 3.8396 to 4.3245, with the lower endpoint again the higher of the two and the width only six per cent greater. At every record length in the sweep, the honest interval starts further up than the symmetric one.
That is the fact a width comparison hides, and it settles the obvious objection. If the symmetric interval merely needed to be wider, multiplying its half-width by a constant would fix the coverage, and the constant is easy to find: at twenty-five blocks it would need about 1.9. But that repair adds as much at the bottom as at the top, and the bottom is where the interval is already too generous — it would produce an interval covering 95% by being wrong in both directions at once rather than in one. The likelihood’s own interval instead moves, because it is reading a distribution that is skewed rather than a distribution that is broad, and an interval that inherits the shape of what it is estimating rather than a normal approximation to it is a distinction this collection has priced elsewhere.
What neither interval is an interval about
Both intervals condition on the family, and neither reports that it does.
Every number here assumes the record came from a generalised extreme value law and asks only where its parameters are. The one thing the fit is least sure of is which of the three families it is looking at: at a hundred blocks a light-tailed record’s shape interval excludes zero on 45.0% of records, and calls the wrong family every time it does, because naming the family from a light-tailed record gets worse as the record grows. So the interval quoted for a level is conditional on a determination the same fit could not make, and no part of the 94.7% is a statement about that.
This is not a defect that a better interval fixes; it is the boundary of what an interval is. What could be done is to say so, in the way a credible interval that averages over the population spread rather than substituting an estimate of it says so — 95.2% against a plug-in’s 78.8%, at 31% more width and a heavier tail. Nothing here does that. Both intervals are conditional, one of them is honest about its own likelihood and neither of them is honest about its own family, and the difference between 80.3% and 94.7% is entirely within the narrower question.
What the honest interval costs
The profile interval is not free, and the accounting is worth doing because the cost is where the symmetric interval’s appeal comes from.
At twenty-five blocks the mean widths are 1.6590 and 2.6161 — the profile interval is 58% wider. At two hundred blocks they are 0.4564 and 0.4849, six per cent wider. So the price of covering falls off fast: at the record lengths where the symmetric interval is worst it is much narrower, and at the lengths where it is least bad the two are nearly the same size.
That is not a trade to weigh up, because the narrow interval is not delivering 95% coverage at a smaller width. It is delivering 80.3% coverage, which is a different product. A reader who wanted an 80% interval could have one from the profile route at a smaller width still, and would then know what they had. The comparison between an interval’s width and the coverage it actually delivers is the only comparison that means anything, and taken that way the symmetric interval is not cheaper — it is mislabelled.
It is the same lesson as a normal interval used where the spread was estimated: the cheap object is not a smaller version of the right one, it is a different claim wearing the right one’s label.
The computational cost is real and is the reason the symmetric one is reported first. A delta interval is one inverted Hessian; a profile interval is about twenty-eight two-parameter maximisations over the whole record. That is the entire justification, and it is a justification from 1980s arithmetic being applied on machines that are five orders of magnitude faster.
What could have produced this without the claim being true
The counting could be too coarse. Three hundred records gives a standard error of 1.3 points on a counted 94.7% and 2.3 points on a counted 80.3%. The gap being described is fourteen points at twenty-five blocks and six at two hundred, so no reading here is inside its own noise, and the ordering holds at all four lengths independently.
The differenced Hessian could be wrong. The delta interval’s covariance comes from a central-difference approximation to the observed information, and a bad one would inflate or shrink the interval. But an inflated interval would cover better, and the failure is one-sided rather than two-sided: a covariance error scales both endpoints and cannot move 99.24% of the misses to one side. The gradient it is multiplied by is written out in closed form for exactly this reason.
The profile could be under-searched. Its endpoints are bracketed by nine doublings and located by eleven bisections, and its nuisance maximisation runs to fifty-five evaluations warm-started. An endpoint that had not converged outward would make the interval too narrow and its coverage too low — the direction that would have hidden the finding, not manufactured it. Records where either endpoint failed to bracket at all are dropped from both intervals’ counts, so neither interval is scored on a record the other one was not.
The truth could be the wrong truth. It is under the parent, computed and not fitted. If the intended target were instead the limit law’s own level, both intervals would be scored against a slightly different number and both would move together; the difference between them would not.
Where this does not hold
Every number above is measured on a normal parent, whose true shape is zero. That is not a neutral choice. A return level’s exponent sits on one side of one when is negative and the other when it is positive, so a Fréchet parent puts the skew of the likelihood on the other side of zero, and the direction of the symmetric interval’s failure may reverse with it. The machinery takes a parent option and would need only a rerun. It was not done: this is the most expensive sweep in the field and a second parent doubles it, and a claim about the direction of a failure on one parent is a claim about that parent.
The second limit is that the level being covered is itself an extrapolation. The hundred-block level of 4.0330 sits just above 4.0062, the largest reading an average fifty-block record contains. So this is coverage of a quantity that a record barely reaches, and the intervals are being asked to bracket something that is at the boundary of observation rather than well inside it.
Push the period out and the situation gets worse rather than better in a way the coverage figures do not show. At a thousand blocks the quoted level is above everything in the record on 99.5% of them, and the shape’s lever on the level is at its largest. Nothing here counts coverage at that period, and the reason is that a profile interval there is a wider and slower object whose endpoints frequently fail to bracket — which is itself a finding of a kind, and one this sweep is not designed to report.
The last thing to carry is where the skew comes from, because it is not a property of return levels. It is a property of a parameter that enters through an exponent, and the parameter is the one this field’s first measurement found is estimated with a bias that a longer record does not remove and a spread that a longer record does. Any quantity read off a fitted extreme-value law inherits both. The intervals differ because one of them is built to inherit the shape of the likelihood and the other is built to be symmetric, and only one of those descriptions is a description of the thing being estimated.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The maximum converges slowly — both name block maxima, closed form, generalised extreme value, return level, shape parameter
- A coverage table with its own error — both name closed form, confidence interval, coverage, standard error
- A flat point with more than one direction — both name closed form, confidence interval, coverage, delta method
- The assumption that identifies the mechanism — both name closed form, confidence interval, maximum likelihood, profile likelihood
- The clustering the tail has — both name closed form, return level, shape parameter, standard error
- The draws aimed at the tail — both name closed form, confidence interval, coverage, standard error
Named objects
A flat tag is an object no other essay names yet.
Block maximaClosed formConfidence intervalCoverageDelta methodExtrapolationGeneralised extreme valueInformation matrixLikelihoodMaximum likelihoodProfile likelihoodReturn levelSampling distributionShape parameterStandard error