The tail past the last observation

The clustering the tail has

Every threshold method counts exceedances as though they were independent pieces of information, and in a dependent series they arrive in clusters. Ignoring that overstates a return level by the reciprocal of the extremal index — ×3.527 counted where the mean cluster holds four — and leaves a reported standard error 2.151 times too small.

Worth reading first: Three shapes, one limit.

Four thousand steps of a stationary series, a threshold at its 0.98 quantile, and eighty readings above it. Every closed-form standard error in a threshold analysis is computed from that eighty. At the strongest dependence measured here the eighty exceedances are about 20.3 independent pieces of information, and the standard error that counts them all reports 2.151 times less spread than the estimate actually has.

The level is worse. Treating exceedances as independent when they are not overstates a return level by the reciprocal of the extremal index — in closed form, with nothing approximated in the ratio — and counted against the process’s own block maxima it comes out at ×3.527 where the mean cluster holds four. That is not a correction of a few per cent applied to a number already carrying an error bar. It is a factor.

Both of those are consequences of one quantity, and this essay is mostly about whether that quantity can be estimated. The answer is that it can be estimated twice, that neither estimator is unbiased, and that the two are biased in opposite directions — which is a more useful thing to know than either of them being right.

A process whose answer is arithmetic

The dependence used here is a max-autoregression,

Xt=max(αXt1,  (1α)Zt)X_t = \max\left(\alpha X_{t-1},\; (1 - \alpha) Z_t\right)

with ZZ unit Fréchet innovations. It is stationary, its marginal is unit Fréchet, and its extremal index is θ=1α\theta = 1 - \alpha exactly. That last fact is why it is the process here rather than a more familiar one. Every counted cluster size below is scored against a number that is arithmetic, and a dependence whose θ\theta had to be simulated could not tell an estimator’s error from the simulation’s — which is the whole discipline this collection runs on.

The extremal index is the reciprocal of the mean cluster size, and it is the factor by which dependence reduces the number of independent extremes. A series with θ=0.25\theta = 0.25 has exceedances arriving in episodes of four, so eighty of them are twenty episodes, and twenty is what every count in the analysis should have been.

The exceedances arrive together. 300 steps of a max-autoregression with dependence 0.75, drawn on a logarithmic scale because its marginal has no variance. The rule marks the 0.9 quantile: 30 of the 300 readings are above it and they fall into 5 clusters, the largest holding 11. The mean cluster holds 6.000, and its reciprocal — 0.167 — is the runs estimator of the extremal index, whose true value for this process is exactly 1 − 0.75 = 0.25. Every threshold method in the collection assumes exceedances are independent pieces of information; here 30 of them are 5.
Fig. 1 Three hundred steps of the process at its strongest dependence, on a logarithmic scale because its marginal has no variance. Thirty of the three hundred readings sit above the 0.9 quantile, and they fall into five clusters, the largest holding eleven.

In one realised series at α=0.75\alpha = 0.75, thirty exceedances fall into five clusters and the largest holds eleven. A threshold analysis handed that series records thirty exceedances and computes every error bar from thirty. What it has is five episodes.

Without dependence they arrive alone. 300 steps of a max-autoregression with dependence 0, drawn on a logarithmic scale because its marginal has no variance. The rule marks the 0.9 quantile: 30 of the 300 readings are above it and they fall into 26 clusters, the largest holding 2. The mean cluster holds 1.154, and its reciprocal — 0.867 — is the runs estimator of the extremal index, whose true value for this process is exactly 1 − 0 = 1.00. Every threshold method in the collection assumes exceedances are independent pieces of information; here 30 of them are 26.
Fig. 2 The same process with the dependence set to zero. The same thirty exceedances now fall into twenty-six clusters, the largest holding two, which is what independence looks like at this threshold and this run length.

With the dependence removed the same thirty exceedances become twenty-six clusters, the largest holding two. That is the control the eye needs, and it also shows the first estimator’s problem: twenty-six rather than thirty, because two independent exceedances that happen to land within four steps of each other are counted as one episode.

Two estimators, scored rather than compared

The runs estimator declusters. Exceedances separated by four or more consecutive non-exceedances belong to different clusters, and the estimate is the number of clusters divided by the number of exceedances. It carries a constant nobody derives — the run length of four — which is chosen and not estimated.

The intervals estimator of Ferro and Segers reads the gaps between exceedances instead and is closed form in the inter-exceedance times. It needs no declustering constant at all. So the two are genuinely different routes: one with an arbitrary constant in it and one without.

Two estimators of one index, missing in opposite directions. Two estimators of the extremal index against the value the process actually has, over 300 records of 4000 steps at the 0.98 threshold. The process is a max-autoregression whose extremal index is 1 − α exactly, which is what makes a comparison of two estimators a statement about them rather than about a simulation. The runs estimator declusters with a gap of 4 and sits below the truth at 4 of the 5 dependent settings — 0.2540 against 0.25 at the strongest. The intervals estimator needs no gap at all and sits above it at all 5, 0.2701 at the same setting. Under independence the runs estimator reports 0.9396 rather than one, because two independent exceedances inside 4 steps are counted as one cluster.
Fig. 3 Two estimators of the extremal index against the value the process actually has, over three hundred records of four thousand steps. The runs estimator sits below the truth at four of the five dependent settings and the intervals estimator above it at all five.

Against a truth of 1, 0.8, 0.6, 0.5, 0.4, 0.25, the runs estimator reads 0.9396, 0.7632, 0.5833, 0.4890, 0.3952, 0.2540 and the intervals estimator reads 0.9683, 0.8147, 0.6244, 0.5247, 0.4192, 0.2701. The runs estimator is low at four of the five dependent settings; the intervals estimator is high at all five.

Both errors have the same source and it is the run-length rule, even in the estimator that does not use one. Under independence the runs estimator reports 0.9396 rather than one, because a gap of four is not long enough to separate two genuinely independent exceedances every time — the estimator is measuring the rule as well as the process. The intervals estimator is not fooled by that and overshoots instead, at every dependence, which is the direction its own bias correction pushes at these run lengths.

Neither is right, and knowing they bracket the truth is worth more than either number. It is the same reason a statistic read against the wrong one of its three distributions calls unrelated random walks cointegrated: a number is only usable alongside a statement of what it is a number about, and here the statement is a direction. An analyst with both in hand has an interval; an analyst with one has a point estimate whose direction of error is unknown to them.

What a level costs when the clusters are ignored

What a level costs when the clusters are ignored. The factor by which the 20-period return level is overstated when the exceedances are treated as independent, against the mean cluster size, over 300 records of 4000 steps. For this process the overstatement is exactly 1/θ in closed form, and the note beside each bar carries that number. Counted, it runs from 0.995 under independence to 3.527 where the mean cluster holds 4.086 — short of the closed 4.000, because a block of 200 steps is not yet the limit the closed form describes. The second cost is quieter: the standard error a threshold fit reports from all 80.0 exceedances is 2.151 times too small there, because those exceedances are about 20.3 independent pieces of information.
Fig. 4 The factor by which a twenty-period return level is overstated when the exceedances are treated as independent, against the mean cluster size. Counted, it runs from ×0.995 under independence to ×3.527 at the strongest dependence, against closed forms of ×1 through ×4.

For a stationary series, P(max of nu)exp{nθ(1F(u))}P(\max \text{ of } n \le u) \approx \exp\{-n\theta(1 - F(u))\}, so the level exceeded once per TT periods solves nθ(1F(u))=log(11/T)n\theta(1 - F(u)) = -\log(1 - 1/T). An analysis that treats the exceedances as independent sets θ=1\theta = 1, demands a smaller exceedance probability, and therefore quotes a higher level. For a unit Fréchet marginal 1F(u)1/u1 - F(u) \approx 1/u, so the whole thing collapses to a level of nθ/cn\theta/c and the overstatement is exactly 1/θ1/\theta with nothing asymptotic left in the ratio.

Counted against the process’s own two-hundred-step block maxima, the overstatement runs ×0.995, ×1.195, ×1.598, ×1.854, ×2.293, ×3.527 against closed forms of ×1.000, ×1.250, ×1.667, ×2.000, ×2.500, ×4.000. In levels: at the strongest dependence the process’s own twenty-period level is 1105.5, the closed form with θ\theta in it says 974.8, and the analysis that ignores the clustering says 3899.1.

The counted factor falls short of the closed one, and by more as the clusters lengthen. That is a real discrepancy and it is reported rather than smoothed over: a block of two hundred steps is not the limit the closed form describes, and at θ=0.25\theta = 0.25 a block holds only about fifty independent episodes. The closed form is the asymptotic statement and the count is what a finite block delivers, and the gap between them widens exactly where the effective sample inside a block gets smallest. Both are on the figure for that reason.

The quieter cost, which is the standard error

The threshold buys accuracy and spends exceedances. The mean squared error of the estimated shape against the threshold, split into the square of its bias and its spread, over 600 records of 2000 readings from a a normal parent. At the 0.9 quantile 199 exceedances are left, the bias is -0.1708, the spread is 0.0701 and the total error is 0.0341. The bias falls as the threshold rises because the exceedances get closer to being generalised Pareto; the spread rises because there are fewer of them. The sum is smallest at the 0.925 quantile, at 0.0340, of which 80.6% is still bias — so even the best threshold on this grid is one where accuracy, not spread, is the binding constraint.
Fig. 5 The error of a generalised Pareto shape estimate against the threshold, split into bias and spread, for independent readings. Every closed-form spread on this figure is computed from a count of exceedances, which is the quantity dependence makes wrong.

A threshold analysis reports the shape’s standard error as (1+ξ)/k(1 + \xi)/\sqrt{k}, where kk is the number of exceedances. That formula is a statement about kk independent observations, and it is the same formula whether the exceedances arrived one at a time or eleven at a time.

Fit the shape to all eighty exceedances at the strongest dependence and it comes out at 0.8069 with a counted spread across records of 0.4346. The closed form computed from those eighty reports a spread 2.151 times smaller. Fit instead to the twenty cluster maxima and the shape is 0.9002 with a spread of 0.4898, and the closed form is now 1.133 times off — near enough. Under independence the naive figure is 1.083, which is where the formula is meant to sit.

So the cost of not declustering is not that the estimate is much worse; it is that the reported precision is a fiction. The naive shape of 0.8069 and the declustered 0.9002 are both estimating a shape that is exactly one, and neither is good. The difference is that one of them arrives with an error bar that is roughly right and the other arrives with one that is less than half the size it should be.

One number in that paragraph deserves separating out, because it says declustering does two things rather than one. Under independence the naive fit reads 0.9451 and the declustered fit 0.9431, against a shape of exactly one — so both understate by about a twentieth, and that shortfall is the threshold bias measured elsewhere in this field rather than anything to do with dependence. At the strongest dependence the naive fit falls to 0.8069 while the declustered one holds at 0.9002. Declustering therefore recovers most of the extra bias the clustering introduced, and leaves the bias that was there anyway. The error bar is the larger repair, but it is not the only one.

This is the same failure as a standard error that divides by the square root of a count of observations that repeat each other, where a fifty-point series at a lag-one correlation of 0.8 is worth about six independent observations and its 95% interval covers 47%. It is the same failure again as the threshold grid’s own top end, where a closed-form spread of 0.1429 sits against a counted 0.5268 at nineteen exceedances. In all three cases the arithmetic is correct and the nn it is applied to is not the nn the data have.

The count that does not move

There is one column in the sweep that is the same at every dependence, and it is the column an analyst sees.

The threshold is the 0.98 quantile of the record, so two per cent of four thousand steps are above it by construction: the mean number of exceedances is 80.0 at α=0\alpha = 0 and 80.0 at α=0.75\alpha = 0.75 and 80.0 at every setting in between. Dependence does not change how many readings clear a quantile of their own series. It changes only how they are arranged in time.

What moves is the number of episodes those eighty readings represent: 75.2, 61.1, 46.7, 39.1, 31.6, 20.3 cluster maxima across the six settings, which is 94.0%, 76.3%, 58.3%, 48.9%, 39.5% and 25.4% of the exceedances surviving. The mean cluster runs 1.065, 1.315, 1.729, 2.071, 2.580 and 4.086.

So the defect is invisible in exactly the way this collection keeps finding defects to be invisible. Two records with wildly different information content present an identical count, an identical threshold, an identical fitted shape to within a tenth, and identical error bars. Nothing on the face of the analysis distinguishes eighty independent exceedances from twenty episodes of four, and the quantity that does distinguish them has to be estimated separately and usually is not.

That is the same arithmetic as an effective sample size counted in draws rather than in steps, where a walk yields one independent draw every τ\tau steps and the cost is counted in the same unit as a hunt’s. Here the unit is exceedances and the divisor is the mean cluster size, and the effective share is exactly the runs estimator: 20.3 divided by 80.0 is 0.254, which is the estimator’s own reading at that setting.

The practical consequence is that the check has to be run deliberately. There is no threshold at which the clustering announces itself, because raising the threshold raises it for the clusters too — a higher quantile leaves fewer exceedances arranged in the same episodes, so the ratio survives while the count shrinks, which is the worst of both adjustments. How much dependence a block shares is a quantity somebody has to choose before any of the arithmetic downstream means anything, and the choice is not visible in the arithmetic.

What declustering costs

The exceedances arrive together. 300 steps of a max-autoregression with dependence 0.4, drawn on a logarithmic scale because its marginal has no variance. The rule marks the 0.9 quantile: 30 of the 300 readings are above it and they fall into 13 clusters, the largest holding 6. The mean cluster holds 2.308, and its reciprocal — 0.433 — is the runs estimator of the extremal index, whose true value for this process is exactly 1 − 0.4 = 0.60. Every threshold method in the collection assumes exceedances are independent pieces of information; here 30 of them are 13.
Fig. 6 The same process at an intermediate dependence. Thirty exceedances fall into thirteen clusters, the largest holding six, which is the regime where the choice between the two estimators matters most.

Declustering is not free and the price is the obvious one. At the strongest dependence, eighty exceedances become 20.3 cluster maxima — 25.4% of them survive — and the shape’s standard error rises by roughly 1/θ1/\sqrt{\theta} as a result. At α=0.4\alpha = 0.4 the survival is 58.3% and thirty exceedances in a drawn series become thirteen clusters with the largest holding six.

So the trade is a trade rather than a recommendation. Not declustering keeps four times the sample and reports an error that is half what it should be; declustering keeps a quarter of the sample and reports an error that is roughly right. Neither of those is the free option, and the second one is only better because a wrong error bar is worse than a wide one — which is a judgement rather than a measurement, and it is stated here as a judgement.

The one thing the measurements do settle is that the choice cannot be made by looking at how the fits come out. The naive shape at the strongest dependence is 0.8069 and the declustered one is 0.9002; both have spreads near 0.45, so the two fits are entirely consistent with each other and with a truth of one. Nothing in the pair of estimates says which analysis is sound. The difference is only visible in a quantity — the counted spread across three hundred records — that a single record does not contain.

The fit that had no heavy half

While this was being built, the generalised Pareto fit could not see this process at all, and it said nothing about it.

The one-dimensional search that fits the shape needs a bracket, and the bracket was built on the mean excess above the threshold. A generalised Pareto with ξ1\xi \ge 1 has no mean. A unit Fréchet tail has ξ=1\xi = 1 exactly, so the bracket of the form c/yˉc/\bar{y} collapsed toward zero at precisely the setting the shape was largest, the search returned whatever the smallest candidate on its grid gave, and the reported shape was 0.66 against a truth of 1. Nothing threw. The number was plausible, the figure drew, and every heavy tail the machinery was shown would have been understated in the same direction.

It is built on the median excess and on the shape estimate itself now, and the fit recovers 1.0, 0.5, 0.25, 0.0, −0.3, −0.5 and −0.9 correctly. The symptom of the old version was that a region of the parameter space was quietly absent — the same shape as a walk that reaches exactly half its admissible set and reports the two-sided p-value correctly for ever, and the same shape as the negative half of the same fit’s profile, which the essay that priced the threshold describes. Three defects in one field, all of them absences, none of them detectable by asking whether a number that is there is right.

Where this does not hold

One process, one dependence structure. The max-autoregression was chosen because its extremal index is arithmetic, and that is exactly what makes it unrepresentative: it produces clusters by carrying a single large value forward with decay, which is one of several ways a real series clusters. A Gaussian autoregression at high threshold has θ=1\theta = 1 despite being strongly dependent, and nothing measured here would show that. The claim is about what clustering costs when there is clustering, not about how often there is.

One threshold and one run length. Every count above is at the 0.98 quantile with a gap of four. The runs estimator is a function of that gap by construction and its bias under independence — 0.9396 rather than one — is the gap showing through. A longer gap would fix the independent case and lose short clusters everywhere else. Nothing here sweeps it, and a sweep over the gap is the obvious next measurement.

And the counted overstatement is a finite-block number. ×3.527 against a closed ×4.000 is a fifteen per cent gap that this essay attributes to a block of two hundred steps not being the limit. That attribution is consistent with the gap widening monotonically with the cluster length, and it is not proved: a longer block would settle it and would cost proportionally more.

How often the quoted level is above everything ever recorded. The share of 800 records of 50 blocks on which the fitted return level sits above the largest reading in the record itself. The largest of 50 block maxima is, by its plotting position, a 51-block event — so a 100-block level is being read 1.96 times past the longest event the record contains, and it lands above that reading 52.4% of the time. Below the record's own reach the picture is different: a 10-block level is above the largest reading on 0.0% of records. The note beside each bar is the median share of the rise from a typical block maximum to the quoted level that lies above anything observed — 28.0% at a thousand blocks.
Fig. 7 The share of records on which a fitted return level sits above the largest reading in the record itself — 52.4% at a hundred blocks. Every one of those extrapolations assumes the exceedances underneath it were independent.

The last thing worth putting beside all of this is what it multiplies. A hundred-block level read from fifty blocks sits above everything in the record on 52.4% of records, with an error that grows sixfold from ten blocks to a thousand and a shape that carries most of it. The clustering correction is a factor applied to that number, and it is a factor of up to four. So the two largest uncertainties in a threshold analysis of a dependent series are a shape estimated with a bias no record length removes, and a count of exceedances that is not a count of anything independent — and the second one is the one that never appears in the reported interval at all.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AutocorrelationClosed formCluster sizeDeclusteringEffective sample sizeExceedanceExtremal indexFréchet lawGeneralised ParetoIndependencePeaks over thresholdReturn levelShape parameterStandard errorStationarity