The tail past the last observation

Where the extremal index matters to a return level

A return level from a declustered fit depends on how often clusters arrive, which is the extremal index estimated from the same record. For a level reached once in two periods of a twenty-period record the index's error is a quarter to four fifths of the level's variance, and an interval that leaves it out covers 59.0% when clusters are long; put in, 90.0%. For a level reached once in a hundred periods its share is under two and a half per cent and adding it changes no interval at all — what fails there is the shape, fitted from however many clusters the dependence left: 82.3% coverage from forty-six clusters, 65.7% from nineteen.

Worth reading first: Three shapes, one limit.

The clustering the tail has found that a threshold analysis which counts exceedances in a dependent series as independent overstates a return level by the reciprocal of the extremal index, and the run length a declustering chooses found that estimating the index needs a rule for where one cluster ends, and that a rule reading the run length off the data has the smallest worst error of those compared. Both essays measured the index. Neither measured what happens to a return level, and its interval, when the index inside it is the one the record estimated.

That essay named the gap in so many words: two intervals for one return level were priced on independent blocks, where the index is one and known, and an interval for a level with an estimated index inside it had not been measured anywhere in this field. It has an answer with two halves, and which half applies depends on how far past the record the level is.

A level built from clusters

The records are the earlier essays’ processes, chosen because their extremal index and their marginal law are both known exactly, so every return level has a closed form to be scored against. Four thousand observations from a unit Fréchet marginal, read as twenty periods of two hundred: an independent series, two max-autoregressions whose clusters are runs of consecutive exceedances with indices 0.6 and 0.25, and a max-moving-maximum whose cluster members arrive six steps apart, with index 0.4. For each, the level exceeded once in TT periods is exactly 200 θ/c200\,\theta / c with c=−log⁡(1−1/T)c = -\log(1 - 1/T).

The analysis is the one the earlier essay recommended. Take the exceedances of the 98th percentile, decluster them with the automatic rule — the run length read off the gaps between exceedances — fit a generalised Pareto to the cluster maxima, and count the clusters. The clusters arrive at a rate λ\lambda per observation, which is the extremal index times the exceedance rate, and with the clusters a Poisson stream the TT-period level is

zT=u+σξ[(200 λc)ξ−1].z_T = u + \frac{\sigma}{\xi}\left[\left(\frac{200\,\lambda}{c}\right)^{\xi} - 1\right].

Three numbers are estimated: the Pareto’s scale σ\sigma and shape ξ\xi, and the cluster rate λ\lambda. The first two come from the cluster maxima; the third is the cluster count over the record length, with a Poisson variance of 1/C1/C on its logarithm for CC clusters. The index enters the level only through λ\lambda.

Inside the record, the index is the uncertainty

The delta method splits the level’s variance into the part from (σ,ξ)(\sigma, \xi) and the part from λ\lambda, and the split depends on how far the level reaches. The gradient with respect to log⁡λ\log\lambda is σ(200λ/c)ξ\sigma (200\lambda/c)^{\xi}, while the gradient with respect to ξ\xi carries an extra factor of log⁡(200λ/c)\log(200\lambda/c), which grows as the return period does.

How much of a return level's uncertainty is the extremal index's, by how far past the record the level is. Delta-method variance of the level from a declustered peaks-over-threshold fit, averaged over 300 records of 4,000 observations. The cluster rate's share at 1.5, 2, 5, 20, 100, 1000 periods: independent, 0.374, 0.253, 0.084, 0.027, 0.012, 0.005; runs of exceedances, θ = 0.6, 0.558, 0.396, 0.141, 0.039, 0.015, 0.006; spaced clusters, θ = 0.4, 0.845, 0.562, 0.222, 0.057, 0.019, 0.007; runs of exceedances, θ = 0.25, 0.787, 0.788, 0.327, 0.083, 0.024, 0.008.
Fig. 1 The extremal index’s share of a return level’s delta-method variance, against the return period in periods of two hundred observations, for the four processes; three hundred records each, the same fits at every period. The record is twenty periods long.

For a level reached once in two periods — roughly the median of the period maximum, a level the record has seen ten times — the cluster rate carries a quarter of the variance for the independent series, 0.396 for the runs with index 0.6, 0.562 for the spaced clusters and 0.788 for the long runs with index 0.25. The fewer the clusters, the larger the share, because the Poisson error in a count of nineteen is larger than in a count of seventy-seven while the shape matters little at a level the data surround.

The arithmetic behind that is short. With a shape near one, the level is close to σ\sigma times the extrapolation factor a=200λ/ca = 200\lambda/c, so an error in log⁡λ\log\lambda moves the level by the same proportion at every return period: a cluster count wrong by ten per cent makes every level wrong by about ten per cent. An error in the shape moves it by aξa^{\xi}, which is a proportion of log⁡a\log a times the error, and log⁡a\log a grows with the return period — about 1.7 at two periods and about 6 at a hundred for the independent series. The index’s share of the variance therefore falls roughly as one over the square of that logarithm. It is not a statement that the index stops mattering. It is a statement that the shape starts mattering much more.

By five periods the shares are 0.084 to 0.327; by twenty, the record’s own length, 0.027 to 0.083; and at a hundred periods they are 0.012 to 0.024. Past the record the shape is almost everything, because a small error in ξ\xi is an error in the exponent applied to the extrapolation, and the extrapolation is the whole of the level.

Leaving the index out, and putting it in

How often four intervals for the 2-period level cover it, with the extremal index estimated from the record. 95% intervals over 300 records of 4,000 observations, the level exact for each process. independent: index treated as known 86.0%, index's error added 91.0%, added, on the log scale 94.3%; runs of exceedances, θ = 0.6: index treated as known 86.7%, index's error added 91.7%, added, on the log scale 94.7%; spaced clusters, θ = 0.4: index treated as known 75.3%, index's error added 88.7%, added, on the log scale 92.7%; runs of exceedances, θ = 0.25: index treated as known 59.0%, index's error added 90.0%, added, on the log scale 91.3%.
Fig. 2 How often three 95% intervals cover the exact two-period level: the delta interval with the cluster rate treated as known, with its Poisson error added, and with that error carried on the logarithm of the level. Three hundred records of each process.

An interval that treats the estimated cluster rate as known — which is what a fit that reports only the Pareto’s standard errors amounts to — covers the two-period level 86.0% of the time on the independent series and 86.7% on the short runs, and then falls with the index: 75.3% on the spaced clusters and 59.0% on the long runs. Adding the cluster count’s Poisson error lifts these to 91.0%, 91.7%, 88.7% and 90.0%. Carrying the same variance on the logarithm of the level, which keeps the interval positive and skews it the way the level’s sampling distribution is skewed, gives 94.3%, 94.7%, 92.7% and 91.3%.

So inside the record the estimated index is not a refinement. It is most of what the interval has to say, and leaving it out makes a 95% interval into a 59% one on exactly the processes where declustering was needed. The error is invisible from the fit’s own output, because the Pareto’s standard errors are computed from the cluster maxima and know nothing about how many clusters there were.

The levels inside a record are not academic. A median period maximum, a level exceeded in one period of three, the alarm threshold a monitoring system is set to trip a few times a year: these are the levels an operator sees reached, argues about and adjusts, and for each of them a declustered fit is the standard tool exactly because the series clusters. The fitted Pareto comes with standard errors for its scale and shape, and the cluster rate comes as a ratio of two counts with no standard error printed beside it. An interval assembled from what is printed is the first row of each group in the figure above.

The size of the miss also says which process an analyst should worry about. Where clusters are short and plentiful — the runs with index 0.6, forty-six clusters a record — leaving the index out costs five points. Where clusters are long and few — index 0.25, nineteen clusters — it costs thirty-one. The processes that most need declustering are the ones on which the declustered fit’s own output is most misleading about its precision, because the count it divides by is smallest there and its Poisson error largest.

Far past the record, it is not

How often four intervals for the 100-period level cover it, with the extremal index estimated from the record. 95% intervals over 300 records of 4,000 observations, the level exact for each process. independent: index treated as known 77.0%, index's error added 77.0%, added, on the log scale 89.7%, block bootstrap, all re-estimated 86.7%; runs of exceedances, θ = 0.6: index treated as known 82.3%, index's error added 82.7%, added, on the log scale 92.3%, block bootstrap, all re-estimated 89.0%; spaced clusters, θ = 0.4: index treated as known 74.3%, index's error added 74.3%, added, on the log scale 85.0%, block bootstrap, all re-estimated 87.3%; runs of exceedances, θ = 0.25: index treated as known 65.7%, index's error added 65.7%, added, on the log scale 79.3%, block bootstrap, all re-estimated 80.7%.
Fig. 3 How often four 95% intervals cover the exact hundred-period level: the three delta intervals and a block bootstrap that cuts the record into blocks of a hundred, resamples them, and repeats the threshold, the declustering, the fit and the count 199 times.

At a hundred periods the interval that treats the index as known covers 77.0%, 82.3%, 74.3% and 65.7% on the four processes. Adding the index’s error changes those to 77.0%, 82.7%, 74.3% and 65.7% — the same to a tenth of a point on three of them. A two-per-cent share of the variance widens an interval by one per cent.

What does move coverage at a hundred periods is the interval’s shape. On the log scale the delta interval covers 89.7%, 92.3%, 85.0% and 79.3%; the block bootstrap, which re-estimates the threshold, the declustering, the shape and the count on every resample, covers 86.7%, 89.0%, 87.3% and 80.7%. Both are wider and both are asymmetric. Neither reaches 95%, and their misses are still nearly all on one side: the interval sitting wholly below the truth, as the two intervals for one return level found for block maxima, where 99.24% of the symmetric interval’s misses were low.

Far past the record is where a level with no data in it found the estimate staying nearly unbiased while its error grew sixfold, and the same thing is happening here with clusters in place of blocks: nothing in the record is near a hundred-period level, so everything about it is borrowed from the shape.

The reason is the shape estimate. Over the three hundred records the median fitted shape is 0.951 for the independent series and 0.824 for the long runs, against the true value of one, and the median level is 0.769 and 0.558 of the truth. A shape fitted from a few dozen cluster maxima is biased low and skewed, and at a hundred periods its error is raised to the power of a large extrapolation. The interval that would fix that is a profile likelihood, which the earlier essay found covering near 95% on independent blocks; it is the shape’s interval that is missing, not the index’s.

The analysis that does not decluster

Everything above declustered first. The alternative an analysis might take — fit every exceedance as though independent, and use the exceedance rate where the cluster rate belongs — is the case the clustering the tail has priced as an overstatement of the level by 1/θ1/\theta. Read through intervals it has two faces as well.

At the two-period level it is simply wrong. Its median level is 1.643, 2.335 and 3.364 times the truth on the three clustered processes — close to the 1/θ1/\theta of 1.67, 2.5 and 4 — and its log-scale interval covers 42.3%, 6.0% and 2.0%, every miss above the truth. On the independent series, where there is nothing to decluster, it covers 94.7%, as the declustered interval does.

At the hundred-period level it is wrong in a less legible way. Its median level is 1.224, 1.415 and 0.976 times the truth on the clustered processes, because a shape fitted to every exceedance of a clustered series is fitted to runs of near-identical values and comes out distorted as well as the rate. Its coverage is 81.3%, 68.3% and 54.0%, with the misses split about evenly between the two sides. The declustered log-scale interval covers 92.3%, 85.0% and 79.3% on the same records. Declustering is worth ten to twenty-five points of coverage past the record and fifty to ninety inside it — and inside it the index’s estimation error, left out, gives back a large share of what declustering bought.

What the clustering costs past the record

So the extremal index’s estimation error is irrelevant to a hundred-period level, and the index itself is not. It acts through the number of clusters it leaves.

Coverage of the hundred-period level against the number of clusters a record leaves. One point a process: runs of exceedances, θ = 0.25, 19.3 clusters, covered 65.7% by index treated as known, 79.3% by log scale, 80.7% by block bootstrap; spaced clusters, θ = 0.4, 28.6 clusters, covered 74.3% by index treated as known, 85.0% by log scale, 87.3% by block bootstrap; runs of exceedances, θ = 0.6, 45.6 clusters, covered 82.3% by index treated as known, 92.3% by log scale, 89.0% by block bootstrap; independent, 77.3 clusters, covered 77.0% by index treated as known, 89.7% by log scale, 86.7% by block bootstrap.
Fig. 4 Coverage of the hundred-period level against the average number of cluster maxima a record of four thousand observations leaves, one point a process, for the delta interval with the index treated as known, the same on the log scale, and the block bootstrap.

The four processes leave, on average, 77.3, 45.6, 28.6 and 19.3 cluster maxima from the same eighty exceedances. Every interval’s coverage falls with that count. The delta interval goes from 82.3% at forty-six clusters to 65.7% at nineteen; on the log scale from 92.3% to 79.3%. (The independent series sits below the trend on the delta interval, at 77.0% from seventy-seven clusters, because its clusters arrive most often: the factor 200λ/c200\lambda/c that the shape is raised against is about 380 for it and about 230 for the runs with index 0.6, so the same hundred periods is a longer extrapolation in clusters.) The shape has to be fitted from the clusters, and a process with index 0.25 has given a quarter of its exceedances’ worth of information about the tail.

That is the same arithmetic declustering costs sample size set out for the shape’s standard error, now read as coverage. A long record of a strongly clustered process is a short record of independent extremes, and its hundred-period interval is as untrustworthy as one from nineteen independent peaks would be — whatever the index estimate’s own precision.

Which interval to report

Inside the record, add the index’s error. For any level the record has seen — the median period maximum, a level reached once in a few periods — the cluster count’s Poisson error is a quarter or more of the uncertainty, and on strongly clustered series most of it. One term in the delta method, (∂z/∂log⁡λ)2/C(\partial z/\partial \log\lambda)^2 / C, takes a two-period interval from 59% to 90% coverage, and the log scale takes it nearer 95%.

Past the record, spend the effort on the shape. At a hundred periods the index’s term is one or two per cent of the variance. The interval that is wrong there is wrong because a symmetric interval in a skewed, extrapolated quantity misses low; a log-scale or bootstrap interval moves coverage by eight to fifteen points and the index’s term by none.

Never let the exceedance rate stand in for the cluster rate. On these processes it costs most of an interval’s coverage at a level the record has seen and a fifth to two fifths of it far past the record, and its misses at the short end all fall on the side a design standard cares about least: it says the level is higher than it is, so it is the error that wastes money rather than the one that floods a town — which is no defence of quoting it.

Report the cluster count beside the level. It is the number that says how much tail the record contains. Two series with the same length and the same exceedance count can carry seventy-seven clusters or nineteen, and a reader shown only the fitted level cannot tell a level read from a long record from one read from a short one.

Decluster with a rule that reads the data, and carry its choice. The bootstrap here re-ran the automatic rule on every resample, so its interval includes the run length’s own variability — which the run length a declustering chooses showed can be the difference between an index of 0.91 and one of 0.36 on spaced clusters. On those clusters the bootstrap covers 87.3% where the delta interval on the log scale covers 85.0%: the run length’s variability is small, and on the only process where it could have been large it was worth two points.

Counted against closed forms

Every coverage is over three hundred records of four thousand observations per process, scored against the exact level 200θ/c200\theta/c. The two-period sweep uses the same records as the hundred-period one, with no bootstrap. The bootstrap resamples forty blocks of a hundred consecutive observations, 199 times a record. The estimated index averages 0.966, 0.570, 0.359 and 0.241 on the four processes against 1, 0.6, 0.4 and 0.25 — the automatic rule reads every process slightly low.

Not claimed: that the Poisson variance of the cluster count is the whole of the index’s uncertainty. It ignores the run length’s own variability and the threshold’s, both of which the bootstrap carries; on these processes the difference between the two at a two-period level was within a few points. Not measured either: the profile-likelihood interval for a level built from cluster maxima, which is the obvious repair for the shape and has not been built in this field for a level with a cluster rate inside it.

Still open: a profile interval with a cluster rate inside it

The essay on block maxima found the profile-likelihood interval covering near 95% where the delta interval covered 80%. A level from cluster maxima has three parameters rather than three of a different kind — the Pareto’s two and the cluster rate — and its profile likelihood can be written with the rate’s Poisson likelihood beside the Pareto’s, maximised over the other two at each candidate level.

Whether that profile interval covers the hundred-period level near 95% from nineteen clusters, whether its asymmetry is enough to stop the misses falling low, and how its width compares with the bootstrap’s at the same coverage, are measurable on the same four processes and have not been measured. They decide whether a clustered record can give an honest interval far past itself at all, or whether nineteen clusters are simply too few.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Confidence intervalCoverageDeclusteringDelta methodEffective sample sizeExtremal indexGeneralised ParetoPeaks over thresholdReturn levelReturn period