Where the extremal index matters to a return level
Worth reading first: Three shapes, one limit.
The clustering the tail has found that a threshold analysis which counts exceedances in a dependent series as independent overstates a return level by the reciprocal of the extremal index, and the run length a declustering chooses found that estimating the index needs a rule for where one cluster ends, and that a rule reading the run length off the data has the smallest worst error of those compared. Both essays measured the index. Neither measured what happens to a return level, and its interval, when the index inside it is the one the record estimated.
That essay named the gap in so many words: two intervals for one return level were priced on independent blocks, where the index is one and known, and an interval for a level with an estimated index inside it had not been measured anywhere in this field. It has an answer with two halves, and which half applies depends on how far past the record the level is.
A level built from clusters
The records are the earlier essays’ processes, chosen because their extremal index and their marginal law are both known exactly, so every return level has a closed form to be scored against. Four thousand observations from a unit Fréchet marginal, read as twenty periods of two hundred: an independent series, two max-autoregressions whose clusters are runs of consecutive exceedances with indices 0.6 and 0.25, and a max-moving-maximum whose cluster members arrive six steps apart, with index 0.4. For each, the level exceeded once in periods is exactly with .
The analysis is the one the earlier essay recommended. Take the exceedances of the 98th percentile, decluster them with the automatic rule — the run length read off the gaps between exceedances — fit a generalised Pareto to the cluster maxima, and count the clusters. The clusters arrive at a rate per observation, which is the extremal index times the exceedance rate, and with the clusters a Poisson stream the -period level is
Three numbers are estimated: the Pareto’s scale and shape , and the cluster rate . The first two come from the cluster maxima; the third is the cluster count over the record length, with a Poisson variance of on its logarithm for clusters. The index enters the level only through .
Inside the record, the index is the uncertainty
The delta method splits the level’s variance into the part from and the part from , and the split depends on how far the level reaches. The gradient with respect to is , while the gradient with respect to carries an extra factor of , which grows as the return period does.
For a level reached once in two periods — roughly the median of the period maximum, a level the record has seen ten times — the cluster rate carries a quarter of the variance for the independent series, 0.396 for the runs with index 0.6, 0.562 for the spaced clusters and 0.788 for the long runs with index 0.25. The fewer the clusters, the larger the share, because the Poisson error in a count of nineteen is larger than in a count of seventy-seven while the shape matters little at a level the data surround.
The arithmetic behind that is short. With a shape near one, the level is close to times the extrapolation factor , so an error in moves the level by the same proportion at every return period: a cluster count wrong by ten per cent makes every level wrong by about ten per cent. An error in the shape moves it by , which is a proportion of times the error, and grows with the return period — about 1.7 at two periods and about 6 at a hundred for the independent series. The index’s share of the variance therefore falls roughly as one over the square of that logarithm. It is not a statement that the index stops mattering. It is a statement that the shape starts mattering much more.
By five periods the shares are 0.084 to 0.327; by twenty, the record’s own length, 0.027 to 0.083; and at a hundred periods they are 0.012 to 0.024. Past the record the shape is almost everything, because a small error in is an error in the exponent applied to the extrapolation, and the extrapolation is the whole of the level.
Leaving the index out, and putting it in
An interval that treats the estimated cluster rate as known — which is what a fit that reports only the Pareto’s standard errors amounts to — covers the two-period level 86.0% of the time on the independent series and 86.7% on the short runs, and then falls with the index: 75.3% on the spaced clusters and 59.0% on the long runs. Adding the cluster count’s Poisson error lifts these to 91.0%, 91.7%, 88.7% and 90.0%. Carrying the same variance on the logarithm of the level, which keeps the interval positive and skews it the way the level’s sampling distribution is skewed, gives 94.3%, 94.7%, 92.7% and 91.3%.
So inside the record the estimated index is not a refinement. It is most of what the interval has to say, and leaving it out makes a 95% interval into a 59% one on exactly the processes where declustering was needed. The error is invisible from the fit’s own output, because the Pareto’s standard errors are computed from the cluster maxima and know nothing about how many clusters there were.
The levels inside a record are not academic. A median period maximum, a level exceeded in one period of three, the alarm threshold a monitoring system is set to trip a few times a year: these are the levels an operator sees reached, argues about and adjusts, and for each of them a declustered fit is the standard tool exactly because the series clusters. The fitted Pareto comes with standard errors for its scale and shape, and the cluster rate comes as a ratio of two counts with no standard error printed beside it. An interval assembled from what is printed is the first row of each group in the figure above.
The size of the miss also says which process an analyst should worry about. Where clusters are short and plentiful — the runs with index 0.6, forty-six clusters a record — leaving the index out costs five points. Where clusters are long and few — index 0.25, nineteen clusters — it costs thirty-one. The processes that most need declustering are the ones on which the declustered fit’s own output is most misleading about its precision, because the count it divides by is smallest there and its Poisson error largest.
Far past the record, it is not
At a hundred periods the interval that treats the index as known covers 77.0%, 82.3%, 74.3% and 65.7% on the four processes. Adding the index’s error changes those to 77.0%, 82.7%, 74.3% and 65.7% — the same to a tenth of a point on three of them. A two-per-cent share of the variance widens an interval by one per cent.
What does move coverage at a hundred periods is the interval’s shape. On the log scale the delta interval covers 89.7%, 92.3%, 85.0% and 79.3%; the block bootstrap, which re-estimates the threshold, the declustering, the shape and the count on every resample, covers 86.7%, 89.0%, 87.3% and 80.7%. Both are wider and both are asymmetric. Neither reaches 95%, and their misses are still nearly all on one side: the interval sitting wholly below the truth, as the two intervals for one return level found for block maxima, where 99.24% of the symmetric interval’s misses were low.
Far past the record is where a level with no data in it found the estimate staying nearly unbiased while its error grew sixfold, and the same thing is happening here with clusters in place of blocks: nothing in the record is near a hundred-period level, so everything about it is borrowed from the shape.
The reason is the shape estimate. Over the three hundred records the median fitted shape is 0.951 for the independent series and 0.824 for the long runs, against the true value of one, and the median level is 0.769 and 0.558 of the truth. A shape fitted from a few dozen cluster maxima is biased low and skewed, and at a hundred periods its error is raised to the power of a large extrapolation. The interval that would fix that is a profile likelihood, which the earlier essay found covering near 95% on independent blocks; it is the shape’s interval that is missing, not the index’s.
The analysis that does not decluster
Everything above declustered first. The alternative an analysis might take — fit every exceedance as though independent, and use the exceedance rate where the cluster rate belongs — is the case the clustering the tail has priced as an overstatement of the level by . Read through intervals it has two faces as well.
At the two-period level it is simply wrong. Its median level is 1.643, 2.335 and 3.364 times the truth on the three clustered processes — close to the of 1.67, 2.5 and 4 — and its log-scale interval covers 42.3%, 6.0% and 2.0%, every miss above the truth. On the independent series, where there is nothing to decluster, it covers 94.7%, as the declustered interval does.
At the hundred-period level it is wrong in a less legible way. Its median level is 1.224, 1.415 and 0.976 times the truth on the clustered processes, because a shape fitted to every exceedance of a clustered series is fitted to runs of near-identical values and comes out distorted as well as the rate. Its coverage is 81.3%, 68.3% and 54.0%, with the misses split about evenly between the two sides. The declustered log-scale interval covers 92.3%, 85.0% and 79.3% on the same records. Declustering is worth ten to twenty-five points of coverage past the record and fifty to ninety inside it — and inside it the index’s estimation error, left out, gives back a large share of what declustering bought.
What the clustering costs past the record
So the extremal index’s estimation error is irrelevant to a hundred-period level, and the index itself is not. It acts through the number of clusters it leaves.
The four processes leave, on average, 77.3, 45.6, 28.6 and 19.3 cluster maxima from the same eighty exceedances. Every interval’s coverage falls with that count. The delta interval goes from 82.3% at forty-six clusters to 65.7% at nineteen; on the log scale from 92.3% to 79.3%. (The independent series sits below the trend on the delta interval, at 77.0% from seventy-seven clusters, because its clusters arrive most often: the factor that the shape is raised against is about 380 for it and about 230 for the runs with index 0.6, so the same hundred periods is a longer extrapolation in clusters.) The shape has to be fitted from the clusters, and a process with index 0.25 has given a quarter of its exceedances’ worth of information about the tail.
That is the same arithmetic declustering costs sample size set out for the shape’s standard error, now read as coverage. A long record of a strongly clustered process is a short record of independent extremes, and its hundred-period interval is as untrustworthy as one from nineteen independent peaks would be — whatever the index estimate’s own precision.
Which interval to report
Inside the record, add the index’s error. For any level the record has seen — the median period maximum, a level reached once in a few periods — the cluster count’s Poisson error is a quarter or more of the uncertainty, and on strongly clustered series most of it. One term in the delta method, , takes a two-period interval from 59% to 90% coverage, and the log scale takes it nearer 95%.
Past the record, spend the effort on the shape. At a hundred periods the index’s term is one or two per cent of the variance. The interval that is wrong there is wrong because a symmetric interval in a skewed, extrapolated quantity misses low; a log-scale or bootstrap interval moves coverage by eight to fifteen points and the index’s term by none.
Never let the exceedance rate stand in for the cluster rate. On these processes it costs most of an interval’s coverage at a level the record has seen and a fifth to two fifths of it far past the record, and its misses at the short end all fall on the side a design standard cares about least: it says the level is higher than it is, so it is the error that wastes money rather than the one that floods a town — which is no defence of quoting it.
Report the cluster count beside the level. It is the number that says how much tail the record contains. Two series with the same length and the same exceedance count can carry seventy-seven clusters or nineteen, and a reader shown only the fitted level cannot tell a level read from a long record from one read from a short one.
Decluster with a rule that reads the data, and carry its choice. The bootstrap here re-ran the automatic rule on every resample, so its interval includes the run length’s own variability — which the run length a declustering chooses showed can be the difference between an index of 0.91 and one of 0.36 on spaced clusters. On those clusters the bootstrap covers 87.3% where the delta interval on the log scale covers 85.0%: the run length’s variability is small, and on the only process where it could have been large it was worth two points.
Counted against closed forms
Every coverage is over three hundred records of four thousand observations per process, scored against the exact level . The two-period sweep uses the same records as the hundred-period one, with no bootstrap. The bootstrap resamples forty blocks of a hundred consecutive observations, 199 times a record. The estimated index averages 0.966, 0.570, 0.359 and 0.241 on the four processes against 1, 0.6, 0.4 and 0.25 — the automatic rule reads every process slightly low.
Not claimed: that the Poisson variance of the cluster count is the whole of the index’s uncertainty. It ignores the run length’s own variability and the threshold’s, both of which the bootstrap carries; on these processes the difference between the two at a two-period level was within a few points. Not measured either: the profile-likelihood interval for a level built from cluster maxima, which is the obvious repair for the shape and has not been built in this field for a level with a cluster rate inside it.
Still open: a profile interval with a cluster rate inside it
The essay on block maxima found the profile-likelihood interval covering near 95% where the delta interval covered 80%. A level from cluster maxima has three parameters rather than three of a different kind — the Pareto’s two and the cluster rate — and its profile likelihood can be written with the rate’s Poisson likelihood beside the Pareto’s, maximised over the other two at each candidate level.
Whether that profile interval covers the hundred-period level near 95% from nineteen clusters, whether its asymmetry is enough to stop the misses falling low, and how its width compares with the bootstrap’s at the same coverage, are measurable on the same four processes and have not been measured. They decide whether a clustered record can give an honest interval far past itself at all, or whether nineteen clusters are simply too few.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A flat point with more than one direction — both name confidence interval, coverage, delta method
- A ratio whose interval has to be the whole line — both name confidence interval, coverage, delta method
- A width promised for a difference — both name confidence interval, coverage, effective sample size
- Testing for a flat point first — both name confidence interval, coverage, delta method
- The count that is not the rows — both name confidence interval, coverage, effective sample size
- The draws aimed at the tail — both name confidence interval, coverage, effective sample size
Named objects
A flat tag is an object no other essay names yet.
Confidence intervalCoverageDeclusteringDelta methodEffective sample sizeExtremal indexGeneralised ParetoPeaks over thresholdReturn levelReturn period