The tail past the last observation

The run length a declustering chooses

The runs estimator of an extremal index carries a constant nobody derives. Where a cluster is a run of neighbouring exceedances the constant barely matters; where a cluster's members fall six steps apart, the estimate is 0.9069 at a run length of six and 0.3649 at seven against an index of 0.40, and a run length of four removes under a tenth of the overstatement declustering exists to remove. A rule that reads the run length off the data has the smallest worst error of the three.

Worth reading first: Three shapes, one limit.

A declustering needs a rule for where one cluster ends and the next begins, and the rule most analyses use is a single integer. Two exceedances that many steps apart or more belong to different clusters; closer than that, they are one episode. The essay that measured what clustering costs used a run length of four throughout and said so, and it ended on the measurement it had not made: the runs estimator is a function of that integer by construction, and nothing had swept it.

Swept, the integer matters in two entirely different ways, depending on what a cluster looks like. On a process whose clusters are runs of consecutive exceedances the estimate drifts gently across run lengths from two to forty, and four was a reasonable choice. On a process whose cluster members arrive six steps apart, the estimate at a run length of six is 0.9069 and at seven it is 0.3649, against an extremal index of exactly 0.40. There is no slope there to be roughly right on. There is a cliff, and four is on the wrong side of it.

What four costs on that process is nearly everything declustering is for. A return level quoted with no declustering at all is overstated by ×2.500; quoted with a run length of four it is overstated by ×2.356. The standard error reported for the fitted shape is 1.812 times too small at four and 1.095 at seven. None of that is visible in the record, and the run length that avoids it can be read off the data by a rule that never needed a run length in the first place.

A cluster whose members are not neighbours

The process here is a max-moving-maximum,

Xt=max(0.4Zt,  0.3Zt6,  0.3Zt12)X_t = \max\left(0.4\,Z_t,\; 0.3\,Z_{t-6},\; 0.3\,Z_{t-12}\right)

with ZZ independent unit Fréchet innovations. Its marginal is unit Fréchet, like the max-autoregression’s, and its extremal index is the largest weight divided by the sum of the weights, so θ=0.4\theta = 0.4 exactly. One large innovation appears three times — at four tenths of its size, then twice more at three tenths, six and twelve steps later — and a threshold high enough to be about extremes sees it as up to three exceedances with five ordinary readings between each. The shape is a familiar one in anything that echoes at a fixed delay: a surge that returns on the next tide, a load that comes back with the next shift.

It is chosen for the same reason as the max-autoregression was. Its index is arithmetic rather than simulated, so every estimate below is scored against a number with no error of its own, which is the only way a comparison of two estimators says something about the estimators rather than about the simulation that produced the reference.

A run length of 4 makes every exceedance its own cluster. 200 steps of a max-moving-maximum, X(t) = max(0.4·Z(t), 0.3·Z(t−6), 0.3·Z(t−12)) with unit Fréchet innovations Z, whose extremal index is exactly 0.40: one large innovation can put three readings above a threshold, 6 steps apart. The rule marks the 0.9 quantile and 19 readings clear it; 12 of the gaps between consecutive exceedances are exactly 6 steps. With a run length of 4, so that two exceedances 4 or more steps apart start separate clusters, they form 19 clusters, shaded, the largest holding 1. The runs estimator reads 1.000 against 0.40.
Fig. 1 Two hundred steps of the spaced process with its nineteen exceedances of the 0.9 quantile marked. With a run length of four no two of them are close enough to share a cluster, so the shading counts nineteen clusters of one and the runs estimator reads 1.000 against an index of 0.40.

In one realised stretch of two hundred steps, nineteen readings clear the 0.9 quantile and twelve of the gaps between consecutive exceedances are exactly six steps long. A run length of four sees nineteen clusters of one. The estimate is 1.000, which is what the estimator reports for a series with no dependence at all — so on this record a declustering at four has looked straight at episodes of three and reported independence.

That is not a small-sample accident and it is not the threshold. Every gap of six between members of one cluster is longer than four, so every member starts a new cluster however long the record is. The runs estimator can only merge exceedances that fall inside its run length, and a dependence whose echoes arrive later than that lies outside what the rule is able to see. More data make the estimate more precise and leave it exactly as wrong.

A run length of 7 finds 7 clusters in 19 exceedances. 200 steps of a max-moving-maximum, X(t) = max(0.4·Z(t), 0.3·Z(t−6), 0.3·Z(t−12)) with unit Fréchet innovations Z, whose extremal index is exactly 0.40: one large innovation can put three readings above a threshold, 6 steps apart. The rule marks the 0.9 quantile and 19 readings clear it; 12 of the gaps between consecutive exceedances are exactly 6 steps. With a run length of 7, so that two exceedances 7 or more steps apart start separate clusters, they form 7 clusters, shaded, the largest holding 3. The runs estimator reads 0.368 against 0.40.
Fig. 2 The same two hundred steps with a run length of seven. The nineteen exceedances now fall into seven clusters, the largest holding three, and the runs estimator reads 0.368 against 0.40.

With a run length of seven the same nineteen exceedances become seven clusters and the estimate is 0.368. The shading now spans the triples the eye picks out. Nothing about the record changed between the two pictures. Only the integer did, and the integer moved the estimate from the value for independence to within a few hundredths of the truth.

This stretch was picked because its clusters do not overlap, which is what makes the spacing legible, and that is a choice about what a picture can show rather than about what the estimator does. The counted behaviour below does not depend on it. At a high threshold, clusters from different innovations rarely overlap in any record; at a low one they often do, and that turns out to matter as much as the run length.

A slope on one process and a cliff on another

Four processes whose index is known exactly are run here: independent readings; two max-autoregressions, at θ=0.6\theta = 0.6 and θ=0.25\theta = 0.25, whose clusters are runs of consecutive exceedances; and the spaced process at 0.4. Each gets three hundred records of four thousand steps with the threshold at the 0.98 quantile, and the runs estimator is read on every record at fifteen run lengths from one to forty.

The run length is a cliff on one process and a slope on the others. The runs estimator of the extremal index at 15 run lengths from 1 to 40, each divided by the index its process actually has, over 300 records of 4000 steps at the 0.98 quantile. On the max-autoregressions, whose clusters are runs of consecutive exceedances, the estimate drifts: 0.2573 at a run length of two and 0.2121 at 40, where the truth is 0.25. On the max-moving-maximum, whose cluster members are 6 steps apart, it falls off a cliff: 0.9069 at six and 0.3649 at seven, against 0.40. At a run length of 4 it reads 0.9424, 2.36 times the truth. Without dependence the estimate is 0.9407 at four and 0.4573 at 40, because independent exceedances also land close together.
Fig. 3 The runs estimator at fifteen run lengths, each process divided by its own index. The consecutive-run processes drift slowly; the spaced process sits above twice its index up to a run length of six and falls through the truth between six and seven.

On the max-autoregression with the longest clusters the estimate is 0.2573 at a run length of two and 0.2121 at forty, against 0.25. That is the whole range of the dial, and across it the estimate moves by less than a fifth of its own value. The run length of four the clustering essay used sits within a hundredth of the truth on this curve, which is why the choice looked harmless there: on a process of that kind, it was.

On the spaced process the estimate is 0.9424 at four, 2.36 times the index, and it barely moves until the run length passes the spacing — 0.9069 at six, then 0.3649 at seven. Once past the cliff it drifts down like the others, because a longer run length also starts merging clusters that came from different innovations.

The independent process shows the dial’s other cost. It reads 0.9407 at four and 0.4573 at forty: with no dependence at all, a long run length merges exceedances that merely happened to land near one another. So a run length long enough to be safe on the spaced process is expensive wherever clusters are short, and a run length short enough for independent data is blind to the spaced process. No single integer is right for all four. The value at a run length of one, which declares every exceedance its own cluster, is 1/θ1/\theta on every dependent process by construction — the overstatement the clustering essay priced, now visible as the left edge of the same curve.

How the error presents matters as much as its size. At a run length of four the spaced process’s estimate has a root mean squared error of 0.5437, and 99.5% of its squared error is bias. The estimate is not noisy. It is precise, it is wrong, and its spread across records says nothing about either. A record read at four is indistinguishable from a record of a process with almost no clustering, and no amount of care with that one number reveals the difference.

What the wrong side of the cliff costs a fit

What a run length inside the spacing costs the fit. On the max-moving-maximum, the generalised Pareto fit to the cluster maxima at four run lengths, over 300 records of 4000 steps at the 0.98 quantile. The marginal is unit Fréchet, so the shape is exactly one and a return level is proportional to the extremal index an analysis uses: quoting it with a run length's estimate in place of 0.40 multiplies it by their ratio. At a run length of one, which is no declustering at all, the factor is ×2.500; at four it is ×2.356; at seven ×0.905. The closed-form standard error of the shape, (1 + ξ)/√k, treats the k cluster maxima as independent. At a run length of four the 75.1 maxima are not, and the counted spread of the shape is 1.812 times the reported one; at seven, with 28.9 maxima, it is 1.095.
Fig. 4 On the spaced process, the return level and the shape’s standard error at four run lengths. At a run length of one or four the level is more than doubled and the reported spread is little more than half the counted one; at seven and thirteen the spread comes within a tenth and the level tips slightly low.

The marginal is unit Fréchet, so the shape is exactly one and a return level is proportional to the extremal index an analysis puts into it — the closed form the clustering essay derived, with nothing asymptotic left in the ratio. Quote a level with an estimated index in place of 0.40 and the level is off by exactly their ratio. With no declustering the factor is ×2.500. With a run length of four it is ×2.356, which means declustering at four removed under a tenth of the overstatement it was run to remove.

At a run length of seven the factor is ×0.905 and at thirteen it is ×0.864. Past the cliff, the error changes sign. A run length long enough to merge the echoes of one innovation also merges some neighbouring innovations, which counts too few clusters and therefore quotes too low a level. For a level used to set the height of a defence that is the direction nobody budgets for. It is a tenth rather than a factor of two, but it is not zero, and it grows with the run length.

The standard error fails in step with the level. Fit a generalised Pareto to the cluster maxima and set the counted spread of its shape across records beside (1+ξ)/k(1 + \xi)/\sqrt{k}. At a run length of four each record yields 75.1 cluster maxima and the counted spread is 1.812 times the reported one — no better than fitting every exceedance, where the ratio is 1.807. At seven the maxima fall to 28.9 and the ratio to 1.095, which is where the formula is meant to sit.

Meanwhile the fitted shape barely moves: 0.9342 at four and 0.9708 at seven, against a truth of one. So an analyst who fitted at both run lengths would see nearly the same shape and an error bar far wider at seven, and the natural reading is that four is the more precise analysis. It is not more precise. It reports a precision the data do not contain, the same arithmetic as an interval that divides by a count of observations that repeat each other, with the count here being cluster maxima that are really echoes of one another.

Two routes that need no run length

The intervals estimator of Ferro and Segers reads the extremal index from the inter-exceedance times directly, with no declustering constant in it at all. On the spaced process it reads 0.4491 — high, as it was high at every dependence in the clustering essay’s sweep, but on the right side of a cliff it was never told about.

Their automatic declustering goes one step further. Turn the intervals estimate into a number of clusters, C=θ^(N1)+1C = \lfloor \hat\theta (N - 1) \rfloor + 1 from NN exceedances, and declare the C1C - 1 longest gaps between exceedances to be gaps between clusters. The run length is then an output rather than an input. Where gaps tie — and inter-exceedance times are integers, so they tie constantly — CC is lowered until the (C1)(C-1)-th longest gap is strictly longer than the CC-th, so the rule never has to split a set of equal gaps arbitrarily.

The run length the data choose. The shortest gap the automatic declustering rule treats as a gap between two clusters, record by record, over 300 records of 4000 steps at the 0.98 quantile; the line runs from the tenth to the ninetieth percentile and the dot marks the median. The rule estimates the extremal index with the intervals estimator, which needs no run length, turns that into a number of clusters, and declares that many of the largest gaps to be gaps between clusters. On the spaced process its median is 9 and nine records in ten choose 7 or more, past the 6-step spacing inside a cluster, and it estimates 0.3621. On the process with the longest runs the choice spreads from 3 to 48. With ties among the gaps left unresolved the same rule has a median of 6 on the spaced process and estimates 0.4498, because it splits clusters whose members are exactly 6 steps apart.
Fig. 5 The shortest gap the automatic rule treats as separating two clusters, from the tenth to the ninetieth percentile across records, with the median marked. On the spaced process nine choices in ten are seven or longer; on the consecutive-run processes the choice spreads widely. The last row is the same rule with ties left unresolved.

On the spaced process the median shortest gap the rule treats as a gap between clusters is 9, and nine records in ten choose 7 or more. The rule lands past the spacing on its own and estimates 0.3621. On the process with the longest runs the choice spreads from 3 to 48, which looks alarming and is not: across that entire range, the runs curve for that process moves by less than a fifth of its value. The rule is imprecise exactly where precision in the run length buys nothing and precise where it buys everything, which is what a rule that reads the clusters should do and what a convention cannot.

Why this works is worth stating exactly. The intervals estimator has no constant, so it measures how large the clusters are without having to know what shape they take. The automatic rule turns that size into a count and lets the gaps sort themselves: on the spaced process the gaps between clusters are long and the gaps inside a cluster are all six, so the longest C1C - 1 gaps are the right ones whenever CC is roughly right. It is the same move as a smoothing bandwidth read off the readings themselves instead of fixed in advance, and it has the same limitation: the rule is only as good as the quantity it reads, and here that quantity is an estimator already known to run high.

Judged by the process that suits each rule least

Three usable rules, judged on the process that suits each least. The root mean squared error of four ways to estimate the extremal index, on four processes whose index is known exactly, over 300 records of 4000 steps at the 0.98 quantile. A fixed run length of four has the smallest error of the three usable rules on 3 of the four processes, and an error of 0.5437 on the spaced one, 11.1 times the automatic rule's 0.0490. The largest error each rule gives on any of the four is 0.5437 for the fixed run length, 0.1004 for the intervals estimator and 0.0783 for the automatic rule. The last bar in each group is the run length that happened to do best, which only a process with a known answer can reveal.
Fig. 6 Root mean squared error of three usable rules, and of the run length that happened to do best, on each of the four processes. A fixed run length of four is the most accurate usable rule on three of them and has eleven times the automatic rule’s error on the fourth.

A fixed run length of four has the smallest error of the three usable rules on three of the four processes: 0.0635 on independent readings, 0.0594 and 0.0475 on the two max-autoregressions. On the spaced process its error is 0.5437, 11.1 times the automatic rule’s 0.0490. Judged instead by the worst process each rule meets, the fixed run length’s error is 0.5437, the intervals estimator’s 0.1004 and the automatic rule’s 0.0783.

The automatic rule beats the intervals estimator it is built from on all three dependent processes, and loses to it only on independent readings, 0.0773 against 0.0675. Where it is not the best usable rule it is never more than 1.32 times the best one’s error, and on the spaced process it comes within four thousandths of the run length that happened to do best — which was seven, a number only a process with a known answer can reveal.

So a fixed run length is a bet on the shape of the clusters. An analyst who knew their clusters were runs of neighbours would do well to make it: four is better than either alternative on three processes here. What decides between rules is what happens when the bet is wrong, and the answers are a factor of eleven against a factor of at most 1.32. That is the same kind of choice as how long a block a multiplier bootstrap shares, where a length set in advance walks the rejection rate straight through its nominal level. A constant that encodes an assumption about dependence is safe exactly as far as the assumption is.

It differs in one respect from a restricted-mean horizon chosen after looking. The automatic rule does look at the data, but at the gaps, not at the estimate an analyst hopes to report, and its choice is remade on every record here, so its variability is already inside the errors quoted above. Picking whichever run length made the level look most reasonable would not be.

The threshold moves the cliff

The right run length depends on the threshold as well. Four estimates of the spaced process's extremal index at three thresholds, over 300 records of 4000 steps each; its index is 0.40, as a limit, at every threshold. A run length of four reads 0.8564, 0.9424, 0.9726 at the 0.95, 0.98, 0.99 quantiles, more than double the truth at all three. A run length of seven reads 0.3029, 0.3649, 0.3907: at the 0.95 quantile the 199.7 exceedances are dense enough that clusters from different innovations fall within seven steps of each other and merge. The intervals estimator reads 0.4780, 0.4491, 0.4513 and the automatic rule 0.3029, 0.3621, 0.3784.
Fig. 7 Four estimates of the spaced process’s index at three thresholds, each with its error against 0.40. A run length of four is more than double the truth at every threshold; at the 0.95 quantile a run length of seven and the automatic rule both read 0.3029.

At the 0.95, 0.98 and 0.99 quantiles a run length of four reads 0.8564, 0.9424 and 0.9726. It gets worse as the threshold rises, because fewer exceedances means fewer accidental merges of neighbouring clusters to pull it down. A run length of seven reads 0.3029, 0.3649 and 0.3907, and the automatic rule reads 0.3029, 0.3621 and 0.3784 — low at every threshold and far low at the bottom one. The intervals estimator reads 0.4780, 0.4491 and 0.4513, high at all three.

At the 0.95 quantile one reading in twenty is an exceedance, 199.7 a record, and each innovation’s echoes span twelve steps. Clusters from different innovations overlap often enough that any declustering which sees the triples also joins some of them, and a run length of seven and the automatic rule arrive at exactly the same estimate. The extremal index is a limit as the threshold rises. At a threshold that low, the clustering a record displays is not the clustering the limit counts, and no choice of run length recovers it.

At the 0.99 quantile the problem inverts. There are 39.7 exceedances a record, the estimates sit closer to the index, and the automatic rule’s choice spreads from 8 to 60 with a median of 22. The data now contain too few gaps to say where clusters end. That is the dial every threshold analysis already has — bias bought down with exceedances — turning up a second time inside the declustering, and the run length does not escape it.

Agreement that was built in

The first version of the automatic rule here took the C1C - 1 longest gaps literally and did not resolve ties. On the spaced process it estimated 0.4498, against the intervals estimator’s 0.4491, and its median shortest gap between clusters was 6. On at least half the records it had declared a gap of six or less to be a gap between clusters, splitting the very episodes it was meant to find.

The two estimates agreed to within a thousandth, and that agreement looked like two methods confirming each other. It was worth nothing. Without the tie rule the automatic estimate is C/NC/N, and CC is set from the intervals estimate, so the rule reproduces the intervals estimate whatever the clusters are. The number that would have been checked, the estimate, was the one guaranteed to look right; the number that was wrong, the run length, is not usually reported at all. With ties resolved the estimate moved to 0.3621 and the shortest gap between clusters to nine.

Nothing threw and nothing looked odd. Two routes to one number count as two routes only when neither is built from the other, and the automatic rule is built from the intervals estimator by definition — so its agreement with that estimator is evidence about the arithmetic, never about the clusters. The quantity that caught the defect was the run length itself, set against a spacing the process was known to have.

Still open, and what comes next

One spacing. Every echo here arrives exactly six steps after its innovation. Real echoes do not — a tide is not a fixed number of readings — and a mixture of spacings turns the cliff into a ramp. Whether the fixed run length’s worst case stays eleven times the automatic rule’s on a ramp is unmeasured, and it is the obvious variation to run.

The fit under the automatic rule. The level and the standard error were measured at fixed run lengths only. Fitting the cluster maxima the automatic rule chooses, and carrying the variability of its estimated index into the level, is the next measurement. The two intervals for one return level were priced on independent blocks, where the index is one and known; an interval for a level with an estimated index inside it has not been measured anywhere in this field.

The index at a finite threshold. At the 0.95 quantile every rule that sees the triples reads about 0.30. Whether that is a failure of estimation or the index this process genuinely displays at that threshold is not separated here. For the max-moving-maximum the finite-threshold index can be computed exactly, and it was not; until it is, the low readings at the bottom threshold are attributed to overlapping clusters by argument rather than by count.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCluster sizeDeclusteringExceedanceExtremal indexFréchet lawGeneralised ParetoIndependenceMean squared errorPeaks over thresholdReturn levelShape parameterStandard errorThreshold