The run length a declustering chooses
Worth reading first: Three shapes, one limit.
A declustering needs a rule for where one cluster ends and the next begins, and the rule most analyses use is a single integer. Two exceedances that many steps apart or more belong to different clusters; closer than that, they are one episode. The essay that measured what clustering costs used a run length of four throughout and said so, and it ended on the measurement it had not made: the runs estimator is a function of that integer by construction, and nothing had swept it.
Swept, the integer matters in two entirely different ways, depending on what a cluster looks like. On a process whose clusters are runs of consecutive exceedances the estimate drifts gently across run lengths from two to forty, and four was a reasonable choice. On a process whose cluster members arrive six steps apart, the estimate at a run length of six is 0.9069 and at seven it is 0.3649, against an extremal index of exactly 0.40. There is no slope there to be roughly right on. There is a cliff, and four is on the wrong side of it.
What four costs on that process is nearly everything declustering is for. A return level quoted with no declustering at all is overstated by ×2.500; quoted with a run length of four it is overstated by ×2.356. The standard error reported for the fitted shape is 1.812 times too small at four and 1.095 at seven. None of that is visible in the record, and the run length that avoids it can be read off the data by a rule that never needed a run length in the first place.
A cluster whose members are not neighbours
The process here is a max-moving-maximum,
with independent unit Fréchet innovations. Its marginal is unit Fréchet, like the max-autoregression’s, and its extremal index is the largest weight divided by the sum of the weights, so exactly. One large innovation appears three times — at four tenths of its size, then twice more at three tenths, six and twelve steps later — and a threshold high enough to be about extremes sees it as up to three exceedances with five ordinary readings between each. The shape is a familiar one in anything that echoes at a fixed delay: a surge that returns on the next tide, a load that comes back with the next shift.
It is chosen for the same reason as the max-autoregression was. Its index is arithmetic rather than simulated, so every estimate below is scored against a number with no error of its own, which is the only way a comparison of two estimators says something about the estimators rather than about the simulation that produced the reference.
In one realised stretch of two hundred steps, nineteen readings clear the 0.9 quantile and twelve of the gaps between consecutive exceedances are exactly six steps long. A run length of four sees nineteen clusters of one. The estimate is 1.000, which is what the estimator reports for a series with no dependence at all — so on this record a declustering at four has looked straight at episodes of three and reported independence.
That is not a small-sample accident and it is not the threshold. Every gap of six between members of one cluster is longer than four, so every member starts a new cluster however long the record is. The runs estimator can only merge exceedances that fall inside its run length, and a dependence whose echoes arrive later than that lies outside what the rule is able to see. More data make the estimate more precise and leave it exactly as wrong.
With a run length of seven the same nineteen exceedances become seven clusters and the estimate is 0.368. The shading now spans the triples the eye picks out. Nothing about the record changed between the two pictures. Only the integer did, and the integer moved the estimate from the value for independence to within a few hundredths of the truth.
This stretch was picked because its clusters do not overlap, which is what makes the spacing legible, and that is a choice about what a picture can show rather than about what the estimator does. The counted behaviour below does not depend on it. At a high threshold, clusters from different innovations rarely overlap in any record; at a low one they often do, and that turns out to matter as much as the run length.
A slope on one process and a cliff on another
Four processes whose index is known exactly are run here: independent readings; two max-autoregressions, at and , whose clusters are runs of consecutive exceedances; and the spaced process at 0.4. Each gets three hundred records of four thousand steps with the threshold at the 0.98 quantile, and the runs estimator is read on every record at fifteen run lengths from one to forty.
On the max-autoregression with the longest clusters the estimate is 0.2573 at a run length of two and 0.2121 at forty, against 0.25. That is the whole range of the dial, and across it the estimate moves by less than a fifth of its own value. The run length of four the clustering essay used sits within a hundredth of the truth on this curve, which is why the choice looked harmless there: on a process of that kind, it was.
On the spaced process the estimate is 0.9424 at four, 2.36 times the index, and it barely moves until the run length passes the spacing — 0.9069 at six, then 0.3649 at seven. Once past the cliff it drifts down like the others, because a longer run length also starts merging clusters that came from different innovations.
The independent process shows the dial’s other cost. It reads 0.9407 at four and 0.4573 at forty: with no dependence at all, a long run length merges exceedances that merely happened to land near one another. So a run length long enough to be safe on the spaced process is expensive wherever clusters are short, and a run length short enough for independent data is blind to the spaced process. No single integer is right for all four. The value at a run length of one, which declares every exceedance its own cluster, is on every dependent process by construction — the overstatement the clustering essay priced, now visible as the left edge of the same curve.
How the error presents matters as much as its size. At a run length of four the spaced process’s estimate has a root mean squared error of 0.5437, and 99.5% of its squared error is bias. The estimate is not noisy. It is precise, it is wrong, and its spread across records says nothing about either. A record read at four is indistinguishable from a record of a process with almost no clustering, and no amount of care with that one number reveals the difference.
What the wrong side of the cliff costs a fit
The marginal is unit Fréchet, so the shape is exactly one and a return level is proportional to the extremal index an analysis puts into it — the closed form the clustering essay derived, with nothing asymptotic left in the ratio. Quote a level with an estimated index in place of 0.40 and the level is off by exactly their ratio. With no declustering the factor is ×2.500. With a run length of four it is ×2.356, which means declustering at four removed under a tenth of the overstatement it was run to remove.
At a run length of seven the factor is ×0.905 and at thirteen it is ×0.864. Past the cliff, the error changes sign. A run length long enough to merge the echoes of one innovation also merges some neighbouring innovations, which counts too few clusters and therefore quotes too low a level. For a level used to set the height of a defence that is the direction nobody budgets for. It is a tenth rather than a factor of two, but it is not zero, and it grows with the run length.
The standard error fails in step with the level. Fit a generalised Pareto to the cluster maxima and set the counted spread of its shape across records beside . At a run length of four each record yields 75.1 cluster maxima and the counted spread is 1.812 times the reported one — no better than fitting every exceedance, where the ratio is 1.807. At seven the maxima fall to 28.9 and the ratio to 1.095, which is where the formula is meant to sit.
Meanwhile the fitted shape barely moves: 0.9342 at four and 0.9708 at seven, against a truth of one. So an analyst who fitted at both run lengths would see nearly the same shape and an error bar far wider at seven, and the natural reading is that four is the more precise analysis. It is not more precise. It reports a precision the data do not contain, the same arithmetic as an interval that divides by a count of observations that repeat each other, with the count here being cluster maxima that are really echoes of one another.
Two routes that need no run length
The intervals estimator of Ferro and Segers reads the extremal index from the inter-exceedance times directly, with no declustering constant in it at all. On the spaced process it reads 0.4491 — high, as it was high at every dependence in the clustering essay’s sweep, but on the right side of a cliff it was never told about.
Their automatic declustering goes one step further. Turn the intervals estimate into a number of clusters, from exceedances, and declare the longest gaps between exceedances to be gaps between clusters. The run length is then an output rather than an input. Where gaps tie — and inter-exceedance times are integers, so they tie constantly — is lowered until the -th longest gap is strictly longer than the -th, so the rule never has to split a set of equal gaps arbitrarily.
On the spaced process the median shortest gap the rule treats as a gap between clusters is 9, and nine records in ten choose 7 or more. The rule lands past the spacing on its own and estimates 0.3621. On the process with the longest runs the choice spreads from 3 to 48, which looks alarming and is not: across that entire range, the runs curve for that process moves by less than a fifth of its value. The rule is imprecise exactly where precision in the run length buys nothing and precise where it buys everything, which is what a rule that reads the clusters should do and what a convention cannot.
Why this works is worth stating exactly. The intervals estimator has no constant, so it measures how large the clusters are without having to know what shape they take. The automatic rule turns that size into a count and lets the gaps sort themselves: on the spaced process the gaps between clusters are long and the gaps inside a cluster are all six, so the longest gaps are the right ones whenever is roughly right. It is the same move as a smoothing bandwidth read off the readings themselves instead of fixed in advance, and it has the same limitation: the rule is only as good as the quantity it reads, and here that quantity is an estimator already known to run high.
Judged by the process that suits each rule least
A fixed run length of four has the smallest error of the three usable rules on three of the four processes: 0.0635 on independent readings, 0.0594 and 0.0475 on the two max-autoregressions. On the spaced process its error is 0.5437, 11.1 times the automatic rule’s 0.0490. Judged instead by the worst process each rule meets, the fixed run length’s error is 0.5437, the intervals estimator’s 0.1004 and the automatic rule’s 0.0783.
The automatic rule beats the intervals estimator it is built from on all three dependent processes, and loses to it only on independent readings, 0.0773 against 0.0675. Where it is not the best usable rule it is never more than 1.32 times the best one’s error, and on the spaced process it comes within four thousandths of the run length that happened to do best — which was seven, a number only a process with a known answer can reveal.
So a fixed run length is a bet on the shape of the clusters. An analyst who knew their clusters were runs of neighbours would do well to make it: four is better than either alternative on three processes here. What decides between rules is what happens when the bet is wrong, and the answers are a factor of eleven against a factor of at most 1.32. That is the same kind of choice as how long a block a multiplier bootstrap shares, where a length set in advance walks the rejection rate straight through its nominal level. A constant that encodes an assumption about dependence is safe exactly as far as the assumption is.
It differs in one respect from a restricted-mean horizon chosen after looking. The automatic rule does look at the data, but at the gaps, not at the estimate an analyst hopes to report, and its choice is remade on every record here, so its variability is already inside the errors quoted above. Picking whichever run length made the level look most reasonable would not be.
The threshold moves the cliff
At the 0.95, 0.98 and 0.99 quantiles a run length of four reads 0.8564, 0.9424 and 0.9726. It gets worse as the threshold rises, because fewer exceedances means fewer accidental merges of neighbouring clusters to pull it down. A run length of seven reads 0.3029, 0.3649 and 0.3907, and the automatic rule reads 0.3029, 0.3621 and 0.3784 — low at every threshold and far low at the bottom one. The intervals estimator reads 0.4780, 0.4491 and 0.4513, high at all three.
At the 0.95 quantile one reading in twenty is an exceedance, 199.7 a record, and each innovation’s echoes span twelve steps. Clusters from different innovations overlap often enough that any declustering which sees the triples also joins some of them, and a run length of seven and the automatic rule arrive at exactly the same estimate. The extremal index is a limit as the threshold rises. At a threshold that low, the clustering a record displays is not the clustering the limit counts, and no choice of run length recovers it.
At the 0.99 quantile the problem inverts. There are 39.7 exceedances a record, the estimates sit closer to the index, and the automatic rule’s choice spreads from 8 to 60 with a median of 22. The data now contain too few gaps to say where clusters end. That is the dial every threshold analysis already has — bias bought down with exceedances — turning up a second time inside the declustering, and the run length does not escape it.
Agreement that was built in
The first version of the automatic rule here took the longest gaps literally and did not resolve ties. On the spaced process it estimated 0.4498, against the intervals estimator’s 0.4491, and its median shortest gap between clusters was 6. On at least half the records it had declared a gap of six or less to be a gap between clusters, splitting the very episodes it was meant to find.
The two estimates agreed to within a thousandth, and that agreement looked like two methods confirming each other. It was worth nothing. Without the tie rule the automatic estimate is , and is set from the intervals estimate, so the rule reproduces the intervals estimate whatever the clusters are. The number that would have been checked, the estimate, was the one guaranteed to look right; the number that was wrong, the run length, is not usually reported at all. With ties resolved the estimate moved to 0.3621 and the shortest gap between clusters to nine.
Nothing threw and nothing looked odd. Two routes to one number count as two routes only when neither is built from the other, and the automatic rule is built from the intervals estimator by definition — so its agreement with that estimator is evidence about the arithmetic, never about the clusters. The quantity that caught the defect was the run length itself, set against a spacing the process was known to have.
Still open, and what comes next
One spacing. Every echo here arrives exactly six steps after its innovation. Real echoes do not — a tide is not a fixed number of readings — and a mixture of spacings turns the cliff into a ramp. Whether the fixed run length’s worst case stays eleven times the automatic rule’s on a ramp is unmeasured, and it is the obvious variation to run.
The fit under the automatic rule. The level and the standard error were measured at fixed run lengths only. Fitting the cluster maxima the automatic rule chooses, and carrying the variability of its estimated index into the level, is the next measurement. The two intervals for one return level were priced on independent blocks, where the index is one and known; an interval for a level with an estimated index inside it has not been measured anywhere in this field.
The index at a finite threshold. At the 0.95 quantile every rule that sees the triples reads about 0.30. Whether that is a failure of estimation or the index this process genuinely displays at that threshold is not separated here. For the max-moving-maximum the finite-threshold index can be computed exactly, and it was not; until it is, the low readings at the bottom threshold are attributed to overlapping clusters by argument rather than by count.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A level with no data in it — both name closed form, mean squared error, return level, shape parameter
- The maximum converges slowly — both name closed form, return level, shape parameter
- The sample is a condition — both name closed form, independence, threshold
- A coverage table with its own error — both name closed form, standard error
- A cut is not a polynomial, and it does not have to be — both name closed form, threshold
- A penalty is a trace — both name closed form, mean squared error
Named objects
A flat tag is an object no other essay names yet.
Closed formCluster sizeDeclusteringExceedanceExtremal indexFréchet lawGeneralised ParetoIndependenceMean squared errorPeaks over thresholdReturn levelShape parameterStandard errorThreshold