Concept

Generalised Pareto — where it appears

The two-parameter law an exceedance over a high threshold converges to, with the same shape parameter as the corresponding law for maxima. It is what a threshold analysis fits, and the shared shape is why a threshold fit and a block-maxima fit are two estimators of one number.

Named by 4 essays across one field — each of them below, with the objects they name alongside it.

Also named here as peaks over threshold — the same set of essays touches all of them, so they are one junction rather than several.

The threshold buys accuracy and spends exceedances. The mean squared error of the estimated shape against the threshold, split into the square of its bias and its spread, over 600 records of 2000 readings from a a normal parent. At the 0.9 quantile 199 exceedances are left, the bias is -0.1708, the spread is 0.0701 and the total error is 0.0341. The bias falls as the threshold rises because the exceedances get closer to being generalised Pareto; the spread rises because there are fewer of them. The sum is smallest at the 0.925 quantile, at 0.0340, of which 80.6% is still bias — so even the best threshold on this grid is one where accuracy, not spread, is the binding constraint.

The threshold is a dial

A peaks-over-threshold analysis has one knob, and raising it buys accuracy with exceedances. For a normal parent the error is smallest at the 0.925 quantile and 80.6% of it is still bias there — and both diagnostics practitioners use to set the knob lose to a fixed 0.90 rule, one by a factor of 1.590 and one by 11.881.

extreme · Extremes
The exceedances arrive together. 300 steps of a max-autoregression with dependence 0.75, drawn on a logarithmic scale because its marginal has no variance. The rule marks the 0.9 quantile: 30 of the 300 readings are above it and they fall into 5 clusters, the largest holding 11. The mean cluster holds 6.000, and its reciprocal — 0.167 — is the runs estimator of the extremal index, whose true value for this process is exactly 1 − 0.75 = 0.25. Every threshold method in the collection assumes exceedances are independent pieces of information; here 30 of them are 5.

The clustering the tail has

Every threshold method counts exceedances as though they were independent pieces of information, and in a dependent series they arrive in clusters. Ignoring that overstates a return level by the reciprocal of the extremal index — ×3.527 counted where the mean cluster holds four — and leaves a reported standard error 2.151 times too small.

extreme · Extremes
A run length of 4 makes every exceedance its own cluster. 200 steps of a max-moving-maximum, X(t) = max(0.4·Z(t), 0.3·Z(t−6), 0.3·Z(t−12)) with unit Fréchet innovations Z, whose extremal index is exactly 0.40: one large innovation can put three readings above a threshold, 6 steps apart. The rule marks the 0.9 quantile and 19 readings clear it; 12 of the gaps between consecutive exceedances are exactly 6 steps. With a run length of 4, so that two exceedances 4 or more steps apart start separate clusters, they form 19 clusters, shaded, the largest holding 1. The runs estimator reads 1.000 against 0.40.

The run length a declustering chooses

The runs estimator of an extremal index carries a constant nobody derives. Where a cluster is a run of neighbouring exceedances the constant barely matters; where a cluster's members fall six steps apart, the estimate is 0.9069 at a run length of six and 0.3649 at seven against an index of 0.40, and a run length of four removes under a tenth of the overstatement declustering exists to remove. A rule that reads the run length off the data has the smallest worst error of the three.

extreme · Extremes
How much of a return level's uncertainty is the extremal index's, by how far past the record the level is. Delta-method variance of the level from a declustered peaks-over-threshold fit, averaged over 300 records of 4,000 observations. The cluster rate's share at 1.5, 2, 5, 20, 100, 1000 periods: independent, 0.374, 0.253, 0.084, 0.027, 0.012, 0.005; runs of exceedances, θ = 0.6, 0.558, 0.396, 0.141, 0.039, 0.015, 0.006; spaced clusters, θ = 0.4, 0.845, 0.562, 0.222, 0.057, 0.019, 0.007; runs of exceedances, θ = 0.25, 0.787, 0.788, 0.327, 0.083, 0.024, 0.008.

Where the extremal index matters to a return level

A return level from a declustered fit depends on how often clusters arrive, which is the extremal index estimated from the same record. For a level reached once in two periods of a twenty-period record the index's error is a quarter to four fifths of the level's variance, and an interval that leaves it out covers 59.0% when clusters are long; put in, 90.0%. For a level reached once in a hundred periods its share is under two and a half per cent and adding it changes no interval at all — what fails there is the shape, fitted from however many clusters the dependence left: 82.3% coverage from forty-six clusters, 65.7% from nineteen.

extreme · Extremes

Named alongside it

The objects these essays reach for when they reach for this one.

Peaks over thresholdClosed formDeclusteringExceedanceExtremal indexReturn levelShape parameterStandard errorCluster sizeEffective sample sizeFréchet lawIndependence

All concepts