Concept

Cluster size — where it appears

How many consecutive threshold exceedances belong to one episode, under a rule that separates episodes by a stated run of non-exceedances. Its mean is the reciprocal of the extremal index, and it is what turns a count of exceedances into a smaller count of independent extremes.

Named by 4 essays across 2 fields — each of them below, with the objects they name alongside it.

The exceedances arrive together. 300 steps of a max-autoregression with dependence 0.75, drawn on a logarithmic scale because its marginal has no variance. The rule marks the 0.9 quantile: 30 of the 300 readings are above it and they fall into 5 clusters, the largest holding 11. The mean cluster holds 6.000, and its reciprocal — 0.167 — is the runs estimator of the extremal index, whose true value for this process is exactly 1 − 0.75 = 0.25. Every threshold method in the collection assumes exceedances are independent pieces of information; here 30 of them are 5.

The clustering the tail has

Every threshold method counts exceedances as though they were independent pieces of information, and in a dependent series they arrive in clusters. Ignoring that overstates a return level by the reciprocal of the extremal index — ×3.527 counted where the mean cluster holds four — and leaves a reported standard error 2.151 times too small.

extreme · Extremes
The rows are held fixed; only the clusters move. Counted coverage of four 95% intervals for a slope, at five cluster counts with the row count held at 300 throughout and a within-cluster correlation of 0.1, over 6000 draws apiece, with the sizes equal. The interval that counts rows covers 53.42% at 5 clusters — a second closed form says 2Φ(z/√D) − 1 = 54.44% for a design effect of 6.900, and reads nothing about clusters at all. The cluster-robust interval read against a normal covers 74.43% there and 94.20% at 100 clusters; read against a t on G − 1 it covers 85.08% and 94.47%. The number of independent things is the cluster count, and every quantity here is blind to how many rows were typed.

The count that is not the rows

Three hundred rows in five clusters of sixty carry 6.9000 times the variance an independent-rows calculation reports, and the interval that counts rows covers 53.42%. The same five unequal sizes laid out two ways give design effects of 9.3158 and 5.4652.

sandwich · Misspecification
A run length of 4 makes every exceedance its own cluster. 200 steps of a max-moving-maximum, X(t) = max(0.4·Z(t), 0.3·Z(t−6), 0.3·Z(t−12)) with unit Fréchet innovations Z, whose extremal index is exactly 0.40: one large innovation can put three readings above a threshold, 6 steps apart. The rule marks the 0.9 quantile and 19 readings clear it; 12 of the gaps between consecutive exceedances are exactly 6 steps. With a run length of 4, so that two exceedances 4 or more steps apart start separate clusters, they form 19 clusters, shaded, the largest holding 1. The runs estimator reads 1.000 against 0.40.

The run length a declustering chooses

The runs estimator of an extremal index carries a constant nobody derives. Where a cluster is a run of neighbouring exceedances the constant barely matters; where a cluster's members fall six steps apart, the estimate is 0.9069 at a run length of six and 0.3649 at seven against an index of 0.40, and a run length of four removes under a tenth of the overstatement declustering exists to remove. A rule that reads the run length off the data has the smallest worst error of the three.

extreme · Extremes
One estimator, three answers, and only the reference changes. Coverage of the cluster-robust 95% interval for the slope against the number of clusters, at 30 rows in each. The estimator is identical in all three curves; what differs is the number it is compared against. At 5 clusters it covers 75.05% against a normal, 85.30% against a t on 4 degrees of freedom and 87.95% against a t on 3. At 80 clusters the three agree to within a point. The correction costs nothing: the same standard error, a different table.

The reference the sandwich is read against

The cluster-robust interval covers 75.05% at five clusters and 93.58% at eighty. The same estimate read against a t on G − 2 covers 87.95% at five, and the estimator is unchanged — three hundred rows grouped into five clusters cover 74.28% where the same three hundred grouped into seventy-five cover 94.63%.

sandwich · Misspecification

Named alongside it

The objects these essays reach for when they reach for this one.

IndependenceClosed formCluster-robust standard errorClustered samplingCoverageDeclusteringDegrees of freedomDesign effectEffective sample sizeExceedanceExtremal indexFréchet law

All concepts