The tail past the last observation

Three shapes, one limit

A normalised sum has one limit and a normalised maximum has three, indexed by a single number. Twenty blocks put the sign of that number right 97.3% of the time — and naming the family from a light-tailed record gets worse as the record grows, from 83.0% at twenty blocks to 4.8% at five hundred.

Worth reading first: Sums of almost anything.

Four hundred records, a hundred readings in every block, twenty blocks apiece. The sign of the estimated shape comes out right on 97.3% of the heavy-tailed records and 99.8% of the bounded ones, which is enough to say whether the quantity has a ceiling or does not. Naming which of the three limit laws the record actually came from is a different question with a much worse answer, and on a light-tailed parent it is right 83.0% of the time at twenty blocks and 4.8% at five hundred.

That last pair is the whole of this essay. A twenty-fivefold longer record turns a call that was mostly right into one that is almost always wrong, and nothing about it is noise: the estimate barely moves across the sweep while the interval around it shrinks by a factor of eight onto a number that is not the truth. A confident wrong answer arriving with more data is rare enough to be worth building an essay on, and it has a neighbour in this collection where an interval for a proportion covers twelve points worse at twenty observations than at nineteen. The mechanism there was arithmetic on a lattice. The mechanism here is that the thing being fitted is not the thing the fit is a law for.

A sum has one limit and a maximum has three

The theorem everybody is taught first says that a normalised sum converges on one shape, and it is remarkable and it is true. The corresponding theorem for a maximum says something weaker and stranger: if the normalised maximum of nn independent readings converges at all, its limit is

G(x)=exp{(1+ξx)1/ξ}G(x) = \exp\left\{-\left(1 + \xi x\right)^{-1/\xi}\right\}

and every parent that converges converges to a member of that one family. There is one free number in it. Positive ξ\xi is the Fréchet law, heavy-tailed and unbounded above. Negative ξ\xi is the Weibull law, which stops: it has an upper endpoint at μσ/ξ\mu - \sigma/\xi, and a fit that reports a negative shape is asserting that the quantity has a ceiling. And ξ=0\xi = 0, the join between them, is the Gumbel law, which is where everything with an exponential-ish tail lands — including the normal.

So the three laws are one law and the analysis is a claim about one number. That is a much better situation than three unrelated families would be, and it is the reason the extreme-value literature can be short. It also means the entire weight of the reading rests on a single estimate, which is what this field is about.

Which domain of attraction a parent belongs to is a property of its tail alone, and for the three parents used here it is arithmetic rather than a fit. A Pareto with tail index α\alpha has ξ=1/α\xi = 1/\alpha, so α=2\alpha = 2 gives ξ=0.5\xi = 0.5. A Beta(1,β)\mathrm{Beta}(1, \beta) stops at one and has ξ=1/β\xi = -1/\beta, so β=2\beta = 2 gives ξ=0.5\xi = -0.5. Anything with an exponential tail has ξ=0\xi = 0 exactly. Those three numbers are what every estimate below is scored against, and none of them was estimated.

Blocks drawn rather than taken

A block maximum here is not the largest of a hundred draws. If XX has distribution function FF then the maximum of bb independent draws has distribution FbF^b exactly, so a single uniform gives a block maximum with no approximation in it at all:

M=F1 ⁣(U1/b)M = F^{-1}\!\left(U^{1/b}\right)

This is the same discipline as enumerating a reference distribution rather than sampling it, one level up, and it is what makes the sweeps affordable: a hundred blocks of a thousand readings costs a hundred uniforms rather than a hundred thousand draws. It also removes a whole class of excuse — nothing below can be blamed on a sampled maximum being an unlucky one.

Two routes to every number is the standing rule here rather than a courtesy — a site about probability that only simulates has no way to tell a right answer from a plausible one — so the two routes are checked against each other rather than assumed to agree. Six thousand block maxima drawn through the inverse are compared with six thousand maxima actually taken over fifty draws apiece, by a two-sample Kolmogorov statistic: 0.01517 against a 0.1% critical value of 0.03560. If the exact draw were subtly wrong — an off-by-one in the exponent, a uniform on the wrong half-open interval — that number would be the one to notice it.

The sign is the easy question

The first question a record is asked is the crude one. Does the estimated shape have the sign the parent’s limit actually has, which is the difference between a quantity that is unbounded above and one that is not.

Twenty blocks is enough to know which way the tail goes. The share of records whose estimated shape has the sign the parent's limit law actually has, over 400 records at each block count, at 100 readings a block. Twenty blocks already get it right 97.3% of the time for the heavy-tailed parent and 99.8% for the bounded one, and both are at 100.0% from 50 blocks up. The light-tailed parent has no sign to get right — its shape is exactly zero — and the line drawn for it is the share estimated above zero, which falls from 26.8% to 0.0% because its estimate is biased downward by an amount more blocks do not remove.
Fig. 1 The share of records whose estimated shape has the sign its parent’s limit law actually has, at five record lengths. Twenty blocks already get it right on 97.3% of heavy-tailed records and 99.8% of bounded ones, and both reach certainty by fifty.

Twenty blocks is enough, and by fifty both signed parents are at 100%. The light-tailed parent’s own tail is the one this collection has already drawn at the scale it is actually read at, and it is the awkward case here for the same reason it is awkward there: it decays fast enough to have no sign and slowly enough that nothing finite reaches its limit. The expectation going in was that this would take hundreds of blocks, and it does not: two thousand readings arranged as twenty maxima of a hundred already separate a heavy tail from a bounded one nearly always. Whatever is hard about extreme-value analysis, it is not this.

The light-tailed parent has no sign to get right, because its shape is exactly zero. The line drawn for it is the share of records whose estimate came out above zero, and it falls from 26.8% at twenty blocks to 0.0% at five hundred — not one record in four hundred. That is not a failure of the crude question. It is the first sight of the answer to the hard one.

The family call, and the direction it moves

The stricter question is the one a practitioner is really asking, because a shape estimate arrives with an interval around it. Form a 95% interval for ξ\xi and record the fit as calling Fréchet, Weibull or Gumbel according to whether that interval sits above zero, sits below zero, or straddles it. Three-way, so a light-tailed parent can be got right rather than merely being unsigned.

More blocks, and the light tail is called wrong more often. The share of records whose three-way family call is right, over 400 records at each block count, at 100 readings a block. A 95% interval for the shape is formed and the record is recorded as calling Fréchet, Weibull or Gumbel according to whether that interval sits above zero, below it, or straddles it. The two signed parents go from 51.5% and 71.0% at 20 blocks to certainty by 100. The light-tailed one goes the other way — 83.0%, 75.5%, 55.0%, 36.5%, 4.8% — because its estimate sits at about -0.1157 whatever the record length, and a longer record only shrinks the interval onto that number.
Fig. 2 The share of records whose three-way family call is right, over four hundred records at each length. The two signed parents reach certainty by a hundred blocks; the light-tailed one falls at every step, from 83.0% to 4.8%.

The two signed parents behave as the sign column promised. They start at 51.5% and 71.0% at twenty blocks — which is much worse than their sign rate, because at twenty blocks the interval is wide enough to straddle zero even when the point estimate is clearly signed — and both are at 100% from a hundred blocks up. For those two, more data is exactly what the textbook says it is.

The light-tailed parent runs 83.0%, 75.5%, 55.0%, 36.5%, 4.8%. It falls at every single step. And the reason it starts high is unflattering: at twenty blocks a normal record is called Gumbel because the interval is too wide to exclude zero, not because anything about the estimate says zero. The call is right for the wrong reason, and lengthening the record removes the reason without fixing the estimate.

What it is called instead is worth naming, because “wrong” is not a direction. At twenty blocks the normal records split 83.0% Gumbel, 16.5% Weibull and 0.3% Fréchet. At five hundred they split 4.8% Gumbel and 95.3% Weibull, with not one record in four hundred calling Fréchet. So the failure is entirely to one side: a light-tailed record read at length is read as a quantity with a finite ceiling, and the more of it there is the more firmly the ceiling is asserted. Nothing in the fit ever proposes the opposite error.

The estimate does not move; the interval closes on it

If this were sampling noise it would show as an estimate wandering. It does not wander.

The shape estimate narrows, and one of them narrows onto the wrong number. The middle ninety per cent of the estimated shape parameter, over 400 records apiece, for three parents whose limits are a Fréchet, a Gumbel and a Weibull. Every band narrows as the record lengthens — the heavy-tailed parent's from 0.9734 wide at 20 blocks to 0.1375 at 500 — and two of the three narrow onto the shape their parent actually has. The light-tailed parent's narrows onto -0.1036 rather than onto zero, because a maximum of 100 normal readings is not yet at its limit and the fit reads the shape of what it was given.
Fig. 3 The middle ninety per cent of the estimated shape at each record length, for three parents. Two bands narrow onto the shape their parent has; the light-tailed parent’s narrows onto −0.1036 rather than onto zero.

The normal parent’s mean estimate reads −0.1482, −0.1133, −0.1157, −0.1079, −0.1036 across the five record lengths — a quantity that has essentially arrived by fifty blocks and then sits there. Its spread falls from 0.2295 to 0.0289, a factor of eight. So the interval is doing exactly what an interval is supposed to do, and what it is closing on is wrong by about a tenth.

The other two bands are the control that says the estimator is not simply broken. The Pareto parent comes out at 0.5004 with a spread of 0.0426 at five hundred blocks, against a truth of 0.5; the bounded parent at −0.4993 with a spread of 0.0236, against a truth of −0.5. Both are right to three decimal places. The middle ninety per cent of the Pareto’s estimates runs from 0.0848 to 1.0582 at twenty blocks and from 0.4247 to 0.5622 at five hundred, which is a band tightening onto the right answer in the ordinary way.

One more control is worth having, because a band that narrows is a band whose noise is being reported honestly, and a single panel of genuinely normal data wanders enough to look suspicious. All three bands narrow at the same rate in the same picture; only one of them narrows onto the wrong place.

A shape of −0.1036 is not a small thing to be wrong about. A negative shape is an assertion that the quantity has an upper endpoint, and at that value the endpoint sits 9.65 scale units above the location. A hundred-year reading extrapolated under a law with a ceiling and one without are different numbers, and the difference grows with the period asked for.

Would a larger block fix it

The obvious repair is more readings inside each block rather than more blocks. Ten times as many, and the sweep run again.

More blocks, and the light tail is called wrong more often. The share of records whose three-way family call is right, over 400 records at each block count, at 1000 readings a block. A 95% interval for the shape is formed and the record is recorded as calling Fréchet, Weibull or Gumbel according to whether that interval sits above zero, below it, or straddles it. The two signed parents go from 51.5% and 70.3% at 20 blocks to certainty by 100. The light-tailed one goes the other way — 88.0%, 84.3%, 76.0%, 61.8%, 32.0% — because its estimate sits at about -0.0833 whatever the record length, and a longer record only shrinks the interval onto that number.
Fig. 4 The same three-way family call with a thousand readings in each block rather than a hundred. The light-tailed parent still falls at every step, from 88.0% to 32.0%.

It helps and it does not fix. The light-tailed parent now runs 88.0%, 84.3%, 76.0%, 61.8%, 32.0% — better everywhere, still monotone downward, still below half by five hundred blocks. Its shape estimate at five hundred blocks moves from −0.1036 to −0.0728, which is a third of the bias removed for ten times the data.

That is the shape of the answer. The defect is not in the number of blocks and not in the estimator; a tenfold larger block buys about a third of it, which is what a quantity falling like one over the logarithm of the block looks like when it is asked to fall by a factor of ten.

The bias is a fact about the block

Reading the same estimate against block size at a fixed number of blocks separates the two directions, which the first sweep confounds by moving neither.

The bias is a fact about the block, not about the record. The bias in the estimated shape against the size of the block it is estimated from, at a fixed 200 blocks and 300 records apiece. The light-tailed parent's bias falls from -0.1666 at 10 readings a block to -0.0458 at 100000 — a factor of 3.64 for four orders of magnitude, which is the one-over-log-n rate showing up in an estimate rather than in a distance. The heavy-tailed parent's blocks are exactly Fréchet at every size, and its bias never exceeds 0.0258. Adding blocks does not touch either: the same light-tailed bias is -0.1133 at fifty blocks and -0.1036 at five hundred.
Fig. 5 The bias in the estimated shape against the number of readings in each block, at a fixed two hundred blocks. The light-tailed parent’s falls from −0.1666 to −0.0458 over four orders of magnitude; the heavy-tailed parent’s never exceeds 0.026.

At two hundred blocks the light-tailed bias runs −0.1666, −0.1071, −0.0753, −0.0605, −0.0458 as the block grows from ten readings to a hundred thousand. Four orders of magnitude buy a factor of 3.64. The heavy-tailed parent’s bias over the same range never exceeds 0.026 in absolute value and has no trend in it, which is exactly what it should do: a Pareto’s block maxima are Fréchet at every block size, with nothing to converge to.

So the fit is not misreading a Gumbel sample. It is reading correctly a sample that is not Gumbel.

The exact law of a maximum, and the limit it is fitted as. The exact distribution of the maximum of 100 normal readings, normalised by the textbook constants a = 0.3752 and b = 2.3263, drawn against the Gumbel law it converges to. The exact curve is Φ(x)^n and needs no simulation. The largest gap between them is 0.0272, at x = 2.027, where the exact law puts 0.9038 of its probability below and the limit puts 0.8766. That gap closes like one over the logarithm of the block: at a million readings a block it is still 0.0091, which the exponential parent passes before its blocks reach a hundred.
Fig. 6 The exact law of the maximum of a hundred normal readings, normalised, against the Gumbel law a fit treats it as. The largest gap between them is 0.0272, at x = 2.027, where the exact law puts 0.9038 of its probability below and the limit puts 0.8766.

The maximum of a hundred normal readings has an exact distribution — it is Φ(x)100\Phi(x)^{100}, and no simulation is needed to draw it — and that exact distribution is not the Gumbel law. The largest gap between the two is 0.0272, in the upper shoulder where the fit’s likelihood is most informative about the shape. A maximum-likelihood fit handed that sample returns the shape of that curve, which is slightly short-tailed relative to Gumbel, and it returns it more and more precisely as the record lengthens. There is nothing wrong with the estimator. The estimand moved.

What would have produced this without the claim being true

Four things, and three of them are ruled out.

The fit could be wrong. The shape is estimated twice by routes that share no arithmetic: a Nelder–Mead maximisation of the generalised extreme value likelihood, and Hosking’s closed-form probability-weighted-moment estimator, which reads the sample’s L-moments and has no search in it. On the light-tailed parent the two agree in the mean to 0.0017 and correlate 0.8997 record by record. Two independently-written estimators do not agree to the third decimal place on the same wrong answer.

The blocks could be wrong. They are drawn through the inverse rather than taken, so they are checked against taken maxima by the Kolmogorov statistic above and pass by a factor of two.

The interval could be wrong. It is a Wald interval built from a differenced observed information matrix, and the differencing is a genuine approximation. But an interval that was systematically too narrow would inflate the family call’s failure rate, and the failure being explained here is a point estimate that sits at −0.10 rather than at zero. A correct interval around a wrong centre would give the same picture more slowly.

The counting could be too coarse. Every share above is read over four hundred records, so a counted share carries a standard error of at worst 2.5 points and under 1.5 at the shares quoted. That is nowhere near the twenty-five point falls being described, and it is worth stating because a figure drawn from a sample can be right by luck in a way a figure drawn from a rule cannot. The monotone fall appears at every one of the five record lengths, in both block sizes swept, which is ten independent readings of the same direction.

And one is not ruled out. The Wald interval is the reading a package reports and it is the wrong shape for a likelihood that is skewed — which is the subject of the essay in this field that counts two intervals for one level, where a symmetric interval for a return level covers 80.3% rather than 95%. A profile interval for the shape would be asymmetric and would straddle zero for longer, so the light-tailed parent’s call would fall more slowly than it does here. It would still fall, because the point estimate is what is biased; but the specific numbers 83.0 and 4.8 are numbers about the Wald reading, and they are quoted as that.

What this does not say

The rule being scored is a rule, and a stupid rule can beat it on any one parent. A procedure that always answered Gumbel would score zero on both signed parents and a hundred per cent on the light-tailed one at every record length. So the three columns have to be read together, and the honest statement is that no block count in the sweep gets all three parents above 80% at once.

The bigger caveat is that “named the family correctly” is not the quantity anybody uses. A record is not fitted in order to publish a family name; it is fitted in order to read a level off it, and a shape that is wrong by a tenth matters exactly as much as the extrapolation it is put through. That is not something this sweep can say, and it is the reason the essay that reads a level with no data in it exists — where the estimate turns out to be nearly unbiased and the error grows sixfold anyway.

And the whole measurement is of records whose parent is known. Every bias reported here is a difference from a number that is arithmetic. A real record arrives with no such number, which is precisely why the rate at which the middle and the tail converge has to be priced rather than assumed, and it is why the light tail’s slow arrival is worth an essay of its own.

The rate this essay did not derive

One thing was opened here and left open. The bias is measured at five block sizes and shown to fall; multiplied by the logarithm of the block it reads −0.3835, −0.4931, −0.5201, −0.5574, −0.5273, which is near constant from a hundred readings up and visibly not flat at ten. So the rate is consistent with one over the logarithm of the block and is not shown to be it — unlike the same product for the distance between the two laws, where five orders of magnitude leave it constant to a quarter of a per cent.

Settling it means deriving the second-order term for a normal maximum rather than measuring it, and the reason to want that is not tidiness. A bias whose rate is known can be corrected, and correcting it would move the light-tailed family call further than any amount of extra data does. What is here instead is the weaker claim the measurement supports: the bias is about the block, it falls slowly, and a longer record does not touch it.

Two shapes out of three are read correctly from twenty blocks, and the third is read incorrectly with rising confidence for ever. Which is the one every record of rainfall, wind speed and river level in the world is assumed to have.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Block maximaBlock sizeCentral limit theoremDomain of attractionExtreme value theoryFréchet lawGeneralised extreme valueGumbel lawMaximum likelihoodProbability-weighted momentsSampling distributionShape parameterTail indexUpper endpointWeibull law