Three shapes, one limit
Worth reading first: Sums of almost anything.
Four hundred records, a hundred readings in every block, twenty blocks apiece. The sign of the estimated shape comes out right on 97.3% of the heavy-tailed records and 99.8% of the bounded ones, which is enough to say whether the quantity has a ceiling or does not. Naming which of the three limit laws the record actually came from is a different question with a much worse answer, and on a light-tailed parent it is right 83.0% of the time at twenty blocks and 4.8% at five hundred.
That last pair is the whole of this essay. A twenty-fivefold longer record turns a call that was mostly right into one that is almost always wrong, and nothing about it is noise: the estimate barely moves across the sweep while the interval around it shrinks by a factor of eight onto a number that is not the truth. A confident wrong answer arriving with more data is rare enough to be worth building an essay on, and it has a neighbour in this collection where an interval for a proportion covers twelve points worse at twenty observations than at nineteen. The mechanism there was arithmetic on a lattice. The mechanism here is that the thing being fitted is not the thing the fit is a law for.
A sum has one limit and a maximum has three
The theorem everybody is taught first says that a normalised sum converges on one shape, and it is remarkable and it is true. The corresponding theorem for a maximum says something weaker and stranger: if the normalised maximum of independent readings converges at all, its limit is
and every parent that converges converges to a member of that one family. There is one free number in it. Positive is the Fréchet law, heavy-tailed and unbounded above. Negative is the Weibull law, which stops: it has an upper endpoint at , and a fit that reports a negative shape is asserting that the quantity has a ceiling. And , the join between them, is the Gumbel law, which is where everything with an exponential-ish tail lands — including the normal.
So the three laws are one law and the analysis is a claim about one number. That is a much better situation than three unrelated families would be, and it is the reason the extreme-value literature can be short. It also means the entire weight of the reading rests on a single estimate, which is what this field is about.
Which domain of attraction a parent belongs to is a property of its tail alone, and for the three parents used here it is arithmetic rather than a fit. A Pareto with tail index has , so gives . A stops at one and has , so gives . Anything with an exponential tail has exactly. Those three numbers are what every estimate below is scored against, and none of them was estimated.
Blocks drawn rather than taken
A block maximum here is not the largest of a hundred draws. If has distribution function then the maximum of independent draws has distribution exactly, so a single uniform gives a block maximum with no approximation in it at all:
This is the same discipline as enumerating a reference distribution rather than sampling it, one level up, and it is what makes the sweeps affordable: a hundred blocks of a thousand readings costs a hundred uniforms rather than a hundred thousand draws. It also removes a whole class of excuse — nothing below can be blamed on a sampled maximum being an unlucky one.
Two routes to every number is the standing rule here rather than a courtesy — a site about probability that only simulates has no way to tell a right answer from a plausible one — so the two routes are checked against each other rather than assumed to agree. Six thousand block maxima drawn through the inverse are compared with six thousand maxima actually taken over fifty draws apiece, by a two-sample Kolmogorov statistic: 0.01517 against a 0.1% critical value of 0.03560. If the exact draw were subtly wrong — an off-by-one in the exponent, a uniform on the wrong half-open interval — that number would be the one to notice it.
The sign is the easy question
The first question a record is asked is the crude one. Does the estimated shape have the sign the parent’s limit actually has, which is the difference between a quantity that is unbounded above and one that is not.
Twenty blocks is enough, and by fifty both signed parents are at 100%. The light-tailed parent’s own tail is the one this collection has already drawn at the scale it is actually read at, and it is the awkward case here for the same reason it is awkward there: it decays fast enough to have no sign and slowly enough that nothing finite reaches its limit. The expectation going in was that this would take hundreds of blocks, and it does not: two thousand readings arranged as twenty maxima of a hundred already separate a heavy tail from a bounded one nearly always. Whatever is hard about extreme-value analysis, it is not this.
The light-tailed parent has no sign to get right, because its shape is exactly zero. The line drawn for it is the share of records whose estimate came out above zero, and it falls from 26.8% at twenty blocks to 0.0% at five hundred — not one record in four hundred. That is not a failure of the crude question. It is the first sight of the answer to the hard one.
The family call, and the direction it moves
The stricter question is the one a practitioner is really asking, because a shape estimate arrives with an interval around it. Form a 95% interval for and record the fit as calling Fréchet, Weibull or Gumbel according to whether that interval sits above zero, sits below zero, or straddles it. Three-way, so a light-tailed parent can be got right rather than merely being unsigned.
The two signed parents behave as the sign column promised. They start at 51.5% and 71.0% at twenty blocks — which is much worse than their sign rate, because at twenty blocks the interval is wide enough to straddle zero even when the point estimate is clearly signed — and both are at 100% from a hundred blocks up. For those two, more data is exactly what the textbook says it is.
The light-tailed parent runs 83.0%, 75.5%, 55.0%, 36.5%, 4.8%. It falls at every single step. And the reason it starts high is unflattering: at twenty blocks a normal record is called Gumbel because the interval is too wide to exclude zero, not because anything about the estimate says zero. The call is right for the wrong reason, and lengthening the record removes the reason without fixing the estimate.
What it is called instead is worth naming, because “wrong” is not a direction. At twenty blocks the normal records split 83.0% Gumbel, 16.5% Weibull and 0.3% Fréchet. At five hundred they split 4.8% Gumbel and 95.3% Weibull, with not one record in four hundred calling Fréchet. So the failure is entirely to one side: a light-tailed record read at length is read as a quantity with a finite ceiling, and the more of it there is the more firmly the ceiling is asserted. Nothing in the fit ever proposes the opposite error.
The estimate does not move; the interval closes on it
If this were sampling noise it would show as an estimate wandering. It does not wander.
The normal parent’s mean estimate reads −0.1482, −0.1133, −0.1157, −0.1079, −0.1036 across the five record lengths — a quantity that has essentially arrived by fifty blocks and then sits there. Its spread falls from 0.2295 to 0.0289, a factor of eight. So the interval is doing exactly what an interval is supposed to do, and what it is closing on is wrong by about a tenth.
The other two bands are the control that says the estimator is not simply broken. The Pareto parent comes out at 0.5004 with a spread of 0.0426 at five hundred blocks, against a truth of 0.5; the bounded parent at −0.4993 with a spread of 0.0236, against a truth of −0.5. Both are right to three decimal places. The middle ninety per cent of the Pareto’s estimates runs from 0.0848 to 1.0582 at twenty blocks and from 0.4247 to 0.5622 at five hundred, which is a band tightening onto the right answer in the ordinary way.
One more control is worth having, because a band that narrows is a band whose noise is being reported honestly, and a single panel of genuinely normal data wanders enough to look suspicious. All three bands narrow at the same rate in the same picture; only one of them narrows onto the wrong place.
A shape of −0.1036 is not a small thing to be wrong about. A negative shape is an assertion that the quantity has an upper endpoint, and at that value the endpoint sits 9.65 scale units above the location. A hundred-year reading extrapolated under a law with a ceiling and one without are different numbers, and the difference grows with the period asked for.
Would a larger block fix it
The obvious repair is more readings inside each block rather than more blocks. Ten times as many, and the sweep run again.
It helps and it does not fix. The light-tailed parent now runs 88.0%, 84.3%, 76.0%, 61.8%, 32.0% — better everywhere, still monotone downward, still below half by five hundred blocks. Its shape estimate at five hundred blocks moves from −0.1036 to −0.0728, which is a third of the bias removed for ten times the data.
That is the shape of the answer. The defect is not in the number of blocks and not in the estimator; a tenfold larger block buys about a third of it, which is what a quantity falling like one over the logarithm of the block looks like when it is asked to fall by a factor of ten.
The bias is a fact about the block
Reading the same estimate against block size at a fixed number of blocks separates the two directions, which the first sweep confounds by moving neither.
At two hundred blocks the light-tailed bias runs −0.1666, −0.1071, −0.0753, −0.0605, −0.0458 as the block grows from ten readings to a hundred thousand. Four orders of magnitude buy a factor of 3.64. The heavy-tailed parent’s bias over the same range never exceeds 0.026 in absolute value and has no trend in it, which is exactly what it should do: a Pareto’s block maxima are Fréchet at every block size, with nothing to converge to.
So the fit is not misreading a Gumbel sample. It is reading correctly a sample that is not Gumbel.
The maximum of a hundred normal readings has an exact distribution — it is , and no simulation is needed to draw it — and that exact distribution is not the Gumbel law. The largest gap between the two is 0.0272, in the upper shoulder where the fit’s likelihood is most informative about the shape. A maximum-likelihood fit handed that sample returns the shape of that curve, which is slightly short-tailed relative to Gumbel, and it returns it more and more precisely as the record lengthens. There is nothing wrong with the estimator. The estimand moved.
What would have produced this without the claim being true
Four things, and three of them are ruled out.
The fit could be wrong. The shape is estimated twice by routes that share no arithmetic: a Nelder–Mead maximisation of the generalised extreme value likelihood, and Hosking’s closed-form probability-weighted-moment estimator, which reads the sample’s L-moments and has no search in it. On the light-tailed parent the two agree in the mean to 0.0017 and correlate 0.8997 record by record. Two independently-written estimators do not agree to the third decimal place on the same wrong answer.
The blocks could be wrong. They are drawn through the inverse rather than taken, so they are checked against taken maxima by the Kolmogorov statistic above and pass by a factor of two.
The interval could be wrong. It is a Wald interval built from a differenced observed information matrix, and the differencing is a genuine approximation. But an interval that was systematically too narrow would inflate the family call’s failure rate, and the failure being explained here is a point estimate that sits at −0.10 rather than at zero. A correct interval around a wrong centre would give the same picture more slowly.
The counting could be too coarse. Every share above is read over four hundred records, so a counted share carries a standard error of at worst 2.5 points and under 1.5 at the shares quoted. That is nowhere near the twenty-five point falls being described, and it is worth stating because a figure drawn from a sample can be right by luck in a way a figure drawn from a rule cannot. The monotone fall appears at every one of the five record lengths, in both block sizes swept, which is ten independent readings of the same direction.
And one is not ruled out. The Wald interval is the reading a package reports and it is the wrong shape for a likelihood that is skewed — which is the subject of the essay in this field that counts two intervals for one level, where a symmetric interval for a return level covers 80.3% rather than 95%. A profile interval for the shape would be asymmetric and would straddle zero for longer, so the light-tailed parent’s call would fall more slowly than it does here. It would still fall, because the point estimate is what is biased; but the specific numbers 83.0 and 4.8 are numbers about the Wald reading, and they are quoted as that.
What this does not say
The rule being scored is a rule, and a stupid rule can beat it on any one parent. A procedure that always answered Gumbel would score zero on both signed parents and a hundred per cent on the light-tailed one at every record length. So the three columns have to be read together, and the honest statement is that no block count in the sweep gets all three parents above 80% at once.
The bigger caveat is that “named the family correctly” is not the quantity anybody uses. A record is not fitted in order to publish a family name; it is fitted in order to read a level off it, and a shape that is wrong by a tenth matters exactly as much as the extrapolation it is put through. That is not something this sweep can say, and it is the reason the essay that reads a level with no data in it exists — where the estimate turns out to be nearly unbiased and the error grows sixfold anyway.
And the whole measurement is of records whose parent is known. Every bias reported here is a difference from a number that is arithmetic. A real record arrives with no such number, which is precisely why the rate at which the middle and the tail converge has to be priced rather than assumed, and it is why the light tail’s slow arrival is worth an essay of its own.
The rate this essay did not derive
One thing was opened here and left open. The bias is measured at five block sizes and shown to fall; multiplied by the logarithm of the block it reads −0.3835, −0.4931, −0.5201, −0.5574, −0.5273, which is near constant from a hundred readings up and visibly not flat at ten. So the rate is consistent with one over the logarithm of the block and is not shown to be it — unlike the same product for the distance between the two laws, where five orders of magnitude leave it constant to a quarter of a per cent.
Settling it means deriving the second-order term for a normal maximum rather than measuring it, and the reason to want that is not tidiness. A bias whose rate is known can be corrected, and correcting it would move the light-tailed family call further than any amount of extra data does. What is here instead is the weaker claim the measurement supports: the bias is about the block, it falls slowly, and a longer record does not touch it.
Two shapes out of three are read correctly from twenty blocks, and the third is read incorrectly with rising confidence for ever. Which is the one every record of rainfall, wind speed and river level in the world is assumed to have.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The clustering the tail has — both name fréchet law, shape parameter
- The run length a declustering chooses — both name fréchet law, shape parameter
Named objects
A flat tag is an object no other essay names yet.
Block maximaBlock sizeCentral limit theoremDomain of attractionExtreme value theoryFréchet lawGeneralised extreme valueGumbel lawMaximum likelihoodProbability-weighted momentsSampling distributionShape parameterTail indexUpper endpointWeibull law