The tail past the last observation

A level with no data in it

The largest of fifty block maxima is a 51-block event by its own plotting position, so a hundred-block level is read 1.96 times past the longest event the record contains — and it lands above the largest reading on 52.4% of records. The estimate stays nearly unbiased out there; what grows is its error, sixfold from ten blocks to a thousand.

Worth reading first: Three shapes, one limit.

Fifty blocks of three hundred and sixty-five readings — a record of fifty years of daily observations, which is a record somebody actually has. Fit a generalised extreme value law to the fifty maxima and read off the level exceeded once per hundred blocks. On 52.4% of eight hundred such records, that level is above the largest reading anywhere in the record. At a thousand blocks it is above it on 99.5%.

That is not a criticism of the method; it is a description of what the method is for. A hundred-year level read from fifty years of data is by construction a statement about something not observed, and the only alternative on offer is to decline to state it. What is worth measuring is how much of the number is data and how much is law, and the honest version of the question has an answer: the largest of mm block maxima is, by its own plotting position, an (m+1)(m+1)-block event. A fifty-block record reaches 51 blocks. A hundred-block level is therefore read 1.96 times past the longest event the record contains, and it is the standard reading.

The truth needs no fit at all

The measurement here is possible because the parent is known, and a return level under a known parent is arithmetic. A block maximum’s own distribution function is F(x)bF(x)^b, so the level exceeded with probability 1/T1/T per block is

zT=F1 ⁣((11/T)1/b)z_T = F^{-1}\!\left((1 - 1/T)^{1/b}\right)

with nothing estimated in it. For a normal parent at 365 readings a block that gives 3.4421 at ten blocks, 4.0330 at a hundred and 4.5454 at a thousand. Every error reported below is a distance from one of those three numbers rather than a disagreement between two estimates, which is what makes it an error rather than a comparison.

The fitted route is the one a practitioner takes:

z^T=μ^+σ^ξ^{(log(11/T))ξ^1}\hat z_T = \hat\mu + \frac{\hat\sigma}{\hat\xi}\left\{\left(-\log(1 - 1/T)\right)^{-\hat\xi} - 1\right\}

Three estimated parameters, and the shape enters through an exponent — which is the fact that governs everything from here to the end of the field.

What the curve does past the record

The record stops here, and the curve does not. The level exceeded once in T blocks, against T, for a normal parent at 365 readings a block. The truth is closed form — the block maximum's own distribution function is Φ(x) raised to the 365, so the T-block level is Φ⁻¹ of the 365-th root of (1 − 1/T), with nothing fitted in it — and the fitted mean over 800 records of 50 blocks sits on top of it, 4.0186 against 4.0330 at 100 blocks. What moves is not the level but its error, which grows from 0.0956 at 10 blocks to 0.6383 at a thousand while the level itself moves only from 3.4421 to 4.5454. The rule marks the largest reading an average record contains, 4.0062: everything to the right of where it crosses is read from a fit rather than from data.
Fig. 1 The level exceeded once in T blocks, against T, for a normal parent at 365 readings a block. The fitted mean over eight hundred records sits on the closed-form truth — 4.0186 against 4.0330 at a hundred blocks — and the rule marks the largest reading an average record contains, 4.0062.

The two curves are on top of each other. At a hundred blocks the fitted mean is 4.0186 against a truth of 4.0330, a bias of −0.0144; at a thousand it is 4.5496 against 4.5454, a bias of +0.0042. Nothing about the extrapolation goes systematically wrong as it is pushed out, and a reader looking only at the means would conclude that a fitted law reads a thousand-block level from fifty blocks about as well as it reads a ten-block one.

What is not on top of itself is the error. The root mean squared error of the fitted level runs 0.0956 at ten blocks, 0.1474, 0.2050, 0.2779, 0.3663, 0.5088 and 0.6383 at a thousand — a factor of 6.68. Over the same range the level itself moves only from 3.4421 to 4.5454, which is a rise of 32%. So the quantity being estimated grows by a third and the error on it grows sixfold, and the ratio between them worsens at every step.

The rule on the figure is the other half of the reading. The largest reading in an average record is 4.0062, and the hundred-block level is 4.0330. The crossing is almost exactly at the standard reading, which is another way of saying the same thing the 52.4% says: at T=100T = 100 a record is being asked for a number it is, on average, exactly at the edge of containing.

How much of it is above anything observed

How often the quoted level is above everything ever recorded. The share of 800 records of 50 blocks on which the fitted return level sits above the largest reading in the record itself. The largest of 50 block maxima is, by its plotting position, a 51-block event — so a 100-block level is being read 1.96 times past the longest event the record contains, and it lands above that reading 52.4% of the time. Below the record's own reach the picture is different: a 10-block level is above the largest reading on 0.0% of records. The note beside each bar is the median share of the rise from a typical block maximum to the quoted level that lies above anything observed — 28.0% at a thousand blocks.
Fig. 2 The share of records on which the fitted level sits above the largest reading in the record itself. It is 0.0% at ten blocks, 13.5% at fifty, 52.4% at a hundred and 99.5% at a thousand, and the note beside each bar is the median share of the rise that lies above anything observed — 28.0% at a thousand blocks.

The share above the record’s own maximum runs 0.0%, 0.5%, 13.5%, 52.4%, 81.9%, 98.3%, 99.5%. Ten blocks is a reading the data contain on every single record. Fifty blocks — which is the record’s own reach, the level the plotting position says the largest reading is an instance of — is above the maximum on one record in seven, which is the noise around a boundary rather than a failure. A hundred blocks is a coin toss. A thousand blocks is never in the data.

The second measurement on that figure is finer, and it is the one that resists the objection that “above the maximum” is a crude binary. Take the rise from a typical block maximum to the quoted level, and ask how much of that rise sits above anything observed:

share(T)=z^Txmaxz^Tmedianx\text{share}(T) = \frac{\hat z_T - x_{\max}}{\hat z_T - \operatorname{median} x}

At ten blocks the median share is −99.5%: the level sits below the record’s maximum by about as far as it sits above the record’s middle, so it is squarely inside. At a hundred blocks it is 0.6% — the level is at the edge and essentially none of the rise is extrapolation. At two hundred blocks it is 11.6%, at five hundred 22.2%, and at a thousand it is 28.0%. So even the thousand-block level is three-quarters anchored in observed range and one quarter beyond it, which is a more temperate reading than “99.5% of records quote a level above everything ever seen” and is the same measurement.

Both numbers are worth having because they answer different objections. The binary says the quoted level is not a reading anybody has taken. The share says how much of the distance from the ordinary to the quoted is being covered by the law rather than by the record.

Does fitting a law buy anything

Past fifty blocks there is nothing to check the fit against. Below it there is: the (11/T)(1 - 1/T) quantile of the fifty block maxima themselves exists for TT up to about fifty and is undefined past it. Both routes can be scored against the closed-form truth exactly where they overlap.

Inside the record's reach, the fitted law wins. The root mean squared error of two routes to a return level, at the three periods a record of 50 blocks can answer empirically, over 800 records. The fitted route reads the level off a generalised extreme value fit; the empirical route takes the (1 − 1/T) quantile of the block maxima themselves and stops existing past 50 blocks. The fit is the better estimate at every one of the three — 0.0956 against 0.1178, 0.1474 against 0.1638, 0.2050 against 0.2225 — which is what a law buys where the data reach. What it buys past 50 blocks cannot be read off this figure, because the comparison it would need does not exist: the error of the fitted route at a thousand blocks is 0.6383 and there is nothing to put beside it.
Fig. 3 The error of two routes to a return level at the three periods a fifty-block record can answer empirically. The fitted law is better at all three — 0.0956 against 0.1178, 0.1474 against 0.1638, 0.2050 against 0.2225 — and past fifty blocks the empirical route stops existing.

The fitted route wins at all three: 0.0956 against 0.1178 at ten blocks, 0.1474 against 0.1638 at twenty-five, 0.2050 against 0.2225 at fifty. So a law buys something real — roughly a fifth off the error at the shortest period and a twelfth at the longest — and it buys it by borrowing strength from the whole record rather than from the two or three maxima nearest the quantile being read.

That is the honest case for the method, and it is worth stating precisely because the rest of this essay is about where the case stops. It is a demonstration that the law is a better description of the data than the data are of themselves, inside the range where both can be checked. It is not evidence about anything outside that range, and the reason it cannot be is structural: the comparison requires an empirical quantile, and past the record’s own reach there is no empirical quantile to compare with. The error of the fitted route at a thousand blocks is 0.6383, and there is nothing to put beside it.

This is exactly the boundary the argument about a resampling that cannot see past its own data draws in a different subject, and it is the same boundary a fitted autoregression’s order draws in lag: the fit reproduces what it was fitted on and continues on its own past it, and the continuation is the part being read.

What the number rests on once the data stop

A fitted level is a function of three estimates, and past the record’s reach it stops being a function of them equally. The derivatives of the level in the three parameters are written out rather than differenced, and they say how much of the estimate each parameter is carrying:

zTμ=1,zTσ=yξ1ξ,zTξ=σξ2(yξ1)σξyξlogy\frac{\partial z_T}{\partial \mu} = 1, \qquad \frac{\partial z_T}{\partial \sigma} = \frac{y^{-\xi} - 1}{\xi}, \qquad \frac{\partial z_T}{\partial \xi} = -\frac{\sigma}{\xi^{2}}\left(y^{-\xi} - 1\right) - \frac{\sigma}{\xi} y^{-\xi}\log y

with y=log(11/T)y = -\log(1 - 1/T). Evaluated at the mean fit over the eight hundred records — μ^=2.7828\hat\mu = 2.7828, σ^=0.3182\hat\sigma = 0.3182, ξ^=0.0902\hat\xi = -0.0902 — the three derivatives read 1, 2.0368 and 0.7046 at ten blocks, and 1, 5.1413 and 5.0677 at a thousand. The location’s lever is one at every period, by construction. The shape’s lever grows by a factor of seven.

Multiply each by the spread of its own estimate across records and the sum comes apart into three contributions. The location contributes 0.0538 at every period — the same 0.0538 at ten blocks and at a thousand, because its lever never moves. The scale contributes 0.0773, 0.1429 and 0.1952. The shape contributes 0.0819, 0.2983 and 0.5892, against total errors of 0.0956, 0.2779 and 0.6383.

So the answer to what the number rests on is that it changes hands. At ten blocks the three parameters contribute comparably and the level is a summary of the record. At a thousand blocks the shape carries 92% of the length of the error and the location carries a twelfth of it, and the level has become a statement about one badly-determined exponent with the record supplying an origin for it.

That reallocation is the mechanism behind the sixfold error growth, and it is why the growth cannot be bought off with more blocks. Adding blocks shrinks all three spreads at the same 1/m1/\sqrt{m} rate; it does not touch the levers, and the lever on the worst-determined parameter is the one that grows. It is also why the fitted mean sits so close to the truth: μ^\hat\mu and σ^\hat\sigma are nearly unbiased and the shape’s bias of about a tenth is multiplied by a lever that is small exactly where the level is checkable and large exactly where it is not.

The fitted shape is worth reading on its own. Across eight hundred records of 365-reading blocks it averages −0.0902 where the parent’s limit shape is zero, which is the same bias the essay that counts how often the family is named correctly finds at every record length. Every quoted level in this essay is therefore read off a law with a slight ceiling in it, and the ceiling is not a feature of the world.

Nearly unbiased is not nearly right

The most misleading number in the sweep is the bias, and it is misleading because it is small.

At a hundred blocks the fitted level is short by −0.0144 on average and its error is 0.2779 — the bias is a nineteenth of the error. At a thousand blocks the bias is +0.0042 and the error is 0.6383, so the bias is a hundred and fiftieth of it. Every one of the seven periods has this shape. An analyst who checked the extrapolation by running it on synthetic records and comparing means would find it unimpeachable at every period, and would be checking the one property that is fine.

One estimate, two intervals, and only one of them is lopsided. The mean endpoints of two 95% intervals for the 100-block level, at four record lengths, over 300 records apiece. The true level is 4.0330 and is marked. The delta-method interval is symmetric about the estimate by construction and its upper endpoint at 25 blocks reaches only 4.8406; the profile-likelihood interval reaches 6.2154, with an upper arm 5.69 times its lower one. That lopsidedness is not a defect: the shape parameter enters a return level through an exponent, so the likelihood for a far-out level is skewed and an interval that is not skewed with it has to miss on one side. The asymmetry falls to 1.75 by 200 blocks.
Fig. 4 The mean endpoints of two intervals for the hundred-block level, at four record lengths, against the true level of 4.0330. At twenty-five blocks a symmetric interval reaches 4.8406 above and 3.1816 below, and the likelihood’s own interval reaches 6.2154.

What the small bias hides is how lopsided the estimate is. At twenty-five blocks the mean symmetric interval runs from 3.1816 to 4.8406 around a true level of 4.0330, and the interval the likelihood itself supports runs from 3.5993 to 6.2154 — an upper arm 5.686 times the lower one. A distribution that skewed has a mean near its target and a median well below it, and the mean is the statistic that reports the bias.

The reason is the exponent. The shape enters the return level as yξ^y^{-\hat\xi}, so a shape estimate that is off by a tenth changes the level by a factor rather than by an amount, and by a factor that grows with the period. A symmetric error in ξ^\hat\xi therefore becomes an asymmetric error in z^T\hat z_T, with a long arm upward, and the mean of that asymmetric error is near zero while nearly every individual record sits below the truth. Counting the coverage of both intervals is what turns that observation into a number, and the number is that 99.24% of the symmetric interval’s misses are on one side.

What a reader is entitled to conclude

A hundred-block level with no observations under it is not nothing, and it is not what it is usually taken for. Three statements are supported by the measurements above and a fourth is not.

It is a statement about a shape, conditional on a family. The level is what a generalised extreme value law with the fitted parameters says, and the largest term in it past the record’s reach is an exponent estimated at −0.0902 with a spread of 0.1163. Every claim the level makes inherits both of those. A reader who does not accept the family does not get a weaker version of the number; they get no number at all.

It is better than the alternative inside the range where both exist, by a fifth of the error at ten blocks and a twelfth at fifty, and there is no reason to think the ordering reverses outside that range. That is a real warrant and it is a weak one: it is evidence about the fit’s behaviour where the data are, offered in support of a reading taken where they are not.

Its error is quotable and grows in a known direction. The error at ten blocks is 0.0956 and at a thousand it is 0.6383, and the growth is not noise but the shape’s lever. A level quoted without that error attached has had the one honest part of it removed. What that error is worth as a promise — whether an interval built from it covers at the rate it claims — is the question two intervals for the same level answers, and the answer is that the interval reported first covers 80.3% at twenty-five blocks.

And it is not checkable within the record’s own lifetime. This is the statement that is not supported, and it is arithmetic rather than opinion. A TT-block level is checked by counting exceedances, and over a further mm blocks the expected count is m/Tm/T: another fifty years against a hundred-block level is half an exceedance expected. Observing zero or one is consistent with the level being right, with its being a fifth too high, and with its being a third too low. The quantity is not falsifiable on any horizon a record can offer, which is exactly why the bias and the error above had to be measured against a parent that was chosen rather than found.

The last of those is the reason this field is arranged the way it is. A number nobody can check by waiting is a number whose only guarantee is procedural — the 95% belongs to the procedure and not to the interval in front of anybody — and a procedural guarantee that has never been counted is a promise, not a property. The same substitution runs through a forecast band derived for known parameters and then computed from estimates, which covers 87.3% rather than 95% when somebody counts it.

Where the error comes from

The shape estimate narrows, and one of them narrows onto the wrong number. The middle ninety per cent of the estimated shape parameter, over 400 records apiece, for three parents whose limits are a Fréchet, a Gumbel and a Weibull. Every band narrows as the record lengthens — the heavy-tailed parent's from 0.9734 wide at 20 blocks to 0.1375 at 500 — and two of the three narrow onto the shape their parent actually has. The light-tailed parent's narrows onto -0.1036 rather than onto zero, because a maximum of 100 normal readings is not yet at its limit and the fit reads the shape of what it was given.
Fig. 5 The middle ninety per cent of the estimated shape at five record lengths. At fifty blocks the light-tailed parent’s estimate spans about a third of a unit, and it is that spread the exponent in a return level magnifies.

Almost all of it is the shape. At fifty blocks a light-tailed parent’s shape estimate has a spread of about 0.12 and a middle ninety per cent spanning roughly a third of a unit, and that is the quantity a return level exponentiates. The location and scale are estimated much more stably — they are essentially a mean and a spread of fifty numbers — and neither of them is raised to a power.

So the error growth from 0.0956 to 0.6383 is not seven independent things going wrong at seven periods. It is one badly-determined number entering an expression whose sensitivity to it grows with TT, which is why the growth is smooth and why it does not level off. The essay that counts how often the family behind that number is named correctly is the same measurement one step earlier: at fifty blocks a light-tailed record’s shape is called correctly 75.5% of the time, and the 24.5% that are not are the records whose extrapolation is a different curve.

What is being scored, and what is not

One thing about the design deserves stating because it cuts against the neatness of the closed form.

The truth used here is the parent’s exact return level, not the return level of the limit law. Those are different numbers. A block of 365 normal readings is not Gumbel — it is Gumbel to within a distance that falls like one over the logarithm of the block, and at 365 readings that distance is around two hundredths — so the law being fitted is not the law the data came from, and the fitted level is being scored against the level the data actually have.

That is the right choice: a practitioner wants the level, not the level the asymptotic law would have had. But it means the small biases above are the net of two things, a fit that is slightly short-tailed relative to the parent and a parent that is slightly short-tailed relative to the limit, and nothing here separates them. A sweep at a different block size would, and it was not run.

One tail arrives; the other is still on its way at a million. The Kolmogorov distance between the exact law of a normalised maximum and its Gumbel limit, at six block sizes, for two parents that both have that same limit. Both are closed form: the exact law of a maximum is F(x)^n and no simulation is involved. The exponential parent's distance falls from 0.0280 to 2.707e-7 — a factor of a hundred thousand, which is exactly one over n. The normal parent's falls from 0.0522 only to 0.0091, a factor of 5.74, because its rate is one over log n. At a million readings a block the two differ by a factor of 33556.3.
Fig. 6 The distance between the exact law of a normalised maximum and its Gumbel limit, at six block sizes, for two parents that share that limit. At the block sizes a real record uses, a normal parent’s maximum is still measurably short of its limit.

The second thing not scored is the one that matters most in practice and cannot be measured here at all. Every number in this essay is an error against a truth that exists because the parent was chosen. A real fifty-year record comes with no such truth, and the quantity an analyst would like — how far the quoted level is from the level the world actually has — is not estimable from the record by any method, because it is precisely the thing the record does not contain. What is estimable is the interval, and an interval is a promise about a rate. Whether a stated rate is the rate that is delivered is a question with a countable answer, and it is the only remaining check on a number with no data in it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Block maximaBlock sizeClosed formEstimation errorExtrapolationExtreme value theoryGeneralised extreme valueGumbel lawMean squared errorOrder statisticReturn levelReturn periodSampling distributionShape parameterUpper endpoint