Concept

Estimation error — where it appears

The part of a forecast's error that comes from the coefficients having been fitted rather than known, which costs about σ² per parameter per row of data. It is the term that makes a bigger model worse out of sample at a true null, and it is arithmetic on two integers and a window length.

Named by 29 essays across 13 fields — each of them below, with the objects they name alongside it.

The price of each thing the rule is not told. What each rule gives up against the best model available, at a persistence of 0.85 on a fifteen-candidate table, over 400 draws. Reading down: least squares with the ordinary penalty; the whitening at the true ρ; the same at a ρ̂ estimated per candidate; that rule with the term the Gaussian likelihood carries and it omits; a Bartlett-tapered Ω̂ estimated once from the fullest candidate at L = 8; the same estimated per candidate; and the truncated Ω̂, which exists on only 45.0% of draws and is averaged over those. Knowing ρ recovers 89.9% of what counting rows gives up, estimating it 83.4%, and estimating a whole covariance 75.6%.

A covariance with no parameter in it

The whitening that repairs a criterion is told the dependence is a first-order autoregression and left to find one number. A real dependence is not one number, and the obvious estimate of it is not a covariance matrix.

banded · Dependence
One factor moves and the other does not. The two factors of the same average, each drawn against its own largest value so that they share an axis. The rate at which the five candidates disagree about the tuning parameter rises from 28.6% at 4 values on the list to 43.3% at 8, a factor of 1.52. What a disagreement costs, given that there was one, is 0.00975 ± 0.00224 and 0.00848 ± 0.00113 at the same two points — 0.5 standard errors apart, and the paired comparison on the draws that disagree under both lists puts it the other way. The guess this field was written to test was that a longer list makes disagreements commoner and each one smaller. The first half is right and there is no second half.

A rate times a size

A sweep reported what it costs to let every candidate choose its own tuning parameter and found it flat across the list. It was reporting a product, and the two things multiplied together do not behave the same way at all.

apiece · Order-selection
What each variant loses before anything has been searched for. The mean loss differential of each of the eight variants against the benchmark, over 600 tables of 60 origins, with every fit given 71 rows. The series is an AR(1) and every variant adds a lag whose coefficient is zero, so in population the two forecasts are the same forecast and the difference drawn here is estimation noise and nothing else. The marked line is σ²(q₁ − q₀)/n = -0.01408, which is an expression in how many coefficients each model has and how many rows it was fitted on — it knows nothing about the series, the persistence or which lag the variant added, and every bar is within a fifth of it. This is the amount a reference distribution recentred at each column's own sample mean believes the candidates are already behind by.

A table of nested models

A benchmark and eight variants of it, each adding one thing. Every variant is behind before the search begins, by an amount that can be written down before the data exists — and the two most natural ways of reading the table are wrong in opposite directions.

search · Forecast
A quantile is the dearer reading, everywhere. The error each rule and window delivers on the two error readings, over 400 draws. The lower pair of lines is the implied long-run variance — the instrument the earlier field uses — and the upper pair is the 95% point of the standardised resampled mean, read against the finite-sample truth of 3.889 found by simulating the law directly. The quantile costs more at every one of the eight cells: at the plug-in rule it is 59.1% against 45.3% for the taper. That is not a defect in the bootstrap; a quantile is a statement about the shape of a distribution as well as its scale, and a fixed number of resamples estimates a tail worse than a variance. What matters for the comparison is that the two orderings between the windows are not the same, which the margins figure is about.

The instrument and the reading

Every comparison between two block windows in this collection is an error in an implied long-run variance. Nobody reads a long-run variance. Read on the 95% point a test uses, the same bootstrap costs half as much again.

readout · Bootstrap
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

The width a band is measured in

A tapered covariance band spends 84% of its own weights at two lags and 74% at thirty. Every charge in the collection is a straight line through the origin in those weights, so it is too dear at one end and too cheap at the other.

curve · Criterion
The curvature is in the denominator. The optimism a Bartlett band of each width actually costs, divided by that width, on two ways of measuring the width, over 2000 draws at 120 rows. Measured in the weights the band spends — Σ w(k), which is what the earlier field levies its charges on — the reading falls from 0.9528 at two lags to 0.7486 at thirty, so a charge proportional to the summed weights is too dear at one end and too cheap at the other. Measured in the pairs the band uses — Σ w(k)(1 − k/n), because a lag of k is an average over n − k products — the same readings are flat from 4 lags up, at 0.0084 of χ² per width against 0.2359. The correction has no fitted parameter in it: it is a function of the window, the width and the sample size.

A lag the sample has less of

A sample autocovariance at lag k is an average over n − k products, not n. Count a band's width in the pairs it actually has and the curvature in its charge goes away, on a correction with nothing fitted in it.

curve · Criterion
Two instruments, two block lengths. The block length that would actually have been best on each draw, for each of the two error readings, averaged over 400 samples of 120 rows. For the rectangular window the implied long-run variance wants 18.92 and the 95% point wants 16.05; for the tapered window, 21.82 against 17.74. The quantile wants a shorter block under both windows — a ratio of 0.848 and 0.813. That is the mechanism the whole field turns on: a rule for choosing a block length is a way of guessing a target, and the two instruments do not have the same target. A rule tuned to one is systematically long for the other, and the two windows do not pay the same price for being long.

A length for each instrument

The block length that is best for an implied variance is 18.92; the one best for the 95% point of the same resamples is 16.05. A rule is a way of guessing a target, and there are two targets.

readout · Bootstrap
What a disagreement costs, split on whether it decided anything. The regret from choosing the tuning parameter per candidate, on the draws where the candidates disagreed, split on whether the disagreement changed which candidate the table selects. Over 1200 draws at each list length: when the winner changes the regret is 0.02215, 0.03029, 0.03145; when it does not it is -0.00243, -0.00069, -0.00065 — negative, and small enough that it is inside two standard errors of nothing at every length. The whole of the cost lives in the first column, and the second column is not merely small but slightly the wrong sign: when the table's answer is unaffected, letting each candidate use its own window is a very slightly better rule than making them share one. So a disagreement about the tuning parameter is not a cost. A disagreement that changes the winner is.

The quarrel that changes the winner

A disagreement about the tuning parameter costs 0.031 when it changes which candidate the table selects and −0.0007 when it does not. The distance between the values disagreed about has nothing to do with it.

apiece · Order-selection
A window with two wrong ends. The regret of a rule whitened by a Bartlett-tapered Ω̂, as the window widens, against three rules that need no window at all. At L = 0 the estimate is the identity and the rule is exactly least squares — 0.08556, the same number to five places. It falls to 0.01754 at L = 20 and rises again by L = 30, because a quarter of the sample's lags are then being estimated from it. The automatic bandwidth a practitioner would reach for, 4(n/100) to the power 2/9, which at this sample size is 4, gives 0.02963 — 69% above the best available window. The rule told the dependence is an AR(1) sits at 0.01440 throughout, which is the price of not knowing the form.

The window that has to be chosen, and the term that was dropped

An estimated covariance has a bandwidth in it, and both ends of the dial are wrong for different reasons. The rule a practitioner would reach for is two thirds worse than the best window there is.

banded · Order-selection
One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.

A curve that is a binning

A forecaster with no miscalibration in it at all reads 0.001429 at five bins and 0.014100 at fifty, on the same five hundred forecasts. The closed form is K/n times the forecaster's own irreducible score, and subtracting it returns zero.

calibrate · Calibration
A line in the right width beats two curves. How far each candidate charge sits from the measured optimism across the plateau, in units of each width's own standard error, over 2000 draws. The straight line through the origin in the band's summed weights — which is what the earlier field levies — misses by 0.2382 per width. The same straight line in the pairs the band actually uses, Σ w(k)(1 − k/n), misses by 0.0095. A fitted power law misses by 0.0293 and a fitted decaying rate by 0.0172, both on one fitted constant more. The deferral this field answers asked for a curve; the answer is a line, in a variable with nothing fitted in it.

A line that beats two curves

A deferral asked for a curve. Fitted against the same measurements, a straight line in a variable nobody had to fit describes the plateau better than either curve does with a constant more — and for three windows out of four it does not.

curve · Criterion
Cut the charge and the width follows it. The band width each charge picks, averaged over 400 draws of 120 rows under AR(1) at 0.8, with the standard deviation across draws beside it. Schwarz's charge — half a log n a lag, which is 2.39 here — picks 3.67. Akaike's picks 6.02. Charging the numbers the window actually leaves free, which is half the width, picks 10.12; charging what the optimism measures, 0.767 of that, picks 14.15. A charge and the width it buys are very nearly reciprocal, which is what a likelihood rising at a fixed rate a lag implies and is why the four answers span a factor of 3.86. The width that was actually best on the draw averages 13.90 and moves by 10.30 from draw to draw — three times as much as any rule's answer does.

A width that moves and an error that does not

Four charges give four widths a factor of four apart and four errors half a per cent apart. The derived charge wins, significantly, by a quarter of what was on offer — and none of the four is an estimate of anything.

dimension · Criterion
The quantity that does not depend on the list. The probability that letting each candidate choose its own tuning parameter changes which candidate the table selects — the product of the two moving shares — against the length of the list, over 1200 draws apiece. It is 14.2%, 11.9%, 12.3%: a spread of 2.2% across a list length that moves the disagreement rate by a factor of 1.52. This is the invariant the whole field turns on. Everything downstream of the winner — the coefficients, the regret, whatever a reader is going to quote — is a function of whether the winner changed, and how often that happens is not something the list controls. A longer list changes how often the candidates quarrel and not how often the quarrel matters.

How often it matters

The disagreement rate rises by half across the list and the share of disagreements that decide anything falls by nearly the same factor. Their product — how often the tuning list changes which candidate wins — sits at an eighth and does not move.

apiece · Order-selection
The row count entered twice, and a penalty is one place. Regret against the best available model as the errors are made persistent. Counting rows more than quadruples; both penalty repairs — the trace, and the scalar effective sample size — are worse than it at every persistence measured; and whitening the sample and keeping the ordinary penalty falls, recovering 86.9% of what counting rows gives up at ρ = 0.85. Doing it at an estimated ρ recovers 80.6%, so having to estimate the dependence from the rows being selected on costs 7.3% of what knowing it is worth. Mallows' forms are drawn beside the logarithmic ones and behave the same, which is what rules the linearisation out.

The repair that was exact and made it worse

A penalty computed from the trace is exactly the optimism it estimates, and selecting with it gives up a fifth more than not correcting anything. The row count entered the criterion twice, and a penalty is the second place.

effective · Forecast
The reversal is a property of the instrument. The margin between the two block windows under each of four rules, on three readings of the same resampled means, signed so that a positive bar is the tapered window winning. On the implied long-run variance the taper wins at the best available block length and at one estimated from the sample and loses at a length written into a protocol and at the rule of thumb — which is the reversal the earlier field's whole argument turns on, at 1.48 and 4.52 points. On the 95% point a test actually reads, the taper wins at all four, by 6.13 to 7.08 points. On the coverage the interval actually delivers, the taper wins at all four again, by 2.50 to 5.75 percentage points. Two of the four rules change sign between the first reading and the other two, and the two that change are exactly the two the earlier field's recommendation is about.

The reversal that was the instrument's

On an implied variance the rectangle wins at a protocol length and at the rule of thumb. On the 95% point a test reads, and on the coverage an interval delivers, the taper wins at all four rules.

readout · Bootstrap
The ranking on the left, the weights on the right. Eight moving-average forecasts of an AR(1) at φ = 0.4895, the persistence at which the best of them exactly ties the 60-observation benchmark. On the left, each candidate's expected squared error in units of the series' own variance: the smallest belongs to L = 2, at 1.0156. On the right, the weight each carries in the variance-minimising combination of all eight — and the best of them carries 0.00000. The two ends of the family carry 1.0172 of the weight between them, and the combination they make is worth 0.7817, which is 23.0% below the best single forecast. Both columns are closed forms in φ. Which forecast to keep and which forecasts to use are different questions, and this is a set where the answers share nothing.

The weight that is a vector

Two forecasts have a best combination and one number describes it. Eight have a best combination too, and the vector describing it puts nothing at all on the forecast with the smallest mean squared error.

search · Rank
The threshold buys accuracy and spends exceedances. The mean squared error of the estimated shape against the threshold, split into the square of its bias and its spread, over 600 records of 2000 readings from a a normal parent. At the 0.9 quantile 199 exceedances are left, the bias is -0.1708, the spread is 0.0701 and the total error is 0.0341. The bias falls as the threshold rises because the exceedances get closer to being generalised Pareto; the spread rises because there are fewer of them. The sum is smallest at the 0.925 quantile, at 0.0340, of which 80.6% is still bias — so even the best threshold on this grid is one where accuracy, not spread, is the binding constraint.

The threshold is a dial

A peaks-over-threshold analysis has one knob, and raising it buys accuracy with exceedances. For a normal parent the error is smallest at the 0.925 quantile and 80.6% of it is still bias there — and both diagnostics practitioners use to set the knob lose to a fixed 0.90 rule, one by a factor of 1.590 and one by 11.881.

extreme · Extremes
The distribution the table does not have. 599 series simulated from the smaller model fitted to one comparison's own data, the whole rolling comparison re-run on each, and the ordinary statistic recorded. Under this null the two forecasts are the same forecast in population, so what is left in a sample is the larger model's estimation error and the statistic is centred at -1.134 rather than at zero. Its 95% point is 0.264; the standard normal drawn behind it puts that point at 1.645. Reading this statistic against that curve is not a poor approximation, it is a different distribution: the share of this one above 1.645 is 0.2%.

A distribution drawn from the null

Between nested models the ordinary comparison statistic has a null distribution centred at minus one and a 95% point of a quarter. A correction to its mean repairs the centre and leaves the shape; simulating the null repairs both.

ranking · Bootstrap
What each construction carries, against what there was. The autocorrelation of a resampled error series at five lags, averaged over 60 samples of 40 resamples each. Three facts are in the picture. The residuals lie below the errors at every lag, which is the ceiling a multiplier cannot exceed. The blocked multiplier and the fixed-length block lie on top of each other below it — they attenuate identically, because the attenuation is the join — while the stationary bootstrap, whose runs are geometric rather than fixed, sits above them both. And the sieve is the exception in kind rather than in degree: at lag six it carries 0.0638 where the residuals have 0.0300 and the multiplier has -0.0011, because a fitted model extrapolates past the lags it was told about and a truncated sample sequence cannot.

Errors generated from a fitted model

The one construction that is not bounded by the residuals, because a model extrapolates past the lags it was told about and a truncated sample sequence cannot. It is nearly exact where the only defect is dependence, and it pays for it where there are two.

banded · Reference
The scale moves the width; the curve does not. The band width each charge picks, averaged over 150 draws of 120 rows. The two conventions — a unit a lag and half a log n a lag — pick 6.08 and 3.65 lags. The four charges derived from the measured optimism pick 15.05, 15.60, 15.47 and 15.44, against a best width on the draw of 14.13. So the scale a charge is levied on moves the width by a factor of 4.27 and the shape of the charge moves it by 3.6%. None of the six is an estimate of the draw's own best width: the correlations are -0.006, -0.003, 0.017, -0.001, 0.016, -0.004.

What a better charge buys

Four charges derived from the same measurements pick band widths within six per cent of each other and deliver errors within two per cent of the gap any of them leaves. The scale a charge is levied on decides the width; the shape of the charge decides nothing.

curve · Criterion
What the interval actually covers. The coverage of the two-sided interval each rule and window builds, over 400 samples of 120 rows, against the 95% it promises. Not one of the eight reaches it: the best is 91.0% and the worst is 80.8%, on a promise of 95%. So the first thing this instrument says is that the choice between the two windows is a choice inside a range that is already four to fourteen points short, which neither of the other two readings can express at all. The second is the ordering: the tapered window covers better under every one of the four rules, by 5.00, 2.50, 5.75 and 4.00 points — including at a length written into a protocol and at the rule of thumb, where the implied variance says the rectangle wins.

What the interval covers

Eight rules and windows, and not one of them reaches its promised 95%. The range is 80.8% to 91.0%, and the choice between two block windows is a choice inside a shortfall that is four times larger.

readout · Bootstrap
The record stops here, and the curve does not. The level exceeded once in T blocks, against T, for a normal parent at 365 readings a block. The truth is closed form — the block maximum's own distribution function is Φ(x) raised to the 365, so the T-block level is Φ⁻¹ of the 365-th root of (1 − 1/T), with nothing fitted in it — and the fitted mean over 800 records of 50 blocks sits on top of it, 4.0186 against 4.0330 at 100 blocks. What moves is not the level but its error, which grows from 0.0956 at 10 blocks to 0.6383 at a thousand while the level itself moves only from 3.4421 to 4.5454. The rule marks the largest reading an average record contains, 4.0062: everything to the right of where it crosses is read from a fit rather than from data.

A level with no data in it

The largest of fifty block maxima is a 51-block event by its own plotting position, so a hundred-block level is read 1.96 times past the longest event the record contains — and it lands above the largest reading on 52.4% of records. The estimate stays nearly unbiased out there; what grows is its error, sixfold from ten blocks to a thousand.

extreme · Extremes
The ceiling a multiplier cannot reach past. A wild-type resampling forms e*_t = e_t·w_t with the multiplier independent of the residual, so what comes out has autocovariance γ_resid(k)·γ_w(k) — the residuals' own, multiplied by the multiplier's. Since |γ_w| ≤ 1 the reference distribution's dependence is bounded above by the residuals', and the residuals' is already below the errors'. The two shortfalls compose. For a block of ℓ the multiplier's autocorrelation is exactly the triangle (1 − k/ℓ)⁺, drawn here as the dashed prediction against the realised resamples at ℓ = 5; the bound is attained only at ℓ = n, where the reference distribution is built from one sign.

What a multiplier cannot keep

Two reasons were named for the quarter a blocked resampling falls short, and taking either away makes the gap larger. What is left is a bound — a multiplier can only take dependence out, and the residuals' own is already below the errors'.

effective · Reference
Largest where least is needed. What the pairs correction supplies against what each window's measured profile needs, across this field's plateau, over 2000 draws at 120 rows. Both are stated as the multiplicative rise the charge per unit of width has to take between four lags and thirty. What the correction supplies is arithmetic — (1 − μ(4)/n)/(1 − μ(30)/n), where μ is the mean lag of the weight the band adds — and it runs 1.0795, 1.0580, 1.0456, 1.0539 for the four windows. What the measurement needs runs 1.1076, 1.2928, 1.2296, 1.6550. The two orderings are opposite: the plain Bartlett window has the longest mean lag, so it gets the biggest correction, and the flattest profile, so it needs the smallest. They coincide to 0.9746 of each other, and nowhere else does the correction account for more than 85.0% of the fall.

What the correction assumes

A correction with nothing fitted in it repairs one window of four. The reason is that its size is set by where a window puts its weight and the curvature it must repair is set by something else — and for one window at one sample size the two happen to agree.

curve · Criterion
Reading the draw changes what is charged, not what is tracked. The correlation between the band width each rule picks and the best band width on the same draw, over 400 draws. The three fixed charges read -0.069, -0.066, 0.012. The three that read the sample read -0.012, -0.019, -0.041. None of the six is distinguishable from nothing. The statistic the first plug-in reads does vary — the draw's own summed squared autocorrelation runs from 2.06 to 10.19 with a mean of 4.36 — so the failure is not that the charge stopped moving. It is that what it moves with carries no information about which width this draw wanted.

A charge that reads the draw

Three charges built to read the sample track the best band width on their own draw at −0.012, −0.019 and −0.041, deliver more error than the fixed rule they are calibrated to, and pick a width half again as variable. The statistic moves; the answer does not.

curve · Criterion
What each wrong count costs, 4 steps ahead. Squared forecast error 4 steps ahead at each imposed rank, relative to the correctly specified fit, at 200 observations. With 1 genuine relations, imposing 0 costs 13.3% and imposing 2 costs 4.8%. With 2 genuine relations, imposing 1 costs 15.6% and imposing 3 costs 2.5%. Under-counting is the more expensive mistake in both systems, and it is the one the procedure's level does not bound.

Which mistake about the rank costs

On a system with two relations, imposing none costs 29.2% of squared forecast error and imposing three costs 2.5%. The expensive mistake is under-counting, which is the error the procedure's 5% does not bound — so the guarantee protects the cheap side.

systems · Rank
One of these converges and the other does not. Two measurements on the same fits, against the sample length, for a system with 2 genuine relations. The distance from the fitted plane to the true plane falls from 0.1438 at 100 observations to 0.0075 at 1600 — halving with each doubling, which is the 1/n rate this field's estimates converge at. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and is flat in between. The plane is an estimate; the relation inside it is not.

A space is not a relation

The fitted plane approaches the true one at rate 1/n — 0.1438 at a hundred observations and 0.0075 at sixteen hundred. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and never moves.

systems · Rank
A proxy removes less than its reliability, always. The share of the confounding bias removed by adjusting for a proxy, against how well the proxy measures the confounder. The diagonal is the answer a reader would guess — a covariate that is 80% signal removes 80% of the problem. The curve is what the arithmetic gives: the reliability, times one minus the squared correlation between the treatment and the confounder, divided by one minus the product of those two. That squared correlation is 0.4475. A reliability of 0.8 removes 68.85% and one of 0.6 removes 45.32%. The two agree only at the ends, and the gap is widest where most applied covariates sit.

Adjusting for a shadow

A covariate that is 80% signal removes 68.85% of the confounding, not 80% — the share is λ(1 − ρ²)/(1 − λρ²) and it is below the reliability everywhere. The residual bias is 0.1084 against an effect of 0.5, and at 25,600 rows it is 17.6 standard errors wide.

collider · Conditioning
Two companions on one simulation, two hundredfold apart. How many times as many draws each companion is worth, on the same 4,000 simulated samples of 40 observations. The coverage of the interval is estimated with the observed count as its companion, whose expectation is 12 exactly; they correlate at 0.2665 and the companion is worth 1.08 times the draws. The expected width is estimated with p̂(1 − p̂) as its companion, whose expectation is 0.20475 exactly; they correlate at 0.9977 because the width is a monotone function of it, and the companion is worth 214 times the draws — 856 thousand simulated samples' worth of precision from four thousand.

The check worth more than the check

The same exactly known companion that verifies a simulation can sharpen it. On one set of four thousand draws, one companion is worth 1.08 times the draws and another is worth 214 times them, and the factor is 1 − ρ² with nothing else in it.

method · Routes

Named alongside it

The objects these essays reach for when they reach for this one.

Monte CarloModel selectionDependenceInformation criterionTaperingClosed formLong-run varianceBandwidth selectionPersistenceDegrees of freedomMean squared errorOptimism

All concepts