The best of a set, and what the search costs

A distribution drawn from the null

Between nested models the ordinary comparison statistic has a null distribution centred at minus one and a 95% point of a quarter. A correction to its mean repairs the centre and leaves the shape; simulating the null repairs both.

Worth reading first: Where the bootstrap lies · What the model says next.

The evaluation field ends with a comparison it cannot make. Between nested models — an order-1 autoregression against an order-2, where the smaller model is the larger one with a coefficient set to zero — the ordinary comparison statistic declares the smaller model significantly better on 67% of comparisons at a nominal 5%, and gets more certain of it the longer the evaluation runs.

The repair that field applies is the Clark–West recentring, and the honest description of it, which that essay gives, is that it is a moment correction. It computes what the loss differential’s mean must be under the null and adds it back. Whether the corrected statistic is normal is a separate question, and it was named there as unresolved.

This essay resolves it by not asking. Instead of correcting a statistic so that a table applies to it, the table is replaced: the null model is simulated, the whole comparison is re-run on each simulated series, and the observed statistic is read against the distribution that comes back.

Why there is no table to read

Before the machinery, the shape of the difficulty. Everything the evaluation field measured about a forecast comparison was about the denominator — that the loss differentials are correlated, that a long-run variance has to be estimated, that estimating it costs size at horizons where it is not needed. Every one of those is a statement about a statistic that is asymptotically normal and whose standard error is hard to get right.

The nested case is not a harder instance of that. It is a statistic that is not asymptotically normal at all, and the reason has nothing to do with dependence between origins: it would still be true if every origin were independent and the variance were known exactly.

The nested null is degenerate in a way that has to be stated before anything else makes sense. Under it, the two forecasts are the same forecast in population: the extra coefficient is zero, so the larger model’s population forecast is the smaller model’s. The loss differential’s population variance is zero, and what survives in a sample is not a difference between two forecasts but the larger model’s estimation error, squared.

A statistic whose numerator is a squared estimation error and whose denominator is an estimate of a variance that is zero has no reason to be normal, and it is not.

The distribution the table does not have. 599 series simulated from the smaller model fitted to one comparison's own data, the whole rolling comparison re-run on each, and the ordinary statistic recorded. Under this null the two forecasts are the same forecast in population, so what is left in a sample is the larger model's estimation error and the statistic is centred at -1.134 rather than at zero. Its 95% point is 0.264; the standard normal drawn behind it puts that point at 1.645. Reading this statistic against that curve is not a poor approximation, it is a different distribution: the share of this one above 1.645 is 0.2%.
Fig. 1 A thousand series simulated from the smaller model fitted to one comparison’s own data, the entire rolling comparison re-run on each, and the ordinary statistic recorded. The pale curve behind it is the standard normal it is normally read against.

The distribution is centred at −1.148, not at zero. Its 95% point is 0.238, not 1.645. The share of it above 1.645 is 0.1%, which is why the test read against a normal one-sided in favour of the larger model does not merely reject too rarely — it never rejects at all.

Reading that statistic against a normal is not a poor approximation to the right answer. It is a different distribution, and no amount of data improves the match: the shape is a consequence of the null being degenerate rather than of the sample being small.

Generating the distribution instead

The construction is the obvious one once the problem is stated. Fit the small model to the whole series. Resample its residuals and simulate a new series from it. Re-run the entire rolling comparison — every origin, both models refitted at each — and record the statistic. Do that B times, and the observed statistic has a distribution to be read against.

The null is true in every simulated series by construction, which is what makes this a reference distribution rather than a second estimate of anything. It costs B refits of two models at every origin, and it is the whole of the method.

Four readings of one comparison, where the smaller model is true. 110 nested comparisons — an order-1 model against an order-2 model on 100 origins with a window of 40 — read four ways, at a nominal 5%. The smaller model is true, so every rejection is a false one. The ordinary statistic read one-sided against a normal rejects 0.0%: not rarely, never, because its null distribution is centred at -1.11 and a normal table asks it to reach 1.645. Against its own bootstrap distribution the same statistic on the same comparisons rejects 4.5%.
Fig. 2 Two hundred nested comparisons, ninety-nine bootstrap series each, with the smaller model true. Four readings of the same statistics: two statistics, each against a normal table and against its own simulated distribution.

Read against a normal, the ordinary statistic rejects 0.0% of the time. Read against its own bootstrap distribution — the same statistic, on the same two hundred comparisons — it rejects 4.0%, which is 5% to within the noise of two hundred draws. The Clark–West recentring gives 1.5% against the normal and 2.5% against the bootstrap.

The check a rejection rate cannot make

A test can have the right rejection rate at one point of the scale by luck. The distribution of its p-values is the check that cannot be passed that way, and it is the check this site applies to every test it builds.

The check a rejection rate cannot make. The p-value distributions of four readings of the same 110 nested comparisons, against the diagonal a valid p-value has to follow. The ordinary statistic read against a normal sits 0.500 from flat — it is not slightly conservative, it is concentrated near 1, which is what a statistic whose null distribution is centred at -1.11 does. Read against its own bootstrap distribution it is 0.053 from flat, which is the sampling noise of 110 draws. The recentred statistic is at 0.127: better than the raw reading everywhere and not flat, because a correction to the mean cannot correct a shape.
Fig. 3 The p-value distributions of three readings of the same two hundred comparisons, against the diagonal a valid p-value has to follow.

The ordinary statistic against a normal sits 0.517 from flat by the Kolmogorov–Smirnov distance: not conservative, but piled up near 1, which is what a statistic whose null distribution is centred at −1.157 does when it is read against a curve centred at zero. Against its own bootstrap distribution the same statistic is 0.065 from flat, which is the sampling noise of two hundred draws of a discrete p-value taking a hundred values.

Clark–West sits at 0.133. That is a long way better than 0.517 and it is not flat, and the gap is exactly what “a moment correction cannot correct a shape” means when it is measured rather than asserted. Recentred and read against the bootstrap it reaches 0.060 — the shape repaired by the simulation and the centre by the correction, which is the only combination in the table that is right for both reasons.

What the distribution looks like, in detail

Three numbers describe it and each says something different about why the normal reading fails.

Its centre is −1.104 at the median and −1.157 at the mean, against zero. That is the part Clark–West computes: under the null the larger model’s forecast is the smaller one’s plus estimation noise, so its squared error is larger by the mean square of that noise, and the differential’s expectation is negative by exactly that amount. A correction that adds it back is doing arithmetic rather than fitting anything.

Its spread is 0.828, against 1. A studentised statistic whose denominator is estimating a variance the null says is zero does not have unit variance, and nothing in the standard construction notices.

Its shape is the part no moment corrects. The 5% and 95% points of the simulated distribution are −2.404 and 0.240, which is not symmetric about the centre: the lower tail runs 1.30 below the median and the upper tail 1.34 above it, and the mass is packed differently at the two ends because the statistic is a ratio of two quantities that are functions of the same estimation error.

A correction can move the first of those three. Reading against the simulated distribution handles all three at once, and the p-value distribution is where that difference is visible rather than arguable.

What the repair costs

A correction that buys size by spending power is a bad trade dressed as a good one, so the same four readings have to be measured where the larger model is genuinely right.

Four readings of one comparison, where the larger model is right. 110 nested comparisons — an order-1 model against an order-2 model on 100 origins with a window of 40 — read four ways, at a nominal 5%. The larger model is genuinely right, so these are power. The ordinary statistic read against a normal finds it 15.5% of the time; against its own distribution, 67.3%. The recentred statistic finds it 69.1%, so the distributional repair costs nothing that the moment repair buys.
Fig. 4 The same four readings against a real second-order coefficient of 0.3, where the larger model is the right one and every rejection is a discovery.

The ordinary statistic against a normal finds it 21.5% of the time. Against its own distribution it finds it 67.0%. Clark–West against a normal finds it 64.0%, and against the bootstrap 65.0%.

So the distributional repair costs nothing that the moment repair buys. The two arrive at nearly the same power from opposite directions — one by correcting the statistic and keeping the table, one by keeping the statistic and replacing the table — and only the second of them also has flat p-values under the null.

Four readings of one comparison, where the larger model is right110 nested comparisons — an order-1 model against an order-2 model on 100 origins with a window of 40 — read four ways, at a nominal 5%. The larger model is genuinely right, so these are power. The ordinary statistic read against a normal finds it 1.8% of the time; against its own distribution, 14.5%. The recentred statistic finds it 14.5%, so the distributional repair costs nothing that the moment repair buys.ordinary statistic, normal table1.8%ordinary statistic, own bootstrap14.5%recentred statistic, normal table14.5%recentred statistic, own bootstrap16.4%the statistic, and the distribution it is read against110 comparisons, 49 series each, P = 100these are power
Fig. 5 Drag the coefficient the smaller model is missing, from zero through the middle to a clear difference. At φ₂ = 0.15 the bootstrap reading finds the larger model 20.5% of the time against the normal reading’s 1.5%, which is the region where the choice of reference distribution is the whole of the analysis.

The middle of the range is where it matters, as it always is. At φ₂ = 0.15 the improvement is real and modest, the normal reading finds it in one comparison in sixty-seven, and the bootstrap finds it in one in five.

Why the bootstrap works here and not in the other place

This site has an essay about a bootstrap that does not work: where the bootstrap lies measures the resampled distribution of a maximum and finds it wrong in a way more data does not fix, because the maximum of a resample can never exceed the maximum of the sample.

The difference is worth stating precisely, because “the bootstrap” is being used for two different things. There, the resampling is asked to reproduce the sampling distribution of a statistic at the truth, and it fails because the statistic is a functional the empirical distribution cannot represent. Here, the resampling is asked to produce the distribution of a statistic under a stated null, and the null is a fitted parametric model that can be simulated from exactly.

The second task is easier and is the one being done. Nothing here is estimating the true distribution of anything; it is enumerating what the statistic would have looked like had the small model been right, which is the same construction the exact-inference field uses when it re-runs an assignment rule, with a fitted model in place of a known mechanism. That substitution is the whole of what is being assumed, and it is the reason this is called a reference distribution rather than an exact test.

The check a rejection rate cannot make. The p-value distributions of four readings of the same 110 nested comparisons, against the diagonal a valid p-value has to follow. The ordinary statistic read against a normal sits 0.386 from flat — it is not slightly conservative, it is concentrated near 1, which is what a statistic whose null distribution is centred at 0.80 does. Read against its own bootstrap distribution it is 0.718 from flat, which is the sampling noise of 110 draws. The recentred statistic is at 0.703: better than the raw reading everywhere and not flat, because a correction to the mean cannot correct a shape.
Fig. 6 The p-value distributions where the larger model is right. All three bend upward, which is what power looks like on this axis, and the ordering is the one the rejection rates already reported.

Two moments and a shape, priced separately

“A moment correction cannot correct a shape” is the sentence the p-value distances stand behind, and the three numbers describing the simulated distribution are enough to say how much of Clark–West’s residual failure is the shape and how much is a moment it does not touch.

The recentring repairs the mean and leaves the spread at 0.828. A statistic with that spread, read against a standard normal at 1.645, rejects when it exceeds 1.645/0.828 = 1.99 of its own standard deviations, which for a normal would be 2.3% rather than 5%. So of the 3.5 percentage points the corrected statistic gives up — 1.5% realised against a nominal 5% — about 2.7 points are the variance and the remaining 0.8 is everything else. The largest part of what the correction leaves behind is a second moment rather than a shape.

Standardising the simulated distribution by both of its moments says what the shape is worth on its own. Its 95% point is 0.240 and its mean −1.157, so in units of its own standard deviation the upper tail sits at (0.240 + 1.157)/0.828 = 1.69, against a normal’s 1.645. The lower tail sits at (−2.404 + 1.157)/0.828 = −1.51, against −1.645. The right tail is about 3% heavier than normal and the left about 8% lighter — visible, one-sided, and small next to the factor of 1.21 the spread alone contributes.

That decomposition is worth having because it says what a cheaper repair could achieve. A correction that computed the null’s variance as well as its mean — arithmetic of the same kind, on the same degenerate null — would land within a few per cent of the right critical value and would still be applicable to a published statistic, which the simulation is not. It would not have flat p-values, because the skew is real; it would be far closer than 0.133.

None of that argues against generating the distribution, which handles all three at once and costs a laptop’s afternoon. It argues that the gap between the two methods is smaller than the p-value distances suggest, and that most of it is a quantity nobody has computed rather than a shape nobody can.

What a reader can check without the series

The simulation cannot be applied to a table, and that limitation has a consequence worth stating for the reader on the receiving end of one.

What is available from a published nested comparison is the sign and rough size of the statistic against what the null distribution’s centre would be. A reported statistic of −1.1 in a nested comparison is not weak evidence against the larger model; it is almost exactly what the null predicts, since the simulated median is −1.104. A reported statistic of 0.0 — which reads as “no difference” against a normal table — sits at about the 80th percentile of this null and is mild evidence for the larger model.

So a reader who knows only that the comparison is nested can already relabel the axis. Zero is not the neutral point, 1.645 is not the 5% point, and a statistic anywhere near either is being read against the wrong scale by a full unit — which is the same relabelling a flat p-value is supposed to make unnecessary and, in this one case, does not.

The two repairs are answering different questions

It is worth being precise about what each construction assumes, because they are not two routes to one number in the way this site usually means that phrase.

The recentring assumes the form of the null’s effect on the mean: it computes E[d] under the nested null and subtracts it, which requires the null to be exactly the one written down. It assumes nothing about the distribution and delivers nothing about it either.

The simulation assumes the model: that the fitted first-order autoregression is the process generating the series under the null, so that draws from it are draws from what the world would have looked like. It delivers the whole distribution and is wrong in a different way if that model is wrong.

Where they agree — and here they agree closely on power, 65.0% against 64.0% — the agreement is worth something precisely because the assumptions are different. Where they disagree, which is under the null and in the p-value distribution, the disagreement is the shape term, and the shape term is not a quantity the recentring ever claimed to handle. The uniformity check is what turns that from a statement about what each method promises into a measurement of what each delivers.

What it costs to run

The honest cost of this construction is that it re-runs the entire comparison B times per experiment. At forty observations of estimation window and a hundred origins, that is a hundred refits of two models per bootstrap replication and ten thousand per comparison — which is nothing on a laptop and is why the method is worth having rather than admiring.

What it is not is a repair that can be applied afterwards to a published table. The bootstrap needs the series, not the summary: it re-runs the forecasting exercise, so it can only be done by whoever did the exercise. A moment correction can be applied to a reported statistic; a reference distribution cannot. That is an argument for the correction as a reporting convention and for the simulation as the analysis, and the two are not in competition.

What is claimed here, and what is not

This essay claims a reference distribution for a nested forecast comparison, generated by simulating the fitted null model, and the measurements that say what it is worth: its size, its p-value distribution, and its power against a real improvement, each beside the moment correction that is the standard alternative.

The boundary against the multiplicity essays of this field is that they are about a maximum over a set and this is about one comparison. Both end at a generated reference distribution and they generate it differently: there by resampling rows of observed differentials, which assumes only stationarity, and here by simulating a fitted model, which assumes the model. A set comparison between nested models would need both at once, and that is named as not built.

What stays out and is named as a decision: the choice of bootstrap for the innovations — residuals are resampled here, and a wild bootstrap or a parametric normal draw would each answer a slightly different question about robustness; nested comparisons at horizons past one, where the loss differential acquires the moving-average structure the comparison it comes from measures and the resampling has to carry it; and the case where the two models are non-nested but overlapping, which is a third null with a third distribution and is not covered by either method here.

The checks, and what they are checked against

Two claims are gated in this field’s library. The ordinary statistic read against a normal is required to reject essentially never at a true null, with p-values far from flat, while the same statistic read against its own bootstrap distribution is required to be at its nominal level with p-values that are flat to within the noise of the count. And against a real improvement the bootstrapped statistic is required to find it several times as often as the normal reading does and about as often as the recentred one does — because a repair that bought its size back out of power would be a different result, and the check is written so that it would have said so.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Benchmark forecastBootstrapClark–WestDiebold–MarianoEstimation errorKolmogorov–SmirnovLoss differentialMonte CarloNested modelsNull hypothesisp-valueReference distributionStatistical powerUniformity