A distribution drawn from the null
Worth reading first: Where the bootstrap lies · What the model says next.
The evaluation field ends with a comparison it cannot make. Between nested models — an order-1 autoregression against an order-2, where the smaller model is the larger one with a coefficient set to zero — the ordinary comparison statistic declares the smaller model significantly better on 67% of comparisons at a nominal 5%, and gets more certain of it the longer the evaluation runs.
The repair that field applies is the Clark–West recentring, and the honest description of it, which that essay gives, is that it is a moment correction. It computes what the loss differential’s mean must be under the null and adds it back. Whether the corrected statistic is normal is a separate question, and it was named there as unresolved.
This essay resolves it by not asking. Instead of correcting a statistic so that a table applies to it, the table is replaced: the null model is simulated, the whole comparison is re-run on each simulated series, and the observed statistic is read against the distribution that comes back.
Why there is no table to read
Before the machinery, the shape of the difficulty. Everything the evaluation field measured about a forecast comparison was about the denominator — that the loss differentials are correlated, that a long-run variance has to be estimated, that estimating it costs size at horizons where it is not needed. Every one of those is a statement about a statistic that is asymptotically normal and whose standard error is hard to get right.
The nested case is not a harder instance of that. It is a statistic that is not asymptotically normal at all, and the reason has nothing to do with dependence between origins: it would still be true if every origin were independent and the variance were known exactly.
The nested null is degenerate in a way that has to be stated before anything else makes sense. Under it, the two forecasts are the same forecast in population: the extra coefficient is zero, so the larger model’s population forecast is the smaller model’s. The loss differential’s population variance is zero, and what survives in a sample is not a difference between two forecasts but the larger model’s estimation error, squared.
A statistic whose numerator is a squared estimation error and whose denominator is an estimate of a variance that is zero has no reason to be normal, and it is not.
The distribution is centred at −1.148, not at zero. Its 95% point is 0.238, not 1.645. The share of it above 1.645 is 0.1%, which is why the test read against a normal one-sided in favour of the larger model does not merely reject too rarely — it never rejects at all.
Reading that statistic against a normal is not a poor approximation to the right answer. It is a different distribution, and no amount of data improves the match: the shape is a consequence of the null being degenerate rather than of the sample being small.
Generating the distribution instead
The construction is the obvious one once the problem is stated. Fit the small model to the whole series. Resample its residuals and simulate a new series from it. Re-run the entire rolling comparison — every origin, both models refitted at each — and record the statistic. Do that B times, and the observed statistic has a distribution to be read against.
The null is true in every simulated series by construction, which is what makes this a reference distribution rather than a second estimate of anything. It costs B refits of two models at every origin, and it is the whole of the method.
Read against a normal, the ordinary statistic rejects 0.0% of the time. Read against its own bootstrap distribution — the same statistic, on the same two hundred comparisons — it rejects 4.0%, which is 5% to within the noise of two hundred draws. The Clark–West recentring gives 1.5% against the normal and 2.5% against the bootstrap.
The check a rejection rate cannot make
A test can have the right rejection rate at one point of the scale by luck. The distribution of its p-values is the check that cannot be passed that way, and it is the check this site applies to every test it builds.
The ordinary statistic against a normal sits 0.517 from flat by the Kolmogorov–Smirnov distance: not conservative, but piled up near 1, which is what a statistic whose null distribution is centred at −1.157 does when it is read against a curve centred at zero. Against its own bootstrap distribution the same statistic is 0.065 from flat, which is the sampling noise of two hundred draws of a discrete p-value taking a hundred values.
Clark–West sits at 0.133. That is a long way better than 0.517 and it is not flat, and the gap is exactly what “a moment correction cannot correct a shape” means when it is measured rather than asserted. Recentred and read against the bootstrap it reaches 0.060 — the shape repaired by the simulation and the centre by the correction, which is the only combination in the table that is right for both reasons.
What the distribution looks like, in detail
Three numbers describe it and each says something different about why the normal reading fails.
Its centre is −1.104 at the median and −1.157 at the mean, against zero. That is the part Clark–West computes: under the null the larger model’s forecast is the smaller one’s plus estimation noise, so its squared error is larger by the mean square of that noise, and the differential’s expectation is negative by exactly that amount. A correction that adds it back is doing arithmetic rather than fitting anything.
Its spread is 0.828, against 1. A studentised statistic whose denominator is estimating a variance the null says is zero does not have unit variance, and nothing in the standard construction notices.
Its shape is the part no moment corrects. The 5% and 95% points of the simulated distribution are −2.404 and 0.240, which is not symmetric about the centre: the lower tail runs 1.30 below the median and the upper tail 1.34 above it, and the mass is packed differently at the two ends because the statistic is a ratio of two quantities that are functions of the same estimation error.
A correction can move the first of those three. Reading against the simulated distribution handles all three at once, and the p-value distribution is where that difference is visible rather than arguable.
What the repair costs
A correction that buys size by spending power is a bad trade dressed as a good one, so the same four readings have to be measured where the larger model is genuinely right.
The ordinary statistic against a normal finds it 21.5% of the time. Against its own distribution it finds it 67.0%. Clark–West against a normal finds it 64.0%, and against the bootstrap 65.0%.
So the distributional repair costs nothing that the moment repair buys. The two arrive at nearly the same power from opposite directions — one by correcting the statistic and keeping the table, one by keeping the statistic and replacing the table — and only the second of them also has flat p-values under the null.
The middle of the range is where it matters, as it always is. At φ₂ = 0.15 the improvement is real and modest, the normal reading finds it in one comparison in sixty-seven, and the bootstrap finds it in one in five.
Why the bootstrap works here and not in the other place
This site has an essay about a bootstrap that does not work: where the bootstrap lies measures the resampled distribution of a maximum and finds it wrong in a way more data does not fix, because the maximum of a resample can never exceed the maximum of the sample.
The difference is worth stating precisely, because “the bootstrap” is being used for two different things. There, the resampling is asked to reproduce the sampling distribution of a statistic at the truth, and it fails because the statistic is a functional the empirical distribution cannot represent. Here, the resampling is asked to produce the distribution of a statistic under a stated null, and the null is a fitted parametric model that can be simulated from exactly.
The second task is easier and is the one being done. Nothing here is estimating the true distribution of anything; it is enumerating what the statistic would have looked like had the small model been right, which is the same construction the exact-inference field uses when it re-runs an assignment rule, with a fitted model in place of a known mechanism. That substitution is the whole of what is being assumed, and it is the reason this is called a reference distribution rather than an exact test.
Two moments and a shape, priced separately
“A moment correction cannot correct a shape” is the sentence the p-value distances stand behind, and the three numbers describing the simulated distribution are enough to say how much of Clark–West’s residual failure is the shape and how much is a moment it does not touch.
The recentring repairs the mean and leaves the spread at 0.828. A statistic with that spread, read against a standard normal at 1.645, rejects when it exceeds 1.645/0.828 = 1.99 of its own standard deviations, which for a normal would be 2.3% rather than 5%. So of the 3.5 percentage points the corrected statistic gives up — 1.5% realised against a nominal 5% — about 2.7 points are the variance and the remaining 0.8 is everything else. The largest part of what the correction leaves behind is a second moment rather than a shape.
Standardising the simulated distribution by both of its moments says what the shape is worth on its own. Its 95% point is 0.240 and its mean −1.157, so in units of its own standard deviation the upper tail sits at (0.240 + 1.157)/0.828 = 1.69, against a normal’s 1.645. The lower tail sits at (−2.404 + 1.157)/0.828 = −1.51, against −1.645. The right tail is about 3% heavier than normal and the left about 8% lighter — visible, one-sided, and small next to the factor of 1.21 the spread alone contributes.
That decomposition is worth having because it says what a cheaper repair could achieve. A correction that computed the null’s variance as well as its mean — arithmetic of the same kind, on the same degenerate null — would land within a few per cent of the right critical value and would still be applicable to a published statistic, which the simulation is not. It would not have flat p-values, because the skew is real; it would be far closer than 0.133.
None of that argues against generating the distribution, which handles all three at once and costs a laptop’s afternoon. It argues that the gap between the two methods is smaller than the p-value distances suggest, and that most of it is a quantity nobody has computed rather than a shape nobody can.
What a reader can check without the series
The simulation cannot be applied to a table, and that limitation has a consequence worth stating for the reader on the receiving end of one.
What is available from a published nested comparison is the sign and rough size of the statistic against what the null distribution’s centre would be. A reported statistic of −1.1 in a nested comparison is not weak evidence against the larger model; it is almost exactly what the null predicts, since the simulated median is −1.104. A reported statistic of 0.0 — which reads as “no difference” against a normal table — sits at about the 80th percentile of this null and is mild evidence for the larger model.
So a reader who knows only that the comparison is nested can already relabel the axis. Zero is not the neutral point, 1.645 is not the 5% point, and a statistic anywhere near either is being read against the wrong scale by a full unit — which is the same relabelling a flat p-value is supposed to make unnecessary and, in this one case, does not.
The two repairs are answering different questions
It is worth being precise about what each construction assumes, because they are not two routes to one number in the way this site usually means that phrase.
The recentring assumes the form of the null’s effect on the mean: it computes E[d] under the nested null and subtracts it, which requires the null to be exactly the one written down. It assumes nothing about the distribution and delivers nothing about it either.
The simulation assumes the model: that the fitted first-order autoregression is the process generating the series under the null, so that draws from it are draws from what the world would have looked like. It delivers the whole distribution and is wrong in a different way if that model is wrong.
Where they agree — and here they agree closely on power, 65.0% against 64.0% — the agreement is worth something precisely because the assumptions are different. Where they disagree, which is under the null and in the p-value distribution, the disagreement is the shape term, and the shape term is not a quantity the recentring ever claimed to handle. The uniformity check is what turns that from a statement about what each method promises into a measurement of what each delivers.
What it costs to run
The honest cost of this construction is that it re-runs the entire comparison B times per experiment. At forty observations of estimation window and a hundred origins, that is a hundred refits of two models per bootstrap replication and ten thousand per comparison — which is nothing on a laptop and is why the method is worth having rather than admiring.
What it is not is a repair that can be applied afterwards to a published table. The bootstrap needs the series, not the summary: it re-runs the forecasting exercise, so it can only be done by whoever did the exercise. A moment correction can be applied to a reported statistic; a reference distribution cannot. That is an argument for the correction as a reporting convention and for the simulation as the analysis, and the two are not in competition.
What is claimed here, and what is not
This essay claims a reference distribution for a nested forecast comparison, generated by simulating the fitted null model, and the measurements that say what it is worth: its size, its p-value distribution, and its power against a real improvement, each beside the moment correction that is the standard alternative.
The boundary against the multiplicity essays of this field is that they are about a maximum over a set and this is about one comparison. Both end at a generated reference distribution and they generate it differently: there by resampling rows of observed differentials, which assumes only stationarity, and here by simulating a fitted model, which assumes the model. A set comparison between nested models would need both at once, and that is named as not built.
What stays out and is named as a decision: the choice of bootstrap for the innovations — residuals are resampled here, and a wild bootstrap or a parametric normal draw would each answer a slightly different question about robustness; nested comparisons at horizons past one, where the loss differential acquires the moving-average structure the comparison it comes from measures and the resampling has to carry it; and the case where the two models are non-nested but overlapping, which is a third null with a third distribution and is not covered by either method here.
The checks, and what they are checked against
Two claims are gated in this field’s library. The ordinary statistic read against a normal is required to reject essentially never at a true null, with p-values far from flat, while the same statistic read against its own bootstrap distribution is required to be at its nominal level with p-values that are flat to within the noise of the count. And against a real improvement the bootstrapped statistic is required to find it several times as often as the normal reading does and about as often as the recentred one does — because a repair that bought its size back out of power would be a different result, and the check is written so that it would have said so.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The corner the test is calibrated at — both name benchmark forecast, bootstrap, monte carlo, null hypothesis, reference distribution, statistical power
- A table of nested models — both name benchmark forecast, clark–west, estimation error, loss differential, nested models
- Estimating how many nulls are true — both name monte carlo, null hypothesis, p-value, statistical power, uniformity
- Two ways to combine p-values — both name kolmogorov–smirnov, monte carlo, p-value, statistical power, uniformity
- When the benchmark is a candidate — both name benchmark forecast, monte carlo, nested models, null hypothesis, reference distribution
- The analysis after three arms — both name monte carlo, null hypothesis, reference distribution, statistical power
Named objects
A flat tag is an object no other essay names yet.
Benchmark forecastBootstrapClark–WestDiebold–MarianoEstimation errorKolmogorov–SmirnovLoss differentialMonte CarloNested modelsNull hypothesisp-valueReference distributionStatistical powerUniformity