The interval the joint fit reports
Worth reading first: The observations that repeat each other · A design is a number.
The essay that priced a joint fit over a general covariance measured what practitioners experience as a risk averaged over draws, and set aside what they actually read off the output: the standard error printed beside each coefficient. It named the omission as a decision. A risk and an interval are different kinds of claim, and a fit whose covariance was maximised rather than plugged in might report its uncertainty correctly or might not; putting an unchecked interval next to a measured risk would have mixed the two.
This essay makes the check. The regression is the one the whole field runs: four predictors with lag-one correlations of 0.9, 0.6, 0.3 and nothing, true coefficients of 0.30, 0.20, 0.12 and zero, a hundred and twenty rows, errors drawn from each of the four laws in turn. Six rules fit it and each prints the usual interval — the coefficient plus or minus 1.96 standard errors, the standard error being what generalised least squares reports when the covariance it whitened by is taken as known. The question is how often that interval contains the true coefficient, over a thousand draws under each law.
The answer splits the rules cleanly. Five of them are about as honest as anyone would expect. The sixth is the joint fit, which wins on risk where it wins at all, and whose interval is the least honest of any rule that whitens.
Six intervals, one of them honest by right
Begin with the law the joint fit was found to help: a five-period moving average, whose dependence stops dead after four lags and is therefore exactly a band.
The law’s own covariance is the reference, and it does what the theory says it should: 94.9% on the persistent predictor and 94.1% on the white one. That is the oracle rule — it whitens by the true covariance and so is entitled to treat it as known — and its coverage is the evidence that the computation is right. Everything else is read against it.
Least squares, which treats the errors as independent, covers 65.0% on the persistent predictor and 94.4% on the white one. That is the oldest result in this field: a dependence in the errors matters to the standard error only through a dependence in the predictor, so it ruins the interval for the column that persists and leaves the column that does not almost untouched.
A fitted first-order correlation covers 93.0% and 96.3%, and the two-step band plug-in 86.5% and 97.3%. Both whiten, so both repair most of what least squares got wrong on the persistent predictor, and neither damages the white one.
The joint fit covers 73.2% on the persistent predictor and 71.4% on the white one. It is short on both, and on the white predictor it is short by twenty-four points where every other rule is within three of nominal. That is the result, and the rest of the essay is about where it comes from.
It is worth saying what that coverage means in the form a reader meets it. The white predictor’s true coefficient is zero, so its interval missing the truth is the same event as its t-test rejecting a true null. An analyst testing whether that column matters, at the conventional 5%, would be told it does on 28.6% of samples under the moving average if they used the joint fit — and on 2.7% if they used the two-step band, whose interval over-covers. One of those is a test at about its stated level and erring on the safe side; the other finds an effect that is not there more than one time in four.
It is not one law’s
The moving average is the law on which the joint fit has the most to offer, so it might be the law on which it overreaches. Run the same comparison under the first-order autoregression these regressions are mostly written under.
The oracle covers 94.1% and 95.5%. Least squares covers 56.9% on the persistent predictor and 94.6% on the white one; a fitted ρ̂ 90.6% and 96.0%; the two-step band 83.7% and 96.5%. The joint fit covers 79.4% and 83.5%.
The shortfall is smaller here and has the same shape: the joint fit is the only rule short on the white predictor. Under all four laws the pattern repeats, and the figure below reads the white predictor alone, because that is the coefficient on which no other rule has anything to hide.
The oracle covers 95.5%, 94.1%, 95.3% and 95.4% under the autoregression, the moving average, long memory and the break; the two-step band covers 96.5%, 97.3%, 94.9% and 96.4%. The joint fit covers 83.5%, 71.4%, 85.3% and 79.5%. Under no law does it come within nine points of what it promises, and it falls furthest under the moving average — the one law on which it was measured to beat the two-step on risk.
What it reports and what it delivers
A coverage figure is a symptom. Its cause is always the same: the standard error printed and the actual spread of the estimate across samples disagree. So set the two side by side, for the white predictor under the moving average.
Least squares reports 0.0884 and delivers 0.0885, which is why its interval covers on this column. The two-step band reports 0.0436 and delivers 0.0376 — a little pessimistic, which is why it over-covers. The oracle reports 0.0173 and delivers 0.0177.
The joint fit delivers 0.0289. That is a genuine improvement on the two-step, and it is the improvement the earlier essay measured as a risk: the joint estimate of a white predictor’s coefficient really is tighter. But the standard error it reports is 0.0167 — 97% of what the oracle reports. The joint fit prints nearly the precision of a rule that knows the covariance exactly, and it delivers about 44% of the oracle’s advantage over the two-step.
Put the other way round: its actual standard deviation is 1.585 times the root-mean-square of the standard errors it reports. Every rule that whitens by an estimate misstates its precision a little; the two-step on this column misstates it in the safe direction, at 0.858. The joint fit misstates it by more than half in the unsafe one.
Not always the oracle’s precision
“Nearly the oracle’s precision” is a description of the moving average, and it would be easy to promote it into a mechanism it is not. Under the break it fails.
The law with a break in its persistence is the one law here whose covariance no band contains: the dependence changes half way through the sample, so the true covariance is not Toeplitz and the oracle — whitening by it exactly — has an advantage nothing banded can approach. On the white predictor the oracle reports 0.0286 and delivers 0.0279. The two-step band reports 0.0472 and delivers 0.0441. The joint fit reports 0.0355 and delivers 0.0523.
So under the break the joint fit does not claim the oracle’s precision; it claims about a quarter more standard error than the oracle does. What it still does is claim more precision than it has: it reports a standard error two thirds the size of its spread, while the two-step reports one a little larger than its own. The joint fit’s spread here is no better than the two-step’s — 0.0523 against 0.0441, worse in fact — and its interval is narrower. That is the general statement, and the moving average is its sharpest instance rather than its definition: the joint fit’s reported precision runs ahead of its delivered precision under every law, by a ratio of 1.316, 1.585, 1.291 and 1.421 on the white predictor, against the two-step’s 0.920, 0.858, 0.982 and 0.929.
The persistent predictor, where everybody is short
The white predictor isolates the joint fit’s problem because nothing else is wrong there. The persistent predictor is the harder column, and every feasible rule is short on it.
Under the autoregression the actual spread of the persistent predictor’s estimate is 1.126 times the root-mean-square reported standard error for a fitted ρ̂, 1.332 for the two-step band and 1.430 for the joint fit; the oracle reads 1.007. Under long memory the fitted ρ̂ and the two-step cover 85.3% and 84.9%, and the joint fit 83.0%.
Two different shortfalls are stacked on this column. The first is shared by every rule that estimates the dependence: a predictor with a lag-one correlation of 0.9 regressed against errors with their own persistence is the setting in which an effective sample size collapses furthest, and an estimate of the covariance that is slightly wrong moves the variance of the coefficient a great deal. Under long memory, where no band of eight lags contains the law, the two-step and a fitted ρ̂ are short by the same ten points for the same reason. The second shortfall is the joint fit’s own, and it is the one already measured on the white predictor. On the persistent column it adds a few points to a gap that was already there, which is why the white predictor is the cleaner place to see it.
Why the maximum overstates what it knows
The reported standard error is on the whitened design — the generalised least-squares formula with the covariance treated as fixed. Each rule plugs in its own covariance, and the formula has no term for the fact that the covariance was estimated from the same hundred and twenty rows. That omission is the same for every rule. What differs is what the omission is hiding.
The two-step band’s covariance is tapered, and a taper shrinks every sample autocovariance towards zero. So the two-step believes the errors are less dependent than they are — a fact the field has measured from the likelihood’s side, where the tapered plug-in sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. For a white predictor, believing in less dependence means believing the predictor is less informative relative to the noise, and so a larger standard error. The two-step’s bias in its covariance becomes a cushion in its interval.
The joint fit removes much of the bias. Under the autoregression its first band coefficient averages 0.714, against the two-step’s 0.635 and a true value of 0.8, and its log-determinant averages −103.5 against −64.5. That is what it was built to do, and the essay that compared iteration with maximisation showed that re-reading the correlation from the generalised residuals recovers most of that kind of bias and the likelihood’s own determinant term the rest. But removing the bias removes the cushion as well, and what is left is the estimation noise in eight band coefficients fitted to a hundred and twenty rows, which the formula never priced for any rule.
The estimate it settles on is not only less biased; it is also selected. A maximum over a band of eight free numbers is drawn towards the covariance that makes this sample’s residuals look most independent after whitening, and a sample’s residuals looking more independent than its errors are is exactly the condition under which a standard error comes out too small. The fit that takes the memory out found the same thing one level down, with one correlation and residuals instead of errors; here it is eight coefficients and a likelihood.
The further the maximum goes, the less it covers
That reading makes a prediction: the draws on which the maximum moves furthest from the plug-in should be the draws on which the interval covers least. Sort the thousand draws under the autoregression by the log-determinant of the joint fit’s band — the more negative, the closer the band has been pushed towards a singular matrix — and read coverage of the white predictor in each quarter.
Coverage runs 75.6%, 84.0%, 84.4% and 90.0% from the quarter nearest the edge to the quarter furthest from it. In the nearest quarter the reported standard error is 77% of the oracle’s own — smaller than what a rule that knew the covariance would report — and in the furthest it is 94%. The prediction holds: where the maximum goes furthest, the interval it reports is narrowest and wrongest.
Even the furthest quarter covers only 90.0%, so this is not a story about a few degenerate draws. It is a gradient across all of them.
Counting the lags is not the repair
The cheapest repair available is the one any textbook would reach for first: the joint fit estimated eight band coefficients, so charge them to the residual degrees of freedom and read the interval against Student’s t on what remains. That moves the critical value from 1.96 to 1.982 and the variance estimate up by eight rows’ worth.
It buys a point or two. On the white predictor it moves coverage from 83.5% to 84.9% under the autoregression, 71.4% to 73.5% under the moving average, 85.3% to 86.1% under long memory and 79.5% to 81.0% under the break. Every one is still short by more than eight points.
The reason is visible in the critical value that would actually have given 95% coverage — the 95th percentile of the absolute t-ratio across draws. For the oracle it is 1.92 on the white predictor under the autoregression, close to 1.96 as it should be. For the joint fit it is 3.73, and under the moving average 4.90. A rule whose t-ratio needs a critical value of nearly five is not a rule that miscounted its degrees of freedom by eight. Its t-ratio has heavy tails, because its denominator moves from draw to draw with the covariance it maximised, and the draws where the denominator is smallest are the draws where the numerator is not.
What stays honest
None of this makes the joint fit a bad estimator. On the moving average it delivers the tightest feasible estimate of every coefficient on the table, and on risk it is the one rule that beats the two-step on a law where it should. What it cannot do is report its own precision with the formula every other rule uses.
The rules that do report honestly are worth naming. The oracle is honest by right, and no analyst has it. A fitted first-order correlation covers within two points of nominal on three of four predictors under every law here and falls short on the persistent one by between two and ten points — its one estimated number is cheap to be wrong about. The two-step band is honest on everything but the persistent predictor, partly by design and partly by the accident of its taper.
The joint fit is the one rule here whose estimate improves and whose interval worsens, and the two are the same event. It moves the covariance towards the truth and towards this sample at once, and the formula reads both movements as knowledge.
Still open: an interval that knows the band was maximised
The next measurement is the repair the degrees-of-freedom count could not supply. Two are standard and neither is measured here. A sandwich estimate replaces with an expression that reads the whitened residuals’ own dependence, so a band that whitened too eagerly is caught by what it left behind; it costs nothing to compute and its honesty on a hundred and twenty rows is an open question, since the check that reads a residual’s dependence misses a lag-one correlation of 0.2 four times in five. A parametric bootstrap redraws from the fitted band, refits the joint maximum on each redraw, and reads the spread directly; it prices the maximisation by repeating it, and costs thirty joint fits per interval at least. Which of the two restores 95% under all four laws — and whether either does so without giving back the risk the joint fit earned — is a distinct question with a measurable answer, and it is the one this field should take next.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A family before a fit — both name covariance matrix, dependence, generalised least squares, long memory, nuisance parameter, profile likelihood, tapering, whitening
- The comparison that was not made — both name covariance matrix, degrees of freedom, dependence, monte carlo, nuisance parameter, tapering, whitening
- A charge that depends on the rule — both name dependence, monte carlo, nuisance parameter, profile likelihood, structural break, whitening
- A dependence with a shape — both name covariance matrix, generalised least squares, long memory, nuisance parameter, structural break, whitening
- Nothing in the fit picks the width — both name covariance matrix, degrees of freedom, nuisance parameter, profile likelihood, tapering, whitening
- The window that has to be chosen, and the term that was dropped — both name covariance matrix, degrees of freedom, dependence, generalised least squares, nuisance parameter, tapering
Named objects
A flat tag is an object no other essay names yet.
Confidence intervalCovariance matrixDegrees of freedomDependenceGeneralised least squaresLong memoryMonte CarloNuisance parameterParametric bootstrapProfile likelihoodSandwich estimatorStandard errorStructural breakTaperingWhitening