How slow a return a sample can see
Worth reading first: Three series and a count.
Two essays in this field have been about pairs with nothing between them and the rate at which a procedure invents a relation. The field’s one positive result runs the other way: when a pair is genuinely tied together, the levels regression is not merely acceptable but unusually good, and the fitted relation converges at rather than .
That result assumes the pair has been identified as tied. The test that identifies it is the one with no table, read against a critical value simulated under the null, and by every ordinary standard it works: at two hundred observations it calls two free random walks related 5.00% of the time, which is exactly what it promises.
The question nobody asks of it is what it does to pairs that are tied, and the answer depends on one quantity. A pair is tied by a mechanism that undoes a share of any disagreement each period, and the natural way to state that share is as a half-life: how long a shock to the relation takes to decay to half its size. At two hundred observations a gap that halves in five steps is found 80.40% of the time. One that halves in fifty is found 4.95% of the time, which is the rate at which the same test finds pairs with no mechanism at all.
Beyond a point the curves are flat at the test’s own size, and a flat power curve at the size is not a weak test. It is the absence of a test.
The gap between those two readings is the subject. Everything between them is a pair that has a mechanism, has a relation, and would be described correctly by the positive result — and about which two hundred observations say nothing.
What a half-life is here
The mechanism this field uses to generate a tied pair is the one Granger’s theorem says every such pair can be written with. One series wanders freely. The other’s change each period includes a term pulling it back towards the first by a fraction of however far apart they have drifted. Call that fraction α: at α = −0.2 a fifth of any disagreement is undone in a period, so the gap decays like 0.8 per step and halves in 3.11 of them.
The half-life is the same information stated as a length of time, and it is the form the question actually arrives in. Nobody knows what α = −0.034 means. “A disagreement takes twenty periods to halve” is a sentence about the world, and it can be compared directly with the length of the data — which is what makes the comparison this essay is about possible at all.
The three gaps are drawn on one set of axes deliberately, because each scaled to its own range would make all three look alike and the first one is genuinely different. A gap held inside a narrow band is what a fast return produces, and it is visible without any statistic at all — which is the easy case, and the case textbook examples are chosen from.
The middle series in that figure has a mechanism. It will return; it is doing so; the process that generated it is not the process that generated the third. Over two hundred and sixty steps it drifts about as far as the pair with no mechanism whatever, because a return that takes fifty periods to accomplish half of its work is, inside a sample of a few hundred, a return that has not yet happened.
The same non-rejection, twice
A test’s output is a verdict, and a verdict is the same words whatever produced it. That is the whole of the problem here.
Those two rows are the same number. One of them is a test correctly holding its size and the other is a test completely failing to find something that is there, and the sentence written from either is “no evidence of a long-run relationship”.
That sentence is not false in the second case. It is precisely true and it is uninformative, and the difference matters because of how such a sentence gets used. A non-rejection at a half-life of three steps is real evidence: the test would have found it 99.90% of the time, so not finding it is strong information. A non-rejection at fifty steps is no evidence at all about the mechanism, and nothing in the output distinguishes the two.
The distributions underneath say it more starkly than the rates do.
A test is a rule for separating two distributions. Where there are not two distributions, no rule separates them, and no choice of critical value, no correction, and no amount of care in running the procedure changes that. The limit is in the data.
The boundary moves with the sample, not with its square root
Reading a threshold off the power curves gives a number with a useful shape. At one hundred observations a gap has to halve within about 2.37 steps to be found four times in five. At two hundred the figure is 5.02 and at four hundred it is 9.83.
Each doubling of the sample roughly doubles the slowest return that can be seen. That is a factor of two where an ordinary estimator, converging with the square root of the sample size, would give 1.41, and it is not a coincidence: it is the 1/n convergence of the fitted relation arriving as a statement about what can be detected rather than about what can be estimated. The statistic accumulates evidence about the relation at a rate proportional to the sample rather than to its square root, so the detectable α shrinks like 1/n and the half-life it corresponds to grows like n.
Stated the other way round the rule is easier to carry. The boundary is a constant share of the sample: 2.4% of it at one hundred observations, 2.5% at two hundred and 2.5% at four hundred for four-in-five power, and 3.5%, 3.5% and 3.6% for even chances. So the test sees a return whose half-life is about a fortieth of the data, and no arrangement of the sample changes that fraction — which is why the boundary doubles when the sample does.
That framing also makes the comparison a reader needs cheap. A quarterly series covering fifty years is two hundred observations, and a fortieth of it is five quarters. A relation whose disagreements take a year and a quarter to halve is at the edge of what that series can establish; one that takes three years is not in it at all, however clean the data and however carefully the test is run.
The practical reading is better than the usual square-root arithmetic and still sobering. Doubling a sample is a genuine doubling of reach here. But the half-lives that matter in practice are often measured in years while samples are measured in quarters, and a boundary that reads 5.02 periods at two hundred observations is a boundary a great many real relations sit on the wrong side of.
The flat region is not a tuning problem
It is worth being explicit about what the flat tail of those curves is, because every instinct a reader has about a low-powered test points the wrong way here.
A test with low power against a particular alternative can usually be improved: a better statistic, a one-sided version, a more efficient estimator of a nuisance quantity. What the histogram figure shows is not low power in that sense. The statistic’s distribution under a fifty-period return and its distribution under no mechanism at all are substantially the same distribution. No rule defined on that statistic separates them, because separation is a property of the pair of distributions and not of the rule.
The deeper reason is that the null and the slow alternative are adjacent in a way the two ends of an ordinary hypothesis test are not. α = 0 is the null, and a half-life of fifty steps is α = −0.0138. The alternative is not a different kind of process; it is the same process at a parameter value a hair from the null, and a point null with the alternative pressed against it is the standard shape of a question a finite sample cannot answer. What makes it striking here is the units: −0.0138 sounds like nothing and “halves in fifty periods” sounds like a mechanism, and they are the same number.
This is also why no choice of significance level helps. Loosening the critical value to 10% raises both the null rate and the slow-alternative rate by about the same amount, because both distributions are putting mass in the same place; the ratio a reader cares about barely moves. The test is not mis-calibrated and it is not badly chosen. There is nothing in two hundred observations to calibrate against.
What the studies that do find something report
The pairs a test finds are not a random sample of the pairs that exist, and that has a measurable consequence for the number those studies print.
Two separate effects are stacked in those bars and it is worth keeping them apart.
The estimate is biased before any selection. Over every pair, found or not, the fitted adjustment gives a half-life of 8.9 steps against a truth of 12. A small negative coefficient estimated from two hundred observations is pushed away from zero, for the same reason every autoregressive coefficient near its boundary is biased, and that alone shortens the reported half-life by a quarter.
The selection roughly doubles the effect. Among the 16.5% of pairs the test actually called cointegrated, the fitted half-life is 5.9 steps — less than half the truth. The mechanism is direct: a pair is found when its gap happened to close quickly in that sample, and the speed is then read off the same closing. It is the winner’s curse with a time constant instead of an effect size.
So the literature a field like this produces reports adjustment that is faster than reality, and it reports it about the subset of relations that were fast enough to find. Both halves push the same way, and neither is visible in a single study’s output.
What a defensible reading looks like
Report the power the sample had, beside the verdict. A non-rejection is a claim about evidence and its strength is entirely determined by what could have been found. Stating “at this length the test would have found a half-life of five steps four times in five” turns a bare verdict into an interval of speeds the data can speak about — and it is computable in advance, from the sample length alone, before the data are looked at.
Treat a found relation’s speed as an upper bound on its own pace. Given the selection above, the sensible reading of a reported half-life of five steps is that the true one is longer, by something like the factor measured here. A confidence interval for α computed in the usual way does not carry that, because it conditions on nothing.
Prefer the half-life to the coefficient in anything written down. A reported α of −0.034 invites no comparison with anything; a reported half-life of twenty periods invites immediate comparison with the length of the series, which is the comparison that decides whether the result means anything. This is the same move the threshold essay makes in a different field: state the quantity on the scale where a reader can price it.
And do not read a failure to find as an absence. This is the operational form of the whole essay. Where the boundary falls relative to the process’s actual pace decides whether a null result is informative, and that comparison needs a number the study already has — its own length — and a number it can generate, which is the curve above.
What is claimed here and what is not
One generating mechanism, with everything else held at its simplest. The pairs are built from a single error-correction equation with no lagged differences, no constant in the long-run relation, unit disturbance variances and no contemporaneous correlation between the two innovations. Each of those is a dial that moves the power curve. Richer dynamics generally cost power, so the boundaries above should be read as favourable rather than typical.
“Found” means rejected at 5% against the simulated value. A study that used 10%, or that reported a p-value and left the reader to judge, would have a different rate and the same shape, since both distributions shift together. Nothing here depends on 5% being the right level; it depends on there being a level at all, which is what a verdict written from a rule always implies.
The test is Engle–Granger, which is not the most powerful available. A system-based test that estimates the relation and its adjustment jointly is known to do better, and a test built for the case where the long-run coefficient is known does better still. The shape of the result — a boundary in half-life, moving in proportion to n — is a property of what the data contain rather than of this procedure; the particular numbers are a property of this procedure and would improve under a better one. What would not change is that a flat region exists, because at a slow enough return there is nothing in a finite sample to be powerful about.
The critical values are simulated at each length. They read −3.39, −3.38 and −3.35 at one hundred, two hundred and four hundred observations, and each is generated under two free walks at that length rather than taken from a table. The measured sizes are 5.13%, 4.33% and 5.13%, which is what makes the columns above power curves rather than drifting rejection rates.
The constant fraction is a reading of three points, not a theorem. The boundary lands at 2.4%, 2.5% and 2.5% of the sample across a fourfold range of lengths, which is a strong enough agreement to state and is still three readings. What is derived rather than fitted is the exponent — a detectable α proportional to 1/n follows from the 1/n convergence of the relation itself — and the constant in front of it is measured here rather than computed.
And the interpolation is an interpolation. The detectable half-life is read between two settings of a sweep on a log scale, so 5.02 should be read as “about five”. The claim it supports — that the boundary roughly doubles when the sample doubles — is robust to that, because the two ratios are 2.12 and 1.96 and a reading error of a half-step in either direction leaves both well away from 1.41.
Still open: what to do with a pair on the wrong side of the boundary
The boundary is a fact about evidence rather than about a procedure, so the useful question is what a study does when its relation is slower than its sample can see. Three responses are available and none of them is measured here.
The relation could be imposed rather than estimated. Where theory supplies the long-run coefficient, the gap it implies is an observable series and testing whether that series returns is a much easier problem than testing whether some linear combination does — the search over combinations is most of what the critical value is paying for. How much of the boundary that recovers is a single measurement and it is the one worth making first.
The sample could be widened rather than lengthened. A relation too slow to see in one pair may be visible across many pairs that share it, at the cost of an assumption that they do. That trades a question about time for a question about homogeneity, which is a different subject and not obviously a better bargain.
Or the search could be charged for explicitly. Most of the distance between this test’s critical value and the ordinary one is the price of having chosen the long-run coefficient to make the residual look as stationary as it could — the same manufactured evidence a searched break has to be charged for, in a continuous form. Whether that charge can be reduced by restricting the search, and what the boundary becomes if it is, is a question with a clear shape and no measurement behind it here.
Or the question could be changed to one the data can answer. “Do these series return to a relation” is a statement about the limit; “is the gap between them narrower than it would be without a mechanism, over this horizon” is a statement about the sample, and the second has power where the first has none. Whether that second question is worth asking — whether a horizon-bounded version of the claim is a weaker result or a more honest one — is a judgement this field has not made, and it is the judgement the whole boundary is pressing for.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The cliff that is a slope — both name critical value, half-life, monte carlo, random walk, spurious regression, statistical power
- Which series does the moving — both name cointegration, the error-correction model, random walk, speed of adjustment, spurious regression
- The charge that is not a sum — both name critical value, monte carlo, selection effect, statistical power
- The cost of differencing a pair — both name cointegration, the error-correction model, random walk, spurious regression
- Which mistake about the rank costs — both name cointegration, the error-correction model, monte carlo, random walk
- Which series goes on the left — both name cointegration, critical value, random walk, spurious regression
Named objects
A flat tag is an object no other essay names yet.
CointegrationComposite nullConvergence rateCritical valueEngle–GrangerThe error-correction modelHalf-lifeMonte CarloRandom walkSelection effectSpeed of adjustmentSpurious regressionStatistical powerThe winner's curse