A slow return across many pairs
Worth reading first: Three series and a count.
A relation supplied in advance measured the first way round the boundary on how slow a return a sample can see: stop estimating the long-run coefficient and test the gap a theory implies. It bought a third more half-life at two hundred observations and a confirmed set of coefficients too wide to confirm any of them. It ended on the second way round. A relation too slow to see in one pair of series may be visible across many pairs that share it — twenty exchange rates against their relative prices, forty regional price indices against a national one, a panel of spreads between related instruments.
That trades a question about time for a question about homogeneity. The trade is worth making when its price is known, and both halves have exact-enough answers.
What pooling buys
The panel test used here is the simplest one that reads every pair: compute each pair’s Engle–Granger statistic, average them, and compare the average with its own 5% point, found by averaging the same number of statistics drawn from pairs of unrelated walks. The average of sixteen null statistics is much less spread than one, so its 5% point sits much closer to the null centre — −2.38 for sixteen pairs against −3.38 for one — and a modest shift in every pair’s statistic, too small to push any single pair past −3.38, moves the average past −2.38.
At two hundred observations a single pair reaches 80% power only if its gap halves within 5.1 steps. Four pairs sharing the half-life reach it up to 9.2 steps, and sixteen up to 15.0. At a half-life of twelve steps — a gap that takes a year of monthly data to halve — one pair is found 20.0% of the time, four 54.2%, and sixteen 99.8%.
So widening is worth much more than supplying the coefficient was. Supplying it moved the 80% half-life from five steps to seven; sixteen pairs move it to fifteen, and the gain grows with the number of pairs, though more slowly than the square root by which the average’s spread shrinks. It is the same trade a sample lengthened makes — the boundary there moved in proportion to the sample’s length — bought with more pairs instead of more periods, which is often the only way to buy it: a hundred years of monthly data is not available for most relations, and sixteen related series usually are.
Three routes round one boundary
The field has now priced three ways of seeing a slow return, and putting them beside each other shows what each is made of.
A longer sample moves the boundary in proportion: the 80% half-life was 2.34 steps at a hundred observations, 5.05 at two hundred and 10.11 at four hundred. It costs time, and for most economic and physical relations the time is not available — the series is as long as it is.
A supplied coefficient moves it by a fixed factor of about 1.4 at every length, because what it removes is the charge for the search, and the charge is a fixed shift of the critical value. It costs the ability to tell the supplied relation from its neighbours.
More pairs move it by a factor that grows with the number of pairs — 1.8 for four, 2.9 for sixteen — because what they add is more returns of the gap to observe, exactly what a longer sample adds, taken sideways. It costs the ability to say which pairs.
Each route’s cost is the question it can no longer answer, and in each case it is a question a study usually wants answered. That is not an accident. The boundary is set by how many times the gap is seen coming back, and every route adds returns by giving up something about where they came from: when, in which relation, in which pair.
The first few pairs are worth the most
The panel’s gain is steep at first and flattens. Going from one pair to four moves the 80% half-life from 5.1 to 9.2 steps; going from four to sixteen, four times as many pairs again, moves it only from 9.2 to 15.0. The average statistic’s spread falls like the square root of the number of pairs, but the power it buys is a function of how far each pair’s statistic is pushed, and a slower return pushes each pair’s statistic by less, so ever more pairs are needed for each extra step of half-life.
A planner who can choose the panel’s size should therefore think of it the way a detectable half-life was read off periods: the first few pairs buy most of the reach, and the returns after a dozen or so are small unless the half-life of interest is already near the edge. Adding pairs also adds the risk the next sections price, which does not diminish: every pair added is one more that might not share the relation.
What the pooled test rejects
The price is in the null. The pooled statistic tests “no pair in the panel is tied”, and its alternative is everything else — one pair tied, half of them, all of them. A rejection establishes that the null is false. It does not establish that the pairs share a relation, and a panel of related series is exactly where a few might and the rest might not.
With none of the sixteen tied the average statistic rejects 5.1% of the time, its level. With one tied it rejects 7.8%; with four, 21.2%; with eight, 55.4%; with twelve, 88.0%; with all sixteen, 99.8%. The rejection rate climbs smoothly with the share of pairs tied, which is the right behaviour for a test of “none” and the wrong one for a test of “all”: a panel with half its pairs tied is rejected more often than not, and a study that reads the rejection as “these series share a long-run relation” has described eight pairs that do not.
Why the pairs cannot say which
The obvious follow-up is to look at the pairs individually and report the ones that are tied. At a half-life of twelve steps the individual tests cannot do that, and the reason is the boundary the whole field has been measuring: each pair alone has about a one-in-six chance of being found.
With all sixteen pairs tied, a pair-by-pair test at 5% finds 2.65 of them on average; with none tied, it finds 0.78, the false positives sixteen tests at 5% produce. The two distributions of the count share 42.4% of their mass. A panel in which exactly two pairs pass turns up in 14.9% of panels with no relation at all and in 23.6% of panels in which every pair has one, and the pairs that pass in a fully tied panel are the ones that happened to look most stationary, not the ones with the strongest relation — every pair in it has the same half-life.
A Bonferroni reading, which asks whether any pair passes at 5% divided by sixteen, does worse still as a detector: it finds something in 22.6% of fully tied panels, against the pooled statistic’s 99.8%. Bonferroni is built for the question “which pairs”, and at a half-life of twelve steps no test of two hundred observations can answer that question about a single pair; the pooled statistic answers a coarser question with far more power. It is the trade what the correction corrects described for any family of tests — control over every member is bought with power over each — and the price of control measured the exchange rate. Here the exchange rate is extreme, because the individual tests had almost no power to spend.
A gentler question sits between the two: not which pairs are tied, but how many. Methods that estimate the share of true nulls from the distribution of a family’s p-values answer it without naming any member. At this half-life each tied pair’s p-value is only slightly pulled towards zero, so such an estimate would be imprecise, but it would at least be an estimate of the right quantity — which neither the pooled rejection nor the count of individual findings is.
The statistic chooses the alternative it can see
The average and the Bonferroni reading are two ways of combining the same sixteen statistics, and the difference in their power is not a matter of one being better. Each is built to see a different kind of departure from “no pair is tied”.
The average is powerful against many small departures: every pair’s statistic pushed a little to the left, none far enough to stand out alone. That is exactly the fully tied panel at a half-life of twelve steps, and it is why the average finds it 99.8% of the time. The Bonferroni reading, which is in effect a test on the most extreme of the sixteen statistics, is powerful against a few large departures: one or two pairs tied fast enough to be unmistakable, the rest unrelated. Against that alternative the average does badly, because fifteen null statistics dilute one extreme one.
So choosing the combining rule is choosing the alternative, and a panel study should choose it knowing which alternative it means. If the theory says the same mechanism operates in every member — one law of one price across many markets — the average is the right statistic and a rejection is evidence of a weak, widespread relation. If the theory says a relation exists in a few members and not the rest, the extreme statistic is the right one and the average will miss it. What neither can do is report, from one rejection, which of the two worlds the panel is in.
Two questions, two instruments
The numbers separate the two things a panel study wants to know as cleanly as the imposed test separated the existence of a relation from its coefficient.
Is there a slow relation anywhere in this family of series? The pooled statistic answers this, with power a single pair cannot approach, and a rejection is solid evidence against “none of them”.
Which of these series are tied? Nothing here answers this at a half-life of twelve. The individual tests are the instruments for it and they are not powerful enough; a panel does not lend its power to its members, because the extra power came from averaging the members together.
Do they all share one relation? This is the question panel studies most often mean, and it is not the null of any test here. It is a claim of homogeneity, and homogeneity has to be argued or tested separately — by a test whose null is “every pair is tied”, which is a different construction with its own critical values, or by an economic argument that the pairs are the same mechanism.
The assumption the whole calculation rests on
Every number above treats the sixteen pairs as independent: each has its own two random walks and its own shocks. Real panels are rarely like that. Twenty exchange rates against the same numeraire share the numeraire’s shocks; forty regional price indices share the national economy’s. When the pairs share shocks, their statistics are correlated, the average of sixteen of them is not as concentrated as the average of sixteen independent ones, and a 5% point computed under independence is too close to the centre. The test then rejects “no pair is tied” too often, and the power curves above overstate what a correlated panel can see. False discoveries that arrive together found the same structure in a family of correlated tests: correlation leaves a procedure’s average error roughly where it was and changes how the errors are distributed across runs, so that a single study is more likely to see none or many. An average of correlated statistics is the extreme case, since every member’s error moves the one number that is tested.
How large that distortion is depends on how much of each pair’s variation is common, and it is the first thing any panel of related series would have to measure before reading the pooled statistic. It is also a place where the method used for a single pair points to a repair: the 5% point for a correlated panel can be simulated from a model of the correlation, exactly as the single pair’s was simulated because no table existed for it. The repair is a simulation with the correlation in it; what is not available is a table.
Why the same boundary keeps appearing
It is worth noticing that every route round the boundary in this field has had the same structure. A longer sample, a supplied coefficient, more pairs: each adds information about the one thing a slow return lacks, which is repeated evidence of the gap coming back. A sample of two hundred steps holds about sixteen half-lives of a gap that halves in twelve, and sixteen half-lives are not many returns to observe. Sixteen such pairs hold about two hundred and sixty half-lives between them, which is enough — provided they are sixteen instances of one mechanism rather than sixteen different mechanisms.
That is the homogeneity question again, arriving as arithmetic. The panel can see what one pair cannot because it borrows returns from its other members, and borrowing is legitimate exactly to the extent that the members are the same. The weight that decides priced the same bargain in a different model: partial pooling lends precision across groups on the assumption that their effects come from one population, and the precision it lends is only as good as that assumption.
What a panel study can defensibly report
Report the pooled rejection as a statement about the family. “At least one of these pairs returns to a long-run relation” is what the pooled statistic supports, and at a half-life of twelve steps it supports it strongly.
Report the individual findings with the rate they would have under no relation. Sixteen tests at 5% find 0.78 pairs on average with nothing there; a study reporting two or three “cointegrated pairs” out of sixteen should put that number beside its own.
Do not read the pooled rejection as a shared relation. With half the pairs tied the pooled test rejects 55.4% of the time. Homogeneity is a separate claim and needs its own evidence.
Measure the correlation between pairs before trusting the pooled critical value. The independence assumption is the one most likely to fail in the panels where widening is most tempting.
What is measured here and what is not
Sixteen independent pairs of two hundred observations reach 80% power up to a half-life of 15.0 steps by their average statistic, against 5.1 for one pair and 9.2 for four; at a half-life of twelve the three find the relation 99.8%, 54.2% and 20.0% of the time.
With some of sixteen pairs tied at a half-life of twelve, the pooled test rejects 7.8% of the time with one tied, 21.2% with four, 55.4% with eight and 99.8% with all; a pair-by-pair test finds 2.65 pairs on average in a fully tied panel and 0.78 in a panel with none.
The rates are counted over five hundred panels a point — two thousand for the null — and every pair is generated by the error-correction model this field has used throughout, with independent shocks across pairs. The pooled 5% point is taken from twenty thousand averages of statistics resampled from twenty thousand pairs of unrelated walks.
Not measured: panels whose pairs share shocks, which is the ordinary case and in which the pooled critical value computed here is too lenient; panels whose pairs have different half-lives, where the average weights fast pairs heavily; and the more elaborate panel statistics in the literature, which standardise each pair’s statistic before averaging and would move the numbers without changing what a rejection means.
Still open: pairs that share a shock
The widened test’s power depends on its pairs being independent, and the panels that tempt a study most — many prices against one numeraire, many regions against one nation — are the ones where they are not. The distortion has a clear shape: a common shock makes the pairs’ statistics move together, the average stays as spread as a single statistic’s in the limit of perfect correlation, and a critical value computed under independence then calls unrelated panels tied far more often than 5%.
How large the distortion is for a stated share of common variation, whether removing the cross-sectional average from every series before testing — the usual repair — restores the level without spending the power widening bought, and at what correlation the panel stops being worth more than its best single pair, are measurements with a clear design and none of them has been made here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The rank is a decision — both name cointegration, critical value, multiple comparisons, statistical power
- A coverage table with its own error — both name bonferroni, multiple comparisons, statistical power
- A horizon chosen after looking — both name critical value, multiple comparisons, statistical power
- An order that spends the error rate — both name bonferroni, multiple comparisons, statistical power
- One control, many arms — both name bonferroni, multiple comparisons, statistical power
- Sixteen subgroups and one effect — both name bonferroni, multiple comparisons, statistical power
Named objects
A flat tag is an object no other essay names yet.
BonferroniCointegrationComposite nullCritical valueEngle–GrangerHalf-lifeMultiple comparisonsStatistical power