A shock every pair shares
Worth reading first: Three series and a count.
A slow return across many pairs found that a relation too slow to see in one pair of series can be seen across sixteen. Averaging sixteen Engle–Granger statistics and reading the average against its own 5% point found a gap that halves in twelve steps 99.8% of the time, where one pair of two hundred observations found it 20.0%. Every number there assumed that the sixteen pairs had nothing in common but the relation being tested, and the essay ended by naming the panels that tempt a study most as the ones where that fails: many exchange rates against one numeraire, many regional prices against one nation, many spreads on one market.
In those panels a single shock moves every pair at once. How much that matters has a precise answer, and it is expressed most simply as a count: the number of independent pairs the panel is still worth.
A panel in which one shock reaches every pair
The panel is built the way the independent one was, with one addition. Each pair’s first series is a random walk and its second is pulled back towards it at a rate set by the half-life, or not at all under the null. But the walks’ increments are now partly shared:
where and are the same for every pair on every date and , are each pair’s own. With every loading equal to one, is the share of each series’ variation that is common to the panel. It is zero in the independent panel and one in a panel of sixteen copies of one pair. An exchange rate against the dollar carries the dollar’s own shocks in exactly this way, and a regional price carries the national one.
Nothing about any single pair changes. Each pair is still two unrelated walks under the null — the regression the cliff that is a slope found called significant 83.4% of the time between two independent walks — or a walk and its error-correcting partner under the alternative, and each pair’s statistic still has the distribution the test with no table simulated, with its 5% point at −3.38. What changes is the joint behaviour of the sixteen.
The average spreads back out
Averaging sixteen statistics was worth so much because the average of sixteen independent draws is four times less spread than one, and its 5% point sits correspondingly close to the centre: −2.38 against a single pair’s −3.38. Correlated draws do not average down that far.
With a quarter of the variation common, the panel’s own 5% point is −2.43; with half, −2.58; with three quarters, −2.86; with nine tenths, −3.18, most of the way back to a single pair’s. A test that keeps using the independence point at −2.38 is reading a statistic whose spread has grown against a line drawn for one that had not.
The consequence is the curve at the top. With no common shock the pooled test at the independence point rejects an unrelated panel 4.6% of the time. With a quarter common it rejects 7.2%; with half, 15.5%; with three quarters, 27.3%; with nine tenths, 31.6%. A panel study of sixteen exchange rates whose shocks are half common, analysed as though its pairs were independent, finds a long-run relation that is not there about one time in six.
Four pairs suffer less, rejecting 6.8% at half common and 16.9% at nine tenths. That is not because four pairs are more robust; it is because the independence point for four pairs, −2.71, was never as close to the centre, so there was less concentration for the common shock to undo.
How many pairs the panel is worth
The spread of the average has a direct reading. One pair’s null statistic has some variance; sixteen independent pairs’ average has a sixteenth of it; sixteen identical pairs’ average has all of it. The ratio of the single variance to the average’s is the number of independent pairs the panel is worth — the design effect of a cluster sample, read off the statistic the panel actually tests.
Sixteen independent pairs are worth 15.96, sixteen to sampling error. A quarter common, 12.25. Half common, 5.37. Three quarters, 2.61. Nine tenths, 1.44. The count falls steeply, because the variance of an average of correlated terms is dominated by the correlation as soon as there are more than a few of them: at sixteen pairs, every one of the hundred and twenty pairs of pairs contributes a covariance, and the sixteen variances are outnumbered.
That is the same arithmetic the observations that repeat each other found along a series, where a fifty-point series with a lag-one correlation of 0.8 was worth about six independent observations. There the dependence ran through time; here it runs across the panel. In both, a count that claims independence overstates what the data can say, and the overstatement grows with the count.
A small correlation between statistics is enough
The count can be turned round to say how correlated two pairs’ statistics are. If every two of the sixteen statistics have correlation , the average’s variance is the single variance times , so a panel worth pairs has
At a quarter common that is 0.020; at half common, 0.132; at three quarters, 0.343; at nine tenths, 0.672. The statistics are much less correlated than the series. Two series sharing half their variation have increments correlated at one half, and the Engle–Granger statistics computed from two such pairs are correlated at about an eighth — because each statistic is a nonlinear reading of a whole path, fitted and then tested, and most of what it reads is the pair’s own detail.
And an eighth is enough to cut sixteen pairs to five. The variance of an average of sixteen terms has sixteen variances and two hundred and forty covariances in it, so a correlation that looks negligible between any two members dominates the sum as soon as the members are many. The same formula, , is the design effect of a cluster sample, and the count that is not the rows found three hundred rows in five clusters of sixty carrying 6.9 times the variance an independent-rows calculation reports. A panel of pairs is a single cluster of sixteen, and the correlation inside it is set by the common shock.
The formula also puts a ceiling on widening. If the statistics’ correlation stays at 0.132 as pairs are added, a panel of pairs is worth , which approaches as grows: no panel of any size sharing half its variation is worth more than about eight independent pairs. The independent calculation said the first few pairs buy most of the reach; with a common shock, the later ones buy almost none of it, and a planner choosing between sixteen pairs and sixty is choosing between 5.4 and 6.8.
The practical consequence is that the correlation worth checking is not the one a reader would look at first. The pairwise correlation of the statistics themselves is small and hard to estimate from one panel; the correlation of the series’ increments is large and easy to estimate, and it is the one that predicts where on these curves the panel sits.
What is left at the right critical value
The false alarms can be removed by simulating the 5% point with the common shock in it, exactly as the single pair’s point was simulated because no table existed for it. That restores the level. It does not restore what the panel was supposed to buy.
Read against its own point, the sixteen-pair average finds a relation shared by every pair 99.3% of the time with no common shock, 93.0% with a quarter, 71.8% with half, 43.5% with three quarters and 22.3% with nine tenths — where one pair alone finds it 20.0% of the time. At nine tenths common the panel of sixteen has become one pair, both in the variance count and in what it can see. At half common it is worth five pairs and sees 71.8%, between the 54.2% of four independent pairs and the 99.3% of sixteen with no common shock.
Read against the independence point instead, the same panels appear to keep most of their power — 79.8% at three quarters common and 73.5% at nine tenths. That is the false alarms counted as findings. A panel whose null is rejected 31.6% of the time will reject its alternative often too, and the gap between 73.5% and 22.3% is exactly what the wrong critical value adds.
So there is a clean answer to the question of when a correlated panel stops being worth more than its best single pair: when most of its variation is common. At nine tenths it is worth 1.44 pairs and sees 22.3% where one sees 20.0%, and widening has bought almost nothing.
Removing the average
The usual repair in panel work is to remove, from every series on every date, the average of that kind of series across the panel — every exchange rate less the average exchange rate, every relative price less the average relative price — and to test the pairs of what is left. When every pair feels the common shock equally the average carries all of it, and removing the average removes it.
It does, completely. With the averages removed the sixteen-pair panel’s false alarms at the independence point stay between 4.8% and 5.8% at every common share, its variance count stays between 14.97 and 15.99 pairs, and its power at its own point stays between 98.0% and 100.0%. The common shock was noise to every pair’s gap as well as to its walks, so taking it out leaves each pair’s relation exactly as strong as it was and each pair’s statistic as independent of the others as it was in the independent panel.
With four pairs the repair has a price. The average of four series contains a quarter of each, so every demeaned series carries part of every other, and the demeaning itself ties the four together. With no common shock at all, the demeaned four-pair test rejects an unrelated panel 7.0% of the time at the independence point and finds a shared relation 46.8% of the time at its own point, against 53.8% for the raw series; its variance count is 3.64 pairs of four. Once the common share passes about half, the repair pays for itself — 56.0% against 41.3% at three quarters — but in a small panel with little common variation it removes information along with the shock.
When the pairs do not feel it equally
The repair works exactly because the average is the common shock. If the pairs feel the shock with different strengths — a currency that tracks the dollar closely beside one that barely moves with it — then the average carries the average loading, and every pair is left with its own departure from that average multiplied by the shock.
With loadings running evenly from 0.5 to 1.5, removing the averages leaves 6.1% false alarms at a quarter common, 6.5% at half, 8.0% at three quarters and 13.3% at nine tenths — better than the 30.8% with no repair at all, and still well above the level. And it takes power with it: at nine tenths common the demeaned test’s power at its own 5% point is 82.5% where it was 99.8% with equal loadings.
The residual is predictable from the construction. A pair with loading 1.5 keeps half a unit of the common shock after the average is removed, a pair with loading 0.5 keeps minus half a unit, and those residuals are still shared, with opposite signs, across the pairs on either side of the average. Demeaning turns one common factor into a smaller one with a sign pattern, and a pooled statistic sees the smaller factor exactly as it saw the larger.
What to measure before pooling
The share of common variation is easy to estimate before any test is run. The correlation between two series’ increments is when loadings are equal, so the average pairwise correlation of the panel’s differenced series says roughly where on the curves above the panel sits. A panel of rates quoted against one numeraire carries the numeraire’s movements in every member, which is where that correlation is likely to be large; at a half, the independence point rejects three times too often.
Whether loadings are equal is also checkable: regress each series’ increments on the average increment and look at the slopes. Slopes all near one mean demeaning will remove the common shock; slopes ranging from a half to one and a half mean it will leave the pattern measured above, and a repair that estimates each pair’s loading — a common-factor model rather than a common average — is needed.
This is the panel version of a lesson that recurs wherever tests are combined. False discoveries that arrive together found that correlation among twenty tests leaves the average error rate of a false-discovery procedure roughly where it was and makes the errors arrive in clumps, so that one study is more likely to see none or many. An average of correlated statistics is the extreme of that clumping, because the whole family is tested through a single number, and every member’s shared error moves it.
One combination of the sixteen is immune. A Bonferroni reading — any pair past its single-pair point at 5% divided by sixteen — holds its level under any dependence whatever, because the chance that at least one of sixteen events happens can never exceed the sum of their chances, which is the whole of what the correction corrects. The independent panel found it nearly powerless against a slow shared relation, 22.6% where the average found 99.8%. Under a strong common shock that comparison narrows from the average’s side, since the average at its correct point has fallen to about one pair’s power. The Bonferroni reading’s own power also falls as the pairs move together — sixteen correlated chances to pass are fewer than sixteen independent ones — and how far it falls was not measured here; what is certain is that its level does not move.
Where the pooled test now stands
A panel of related series should report its common share before its pooled statistic. At half common, a sixteen-pair panel read against the independence point calls an unrelated panel tied 15.5% of the time; the number is not a detail of the method but the difference between a 5% test and a 15% one.
The right critical value restores the level and does not restore the panel. At its own point a sixteen-pair panel sharing half its variation sees a twelve-step half-life 71.8% of the time, as five independent pairs would; sharing nine tenths, it sees about what one pair sees.
Removing the cross-sectional average is a complete repair when the loadings are equal and a partial one when they are not. With loadings from a half to one and a half, the demeaned panel still rejects an unrelated panel 13.3% of the time at nine tenths common.
And in a small panel, demeaning costs something even when there is nothing to remove. Four pairs demeaned are worth 3.64 at no common share, and find a shared relation 46.8% of the time where the raw four find it 53.8%.
A relation supplied in advance and how slow a return a sample can see priced the other two routes round the same boundary — a coefficient supplied rather than estimated, and a longer sample. The panel route has now been priced twice, once for independent pairs and once for pairs that share their shocks, and the second price is the one a real panel pays.
How the numbers were made
Counted: the null distributions over 1,200 panels at each common share and the power over 400, every pair of two hundred observations generated by the same error-correction model as the independent panel, with the common shocks added as above. The independence 5% point is the one the independent panel used, taken from twenty thousand averages of statistics resampled from twenty thousand unrelated pairs. The variance counts divide the variance of those twenty thousand single-pair statistics by the variance of the 1,200 panel averages.
Exact, given the construction: with equal loadings, removing each date’s cross-sectional average removes the common shocks from every series exactly, since each series carries the same multiple of them.
Not claimed: that one common factor with a common loading pattern describes real panels. Exchange rates may share several factors, and loadings may change over time; each added factor is another shared error the pooled statistic reads through. Not claimed either that the Engle–Granger average is the best pooled statistic under dependence — the panel literature’s factor-adjusted tests exist for this case — only that the plain average and its independence critical value, which is what a first analysis usually runs, behave as described.
Still open: a factor the average cannot see
Demeaning removes a shock that every pair feels in the same proportion, and the loadings figure shows what it leaves when they do not. The natural next repair estimates each series’ loading on the common factor — from the panel’s own principal component, or from a regression on an observed factor like the numeraire’s price — and removes that. With sixteen pairs and two hundred dates the factor and its loadings can be estimated, but not exactly, and the estimated factor carries a little of every series’ own movement, including the gaps being tested.
How much of the level and power an estimated factor gives back, whether its errors bias the pooled statistic in a predictable direction, and how many pairs a panel needs before a factor can be estimated well enough to be worth removing, have not been measured here. They decide whether a panel with unequal loadings — which is most panels — can be repaired at all by anything short of a model of every pair’s exposure.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The rank is a decision — both name cointegration, critical value, statistical power
- A detector built for the ordering — both name critical value, statistical power
- A horizon chosen after looking — both name critical value, statistical power
- A probe nobody chose — both name effective sample size, statistical power
- A taper and a critical value — both name critical value, statistical power
- Draws that repeat each other — both name critical value, effective sample size
Named objects
A flat tag is an object no other essay names yet.
CointegrationComposite nullCritical valueDesign effectEffective sample sizeEngle–GrangerHalf-lifeStatistical power