Series that move together

A factor estimated from the panel

When sixteen pairs feel a common shock with loadings from a half to one and a half, removing the cross-sectional average leaves the pooled test rejecting an unrelated panel 13.3% of the time. Estimating the factor and each pair's loading instead brings that to 8.9% — and no further, because each loading comes from the same two hundred dates however many pairs there are — and it costs power against the average at the right critical value, 73.8% against 82.5%. Removing the shocks as observed restores 4.5% exactly and finds a shared relation 7.5% of the time: a series tied to its partner carries the partner's factor in its level, not its own shock.

Worth reading first: Three series and a count.

A shock every pair shares took the panel test that sees slow returns across many pairs and gave its pairs a common shock — every exchange rate moved by the numeraire, every regional price by the nation. With nine tenths of each series’ variation common, the average of sixteen pairs’ statistics was worth 1.44 independent pairs and called an unrelated panel tied 31.6% of the time at the critical value computed for independence. Removing every date’s cross-sectional average repaired it completely when every pair felt the shock equally, and left 13.3% false alarms when the pairs’ loadings ran from a half to one and a half.

That essay named the repair that should come next. If the average is wrong because each pair feels the shock with its own strength, then estimate each pair’s strength: find the common factor from the panel itself, regress each series on it, and remove the fitted share. Or, if the factor is observable — the numeraire’s own price — use the observation and skip the estimate. Both are standard. Neither had been counted on these panels, and each turns out to fail in a way the other does not.

Four treatments of the same shock

The panels are the earlier essay’s. Sixteen pairs, two hundred dates, each pair two random walks whose increments are nine tenths common: one common shock moves every first series and another every second series, scaled for pair jj by a loading λj\lambda_j that runs evenly from 0.5 to 1.5 across the pairs. Under the alternative, every pair’s second series is pulled back towards its first with a half-life of twelve dates. The test is the average of the sixteen Engle–Granger statistics, and the question is how often it rejects at 5%.

The four treatments act on the series before the test:

  • nothing removed;
  • the cross-sectional average removed — every series less, at each date, the average of its kind across the panel;
  • an estimated factor removed — the first principal component of the series’ increments taken as the factor’s increments, each series’ loading estimated by regressing its increments on them, and loading times cumulated factor taken off its level;
  • the observed shocks removed — the same regression and removal, using the common shocks themselves, as if the numeraire’s price were on the screen.
Four treatments of a shock sixteen pairs share: false alarms at the independence point, and power at each one's own point. Sixteen pairs, two hundred dates, nine tenths of each series' variation common. every pair loads the shock equally: nothing removed, false alarms 31.6% and power 22.3%; the cross-sectional average removed, false alarms 4.8% and power 99.8%; an estimated factor removed, false alarms 9.2% and power 76.3%; the observed shocks removed, false alarms 4.5% and power 6.5%. loadings run from a half to one and a half: nothing removed, false alarms 30.8% and power 25.8%; the cross-sectional average removed, false alarms 13.3% and power 82.5%; an estimated factor removed, false alarms 8.9% and power 73.8%; the observed shocks removed, false alarms 4.5% and power 7.5%.
Fig. 1 False alarms at the independence 5% point with no pair tied, and power against a half-life of twelve in every pair at each treatment’s own 5% point, for sixteen pairs with nine tenths of their variation common — with equal loadings, and with loadings from a half to one and a half.

Two numbers are read for each. False alarms at the critical value computed for independent pairs, −2.38, say whether a treatment restores the test anyone would actually run. Power at the treatment’s own 5% point — the point its null distribution really has, found by counting — says how much of the panel’s information the treatment kept, with the level set honestly.

The estimated factor halves the residue and stops

With unequal loadings the estimated factor does what the average could not, about halfway. Its false alarms are 8.9% against the average’s 13.3%, and its average statistic is worth 8.27 independent pairs against the average’s 7.49. A loading of 1.5 is removed as 1.5, more or less, rather than as the panel’s mean of 1.

More or less is the problem. Each loading is a regression coefficient estimated from two hundred dates, and the factor it multiplies is a random walk whose variance grows with every date — so an error of a few hundredths in a loading leaves a few hundredths of a wandering common component in every pair, and the sixteen pairs share it. The residue is smaller than the average’s, but it has the same shape: one leftover factor, in every pair, with a pattern of signs.

How often a panel with no relation is called tied, by its number of pairs, under four treatments of its common shock. Nine tenths of the variation common, loadings from a half to one and a half; at 4, 8, 16, 32 pairs: nothing removed, 15.9%, 20.5%, 30.8%, 34.4%; the cross-sectional average removed, 8.0%, 8.9%, 13.3%, 14.7%; an estimated factor removed, 9.5%, 6.1%, 8.9%, 11.2%; the observed shocks removed, 5.4%, 3.6%, 4.5%, 5.3%.
Fig. 2 False alarms at the independence 5% point against the number of pairs, for the four treatments, with nine tenths of the variation common and loadings from a half to one and a half.

Adding pairs does not cure it. The estimated factor’s false alarms run 9.5%, 6.1%, 8.9% and 11.2% at four, eight, sixteen and thirty-two pairs: no trend towards 5%, because more pairs sharpen the factor’s estimate and leave each loading estimated from the same two hundred dates. The average does worse as the panel grows — 8.0%, 8.9%, 13.3%, 14.7% — because its residue is a factor too, and a pooled statistic over more pairs sees a common factor more clearly. With nothing removed the false alarms climb from 15.9% to 34.4%.

So the answer to “how many pairs a panel needs before a factor can be estimated well enough to be worth removing” is that the number of pairs is not what limits it. A principal component is a cross-sectional average with learned weights; it converges as pairs are added. The loadings are time-series regressions; they converge as dates are added, and a panel of two hundred dates has the dates it has.

That is the same limit how slow a return a sample can see found for a single pair, arriving from a different side. A half-life of twelve is hard to see in two hundred dates because two hundred dates hold only about sixteen half-lives; a loading is hard to pin down in two hundred dates because the factor it multiplies has wandered for only two hundred steps. Widening the panel was the response to the first limit. It is no response to the second, because every pair’s loading is its own regression, estimated from its own two hundred dates, however many neighbours it has.

Where the errors go

The earlier essay asked whether an estimated factor’s errors would bias the pooled statistic in a predictable direction. They do not move it. The null median of the estimated-factor panel’s average statistic is −2.028, the average-removed panel’s −2.079, and an independent panel’s −2.044: the centres agree to within a few hundredths. What moves is the spread. The estimated-factor panel’s own 5% point is −2.447, further out than the independence point’s −2.380, because its sixteen statistics are still correlated and their average is still more spread than sixteen independent ones.

That is a useful fact about the repair. Its leftover error is a correlation among the pairs, which a critical value can absorb, not a shift, which a critical value cannot. Read against its own 5% point the estimated-factor test holds its level exactly, and the only question left is how much power it has there.

Power at the right critical value

How often each repair of a shared shock finds a relation every pair shares, read against its own 5% point. A half-life of twelve in every pair, nine tenths common, loadings from a half to one and a half. At 4, 8, 16, 32 pairs: nothing removed, 24.5%, 28.7%, 25.8%, 33.0%; the cross-sectional average removed, 41.0%, 64.5%, 82.5%, 93.8%; an estimated factor removed, 34.8%, 67.8%, 73.8%, 85.0%; the observed shocks removed, 3.5%, 10.0%, 7.5%, 9.5%.
Fig. 3 Power against a half-life of twelve shared by every pair, each treatment read against its own 5% point, by the number of pairs: nine tenths common, loadings from a half to one and a half.

Less than the average’s. At sixteen pairs the average-removed test finds the shared relation 82.5% of the time at its own point and the estimated factor 73.8%; at thirty-two pairs, 93.8% against 85.0%; at four, 41.0% against 34.8%. Only at eight pairs does the estimated factor edge ahead, 67.8% against 64.5%. Nothing removed finds it a quarter to a third of the time.

The estimated factor loses power because it removes too much. The first principal component of the second series’ increments contains a little of every pair’s own error correction, since every pair is being pulled back towards its partner at once, and taking the component out takes a share of the pull with it. The average does the same with equal weights and less flexibility, and loses less. A repair that learns more from the panel also learns, and removes, a little of the relation the panel is being tested for.

With equal loadings, where the average is exactly right, the estimated factor is simply worse: 9.2% false alarms at the independence point against the average’s 4.8%, and 76.3% power against 99.8%. Its average statistic is worth 7.88 independent pairs where the average’s is worth 15.72. Estimating sixteen loadings that are all one is sixteen opportunities to estimate something other than one. And with no common shock at all it still costs: 96.3% power against 99.3%.

Observing the factor removes the relation

The fourth treatment is the one an analyst with the numeraire’s price on the screen would reach for, and on the null it is the best of the four. Removing the observed common shocks leaves 4.5% false alarms with equal loadings and 4.5% with unequal ones, and the average statistic is worth 15.81 of sixteen pairs. Nothing is estimated except each loading against a known factor, and the residue is nearly nothing.

Then it finds a shared relation 7.5% of the time at its own 5% point. At thirty-two pairs, 9.5%. The repair that is perfect under the null has removed almost everything the alternative consists of.

The reason is the error-correction structure the test is looking for. When a pair is tied, its second series is pulled back towards the first, so in the long run it does not wander with its own common shock — that shock is absorbed, date by date, into a gap that closes. It wanders with the first series, and so with the first series’ common shock. The level of a tied second series contains the partner’s factor, not its own. Regressing its increments on its own observed shock recovers a loading, since the shock moves it on the day it arrives, and subtracting loading times the cumulated shock then adds a random walk to its level that the series never had. The pair is untied by its repair.

An unrelated panel has no pull, so its second series really does carry its own shock in its level, and removing it is exactly right. That is why the repair is perfect where nothing is there and destructive where something is: it encodes the null hypothesis in the treatment of the data, and tests whether the data then look like the null. What differencing costs found the same shape in a single pair — a treatment chosen because it fixes the false alarms of unrelated series, applied to series that are related, removes the thing being measured — and the repair that keeps the question was the rule that came out of it.

The obvious amendment is to remove the observed factor only where it belongs under both hypotheses: from the first series of each pair, whose level it drives whether or not the pair is tied, leaving the second series alone. That fails in both directions at once. With no pair tied, the second series still carry their own common shock, untouched, and the panel rejects an unrelated panel 26.8% of the time at the independence point — barely better than removing nothing. With every pair tied, the first series lose the factor and the second series, which carry it in their levels through the pull, keep it; their difference now wanders, and the test finds the relation 5.5% of the time at its own 5% point.

So there is no way to strip an observed factor from a pair that works under both hypotheses. Under the null the second series’ level carries its own shock; under the alternative it carries its partner’s. A treatment of the data has to choose which, and the test exists to find out.

The estimated factor avoids this, partly and by accident: it estimates the factor from the series as they are, error correction included, so its factor for the second series tracks what is common in their levels rather than in their shocks. That is also why it removes a little of the pull. The average sits further along the same line, with weights too crude to remove much of anything that is not genuinely common.

How many pairs each repair is worth

How many independent pairs a panel's average statistic is worth after each repair of its common shock. Nine tenths common, loadings from a half to one and a half. At 4, 8, 16, 32 pairs: nothing removed, 1.43, 1.64, 1.58, 1.75; the cross-sectional average removed, 2.55, 4.78, 7.49, 11.16; an estimated factor removed, 2.78, 5.48, 8.27, 11.57. The dashed line is a panel of independent pairs.
Fig. 4 How many independent pairs the panel’s average statistic is worth — one pair’s null variance over the panel’s — after each repair, by the number of pairs; nine tenths common, loadings from a half to one and a half. The dashed line is a panel of independent pairs.

The count the earlier essay introduced reads the same story from the null side. With nothing removed sixteen pairs are worth 1.58 and thirty-two are worth 1.75: the common shock is the panel. The average gives back 2.55, 4.78, 7.49 and 11.16 of four, eight, sixteen and thirty-two; the estimated factor gives back slightly more, 2.78, 5.48, 8.27 and 11.57. Neither approaches the dashed line, and the gap between them is a few tenths of a pair at every size. The estimated factor’s extra flexibility buys almost nothing on this count, and costs power on the other.

At half common

Everything above is at nine tenths common, where the shock is most of every series. At half common, with the same spread of loadings, the order changes in one place. Removing nothing leaves 14.3% false alarms; the average, 6.5%; the estimated factor, 4.3%; the observed shocks, 4.9%. The estimated factor now restores the level as well as anything, because a factor that is half of each series’ variation leaves a smaller residue for each loading error to multiply.

Power at each one’s own point is 99.0% for the average, 95.5% for the estimated factor and 15.8% for the observed shocks. The average still keeps the most of the relation, and the observed shocks still remove it. The estimated factor’s price in power is smaller here, three and a half points rather than nine, and its gain on the level is real: an analyst who will read the test at the independence point, and cannot simulate its own, gets the right false-alarm rate from the estimated factor at half common and not at nine tenths.

What a panel test with a common shock should do

Remove the average and find its own critical value. With unequal loadings the average leaves 13.3% false alarms at the independence point; read against the panel’s own 5% point, −2.512, it holds its level and finds the shared relation 82.5% of the time — more than any other treatment here. The own point can be found by simulating the panel’s null with its estimated common share and loadings, which the panel supplies; the earlier essay’s effective-pairs count says how far it will sit from the independence point.

Use an estimated factor only when the critical value cannot be simulated. It brings the false alarms at the independence point to between 6% and 11% across panel sizes, about half the average’s excess at sixteen pairs, and it pays for that in power.

Never remove the observed shocks of a series that may be tied. The repair is right only under the null, and a test whose data are treated as if the null were true cannot reject it. An observed factor is still worth having: it measures the common share and each pair’s loading without estimating a factor first, and those are exactly the inputs a critical value for the average-removed test needs.

Report the common share and the loading spread with the test. Every number here moves with both, and a pooled rejection from a panel whose common share is unknown is a statement about a critical value nobody computed. A slow return across many pairs bought its power by assuming the pairs independent; these are the measurements that say how much of that power survives the assumption’s failure.

What was counted

Every rate is over 1,200 simulated panels with no pair tied and 400 with every pair tied at a half-life of twelve, for each panel size and loading pattern, two hundred dates after a burn-in of fifty, the same panels for all four treatments. Critical values are the 5th percentiles of each treatment’s null statistics; the independence point is the 5th percentile of the mean of sixteen independent single-pair statistics. The estimated factor is the leading principal component of the increments’ covariance, found separately for the first and second series of the pairs; loadings are least-squares slopes on it.

Not measured: more than one common factor, a factor whose loadings drift over time, or panels longer than two hundred dates, where the estimated loadings sharpen and the estimated factor should close on the average’s level while keeping more power than the observed shocks. The ordering found here — observed shocks perfect under the null and useless under the alternative, the average best at its own critical value, the estimated factor between them on level and behind on power — is for this panel length and nine tenths common.

Still open: a critical value the panel computes for itself

The best treatment above was the cheapest one, the average, read against a critical value that came from knowing the panel’s common share and loadings. A real panel does not know them; it estimates both, and the critical value simulated from those estimates is a parametric bootstrap of a statistic whose null depends on two nuisance quantities.

Whether a bootstrap critical value built from the panel’s own estimated common share and loadings holds the average-removed test’s level near 5% at sixteen pairs and two hundred dates, how much of the 82.5% power it keeps when its critical value is itself uncertain, and whether the estimated loadings that failed as a repair succeed as inputs to a calibration, are measurable on the same panels and have not been measured here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CointegrationCritical valueDesign effectEffective sample sizeEngle–GrangerThe error-correction modelRandom walkStatistical power