The two worlds that look the same
Worth reading first: One arithmetic, three decisions.
Three causal structures over the same three variables were fitted to one covariance matrix. The largest disagreement between any two of the three matrices, entry by entry, is 4.4·10⁻¹⁶ — the last bit of a double-precision number. The regression of the outcome on the treatment and the covariate returns 0.5000 under all three, to the same tolerance. The effect the treatment actually has is 0.5000 in one of them and 0.8481 in the other two.
That is not a statement about a small sample. Both worlds imply the identical joint distribution of everything anybody can measure, so a sample of a million rows estimates the same matrix to more decimal places and the two worlds still disagree by 0.3481 about the quantity anybody wanted. The spread of the estimate across the three worlds is 4.4·10⁻¹⁶ and the spread of the effect is 0.3481. It is a statement about identification, and it is the one claim in this field that is a theorem rather than a measurement.
What is a measurement is where identification stops being hopeless, and the answer is sharp: it is where an edge is missing. A structure with a hole in it leaves a signature two independence tests can read, and a rule reading that signature finds a real common effect 95.95% of the time at two hundred rows while calling a fork one 0.05% of the time. The whole of what the data can say about causal direction is in that gap, and the price of the signature is that it is an exact zero.
Fitting a causal order to a covariance
The construction is mechanical and it is why the equality is exact rather than close.
Take any ordering of the variables and any complete directed graph consistent with it. Regressing each variable on its predecessors in that order returns coefficients, and the residual variances of those regressions are the noise scales. Both are unique, because each regression’s normal equations have one solution when the predictors are not collinear. So a complete graph on three nodes has exactly six free parameters — three edge weights and three noise scales — a covariance matrix on three nodes has exactly six distinct entries, and the map between them is a bijection.
Every causal order fits every covariance matrix, exactly, and in exactly one way. There is nothing to choose between them on goodness of fit, because all three fit perfectly. There is nothing to choose between them on parsimony, because all three use six parameters.
The three fitted worlds are worth reading, because the reparameterisation is not a relabelling. Starting from a common cause at , , with unit noise everywhere, the same covariance is produced by a mediator world at with noise scales 1.3454 and 0.7433, and by a common-effect world at , , with three noise scales none of which is one. Every number moves. What does not move is the six entries of the matrix.
Two effects, not three, and the coincidence is exact
The three worlds hold 2 distinct effects rather than three, and that is worth a sentence because it looks like an accident of rounding and is not.
The mediator world’s total effect and the common-effect world’s effect are both 0.8481, to fifteen places. Both are the marginal regression coefficient of the outcome on the treatment, — in the mediator world because every route from treatment to outcome is a causal one, so the marginal regression is the total effect; in the common-effect world because the coefficient is the direct edge and the covariate is downstream of everything, so the marginal regression is again clean. Two different reasons, one number, and the number is a function of the shared covariance.
The same identity says something about the first essay of this ladder. The unadjusted bias in the common-cause world is , and the adjusted bias in the other two worlds of that matched triple is exactly minus that. The arithmetic that is right in one world of three reports those as two separate biases with two separate closed forms; on a matched triple they are one quantity with two signs.
The six numbers that cannot tell them apart
A test of conditional independence is the only feature of a joint distribution that a diagram of arrows constrains directly. So the question of what any structure-learning procedure could possibly find reduces to which independences hold, and on these three worlds the answer is none.
The three marginal correlations are 0.6690, 0.7114 and 0.7170. The three partial correlations, each pair with the third variable regressed out, are 0.4472, 0.3244 and 0.4616. All six are the same number in all three worlds to machine precision, which they must be, since each is a function of the shared matrix. None of the six is zero and none is close to zero.
A procedure with no independence to find is not a weak procedure. It has nothing to run on: every test it could perform rejects, in every world, and rejection is uninformative when all the candidates predict it.
Where the data can choose: an edge that is missing
Delete the direct treatment–outcome edge and the three complete graphs become a fork, a chain and a v-structure — the covariate causing both, the covariate between them, and the covariate caused by both. Now the fits stop being free, because a two-edge world has five parameters against a covariance matrix’s six, and the sixth entry has to come out right on its own.
Counting how many of the three structural coefficients each causal order needs to reproduce each world’s covariance is the whole test. A fork’s data is reproduced by a common-cause order on 2 edges and by a mediator order on 2 edges, and needs 3 for a common-effect order. A chain’s data is the same: 2, 2 and 3. A v-structure’s data is reproduced on 2 edges only by a common-effect order and needs 3 from either of the others. Every one of those nine reparameterisations is exact — the largest gap across all nine is 2.22·10⁻¹⁶.
So the fork and the chain describe each other’s data at no extra cost and are not separable at all: they are Markov-equivalent, which is the name for two diagrams implying exactly the same conditional independences, and no test of independence can prefer one. The v-structure is, and the signature is legible in two numbers. A fork implies marginally and 0.0000 conditionally on the covariate; a chain implies and 0.0000; a v-structure implies 0.0000 marginally and conditionally. The two causes of a common effect are independent until their common effect is held fixed, and then they are not — which is the reverse of every other arrangement, and the reason this one shape has a signature while the other two do not.
What the rule finds, and what it costs
The rule that reads that signature is: fail to reject the marginal correlation, reject the conditional one. Run at two hundred rows over two thousand simulated datasets at a level of 0.05 for each of the two tests, it calls a v-structure a common effect 95.95% of the time against a closed form of 94.99%, with a standard error on the count of 0.0044. It calls a fork one 0.05% of the time and a chain one 0.00%.
The closed form is a product of two Fisher-transform power calculations, and reading it apart says where the 95% comes from. Against a v-structure the marginal test rejects 4.05% of the time — it is a true null and the level holds — and the conditional test rejects 100.00% against a power of 99.99%. Against a fork the marginal test rejects 99.95% of the time against a power of 99.99%, so the first half of the rule almost never fails to reject and the false call rate is bounded by how rarely it does.
That is a diagnostic almost too good to be interesting, and the reason it is not is the shape of the first half. The rule requires a failure to reject, and a failure to reject is not evidence that the null holds — the argument the essay about what a p-value does not say makes about effect sizes is exactly the argument here, one level up: the rule cannot tell an exact zero from a small non-zero, and a v-structure with a small real treatment effect is a small non-zero.
The signature is an exact zero, and here is the bill
Give the v-structure a real direct effect from treatment to outcome and walk it up from nothing. The rule keeps calling the world a common effect: 95.8% of the time at an effect of zero, 89.5% at 0.05, 69.8% at 0.1, 43.6% at 0.15, 19.5% at 0.2 and 1.3% at 0.3. The closed forms are 95.0%, 89.1%, 70.9%, 43.9%, 19.5% and 1.1%, and the worst disagreement anywhere in that sweep is 1.46 standard errors.
The effect at which the rule is wrong half the time is 0.1377 at two hundred rows. A direct effect of that size is a marginal correlation of about 0.14 — a real, non-trivial dependence between treatment and outcome, and the rule reports the two as unconnected causes of the covariate half the time it sees it.
Raising the sample size moves it, and moving it is what says what kind of failure this is. At four hundred rows the half-point is 0.1007. The rule has not acquired a new ability; it has resolved a smaller effect, in the ordinary way any test does.
The two halves of the rule scale together, which is why the whole thing behaves like one test rather than like a conjunction. Against the fork’s marginal correlation of 0.3836, the marginal test has power 0.7915 at fifty rows, 0.9784 at a hundred, 0.9999 at two hundred and 1.0000 at four hundred; the conditional test against the v-structure’s partial correlation of the same size has 0.7829, 0.9773, 0.9999 and 1.0000. The conditional test is fractionally the weaker of the two at every sample size, and by exactly the amount conditioning on one variable costs — the Fisher transform’s variance is one over for a marginal correlation and one over for a partial one, which is a single degree of freedom and is the whole difference between the two columns.
A resolution rather than a blind spot
One sample size reads as a threshold and four read as a rate, and the difference is worth the extra three computations.
The check that the counted rates are not confirming their own arithmetic is the same one this ladder’s first essay makes. The closed forms are non-central power calculations on a population correlation; the counts are Fisher tests run on sample correlations computed from simulated rows, and the rows are drawn one variable at a time from the structural equations rather than from any matrix. Nothing on the counting side ever forms a population correlation, and nothing on the closed-form side ever draws a row.
Solving the closed form for the effect at which the rule is wrong half the time gives 0.1392, 0.0985, 0.0695 and 0.0491 at two hundred, four hundred, eight hundred and sixteen hundred rows. Multiplied by the square root of the sample size those are 1.968, 1.970, 1.965 and 1.962 — a spread of 0.008 across an eightfold change in the sample. The effect a collider test mistakes for a zero falls as , and the counted half-point of 0.1377 sits beside the derived 0.1392.
That constant is not a coincidence and it is not deep. It is the critical value of the marginal test, in the units the Fisher transform puts a correlation in, at the point where the test’s power against the true correlation is a half. The rule is wrong about exactly the effects a two-sided test at the 5% level cannot see, which is what a rule built on a failure to reject has to be wrong about.
So the honest description is that structure learning from independences has a resolution, not a blind spot. A blind spot would be a class of worlds it never recovers however much data arrives; a resolution shrinks. Both readings tell somebody not to trust a v-structure found at two hundred rows against an effect of a tenth, and only one of them tells them what a larger sample buys.
Where this does not hold
Two limits, and the first is the one that would change the field’s headline claim rather than qualify it.
The equivalence is an equivalence of second moments. A covariance matrix is the whole of a Gaussian joint law, so two structures that share one share everything and no test of any kind separates them. A non-Gaussian law carries information a covariance does not, and there are model families in which the direction of an arrow is recoverable from observational data alone — a linear structure with non-Gaussian noise is the standard example, and the recovery is real rather than a technicality. Nothing here measures how much data such a recovery takes or how it degrades as the noise approaches normality, and it should: the claim proved above is that Gaussian worlds are inseparable, and somebody who has read only the headline will over-apply it.
The sparse worlds have exactly one edge missing, in one place. The two-against-three edge counting above is a three-node result, and a real search runs over many variables, many triples and many tests at once. What a full constraint-based search costs on these worlds — how many tests, at what multiplicity, and with what error rate over a whole graph rather than one triple — is a question about error rates across a family, which is not measured here.
The third limit is the one this essay shares with everything around it: the rule assumes the covariate was recorded for every row, so nothing here says what a selection into the sample would do to the very independences the rule reads. That is being in the data as a condition, and it is the mechanism by which a v-structure’s signature can appear where no v-structure is.
What identification means, in this exact case
The word is used loosely enough that the arithmetic above is worth restating as a definition.
A quantity is identified if two models that agree on the distribution of the observables must agree on it. The effect here is not, and the counterexample is the three-row table this essay opened on: two models agreeing to 4.4·10⁻¹⁶ on the distribution and disagreeing by 0.3481 on the effect. Nothing about that is statistical. There is no estimator to improve, no standard error to shrink, and no sample size at which the disagreement narrows — which is precisely what the reversal that no amount of data settles said in words and this says in a matrix.
The practical residue is that the assumption has to be argued for outside the data, and that the argument is about which edges are absent. A complete graph is unfalsifiable and uninformative; every claim any procedure here makes is made by a hole. That is the same shape as the essay that refuses “explained” as a causal word — a fit statistic rises with any structure and so distinguishes none — and it is why what randomisation buys is worth its cost: randomising the treatment deletes every arrow into it by construction, which is a hole put there on purpose rather than one argued for afterwards.
What this essay does not settle is what to do when the structure is not known, which is the ordinary case. Two crude rules are available — adjust for everything measured, or adjust for nothing — and both are answers to a question about structure given by somebody who does not have one. The sweep that priced them finds that the first is worse than the second on 65.5% of four thousand random structures, and that neither is the best rule on more than a fifth of them. Two mechanisms with one observable relation is not a problem confined to causal diagrams either: two series with identical long-run behaviour is the same difficulty with time doing the work the covariate does here, and four datasets with one set of summaries is the version where the shared object is a summary rather than a matrix.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two instruments that disagree — both name null hypothesis, observational equivalence, statistical power
- A boundary for giving up — both name sample size, statistical power
- A coverage table with its own error — both name sample size, statistical power
- A distribution drawn from the null — both name null hypothesis, statistical power
- A family before a fit — both name covariance matrix, identification
- An order that spends the error rate — both name null hypothesis, statistical power
Named objects
A flat tag is an object no other essay names yet.
Causal diagramCausal effectConditional independenceCovariance matrixFisher transformIdentificationMarkov equivalenceNull hypothesisObservational equivalencePartial correlationSample sizeStatistical powerStructural modelV structure