What conditioning on a variable does

The two worlds that look the same

Three causal structures were fitted to one covariance matrix and agree with it to 4.4·10⁻¹⁶. The regression returns 0.5000 under all three; the effect they hold is 0.5000, 0.8481 and 0.8481. What separates structures is a missing edge, and the signature of one is a correlation of exactly zero.

Worth reading first: One arithmetic, three decisions.

Three causal structures over the same three variables were fitted to one covariance matrix. The largest disagreement between any two of the three matrices, entry by entry, is 4.4·10⁻¹⁶ — the last bit of a double-precision number. The regression of the outcome on the treatment and the covariate returns 0.5000 under all three, to the same tolerance. The effect the treatment actually has is 0.5000 in one of them and 0.8481 in the other two.

That is not a statement about a small sample. Both worlds imply the identical joint distribution of everything anybody can measure, so a sample of a million rows estimates the same matrix to more decimal places and the two worlds still disagree by 0.3481 about the quantity anybody wanted. The spread of the estimate across the three worlds is 4.4·10⁻¹⁶ and the spread of the effect is 0.3481. It is a statement about identification, and it is the one claim in this field that is a theorem rather than a measurement.

What is a measurement is where identification stops being hopeless, and the answer is sharp: it is where an edge is missing. A structure with a hole in it leaves a signature two independence tests can read, and a rule reading that signature finds a real common effect 95.95% of the time at two hundred rows while calling a fork one 0.05% of the time. The whole of what the data can say about causal direction is in that gap, and the price of the signature is that it is an exact zero.

Fitting a causal order to a covariance

The construction is mechanical and it is why the equality is exact rather than close.

Take any ordering of the variables and any complete directed graph consistent with it. Regressing each variable on its predecessors in that order returns coefficients, and the residual variances of those regressions are the noise scales. Both are unique, because each regression’s normal equations have one solution when the predictors are not collinear. So a complete graph on three nodes has exactly six free parameters — three edge weights and three noise scales — a covariance matrix on three nodes has exactly six distinct entries, and the map between them is a bijection.

Every causal order fits every covariance matrix, exactly, and in exactly one way. There is nothing to choose between them on goodness of fit, because all three fit perfectly. There is nothing to choose between them on parsimony, because all three use six parameters.

The three fitted worlds are worth reading, because the reparameterisation is not a relabelling. Starting from a common cause at a=0.9000a = 0.9000, b=0.5000b = 0.5000, d=0.7000d = 0.7000 with unit noise everywhere, the same covariance is produced by a mediator world at a=0.4972a = 0.4972 with noise scales 1.3454 and 0.7433, and by a common-effect world at a=0.2391a = 0.2391, b=0.8481b = 0.8481, d=0.3043d = 0.3043 with three noise scales none of which is one. Every number moves. What does not move is the six entries of the matrix.

One distribution, two effects. Three causal structures fitted to one covariance matrix over a treatment, a covariate and an outcome. Each reproduces it exactly — the largest entry-wise disagreement across all three is 4.4e-16 — so no sample of any size distinguishes them. The regression of the outcome on the treatment and the covariate returns 0.500 in all three, to within 4.4e-16, because that coefficient is a function of the covariance and of nothing else. The effect the three worlds hold is 0.500, 0.848 and 0.848: adjusting is exactly right in the first and off by −0.348 in the other two. The arithmetic cannot see the difference and the difference is the whole question.
Fig. 1 The three parameterisations of one covariance matrix, the effect each of them holds and what the adjusted regression returns under each. The estimates agree to 4.4·10⁻¹⁶ and the effects differ by 0.3481.

Two effects, not three, and the coincidence is exact

The three worlds hold 2 distinct effects rather than three, and that is worth a sentence because it looks like an accident of rounding and is not.

The mediator world’s total effect and the common-effect world’s effect are both 0.8481, to fifteen places. Both are the marginal regression coefficient of the outcome on the treatment, Σty/Σtt\Sigma_{ty}/\Sigma_{tt} — in the mediator world because every route from treatment to outcome is a causal one, so the marginal regression is the total effect; in the common-effect world because the coefficient bb is the direct edge and the covariate is downstream of everything, so the marginal regression is again clean. Two different reasons, one number, and the number is a function of the shared covariance.

The same identity says something about the first essay of this ladder. The unadjusted bias in the common-cause world is 0.84810.5000=0.34810.8481 - 0.5000 = 0.3481, and the adjusted bias in the other two worlds of that matched triple is exactly minus that. The arithmetic that is right in one world of three reports those as two separate biases with two separate closed forms; on a matched triple they are one quantity with two signs.

The six numbers that cannot tell them apart

A test of conditional independence is the only feature of a joint distribution that a diagram of arrows constrains directly. So the question of what any structure-learning procedure could possibly find reduces to which independences hold, and on these three worlds the answer is none.

The three marginal correlations are 0.6690, 0.7114 and 0.7170. The three partial correlations, each pair with the third variable regressed out, are 0.4472, 0.3244 and 0.4616. All six are the same number in all three worlds to machine precision, which they must be, since each is a function of the shared matrix. None of the six is zero and none is close to zero.

A procedure with no independence to find is not a weak procedure. It has nothing to run on: every test it could perform rejects, in every world, and rejection is uninformative when all the candidates predict it.

Six correlations, three worlds, no difference. Every correlation and every partial correlation among the treatment, the covariate and the outcome, in three worlds parameterised to share one joint distribution. The six readings are 0.6690, 0.7114, 0.7170, 0.4472, 0.3244, 0.4616, and the largest disagreement between the three worlds on any of them is 5.0e-16 — machine noise. None is zero, so there is no conditional independence to test: all three are complete graphs on three nodes, and a complete graph has no missing edge for the data to notice. The effects those three worlds hold are 0.500, 0.848, 0.848.
Fig. 2 The three marginal and three partial correlations, in the three worlds sharing one covariance matrix. Every bar has the same height in all three and none of the six sits at zero.

Where the data can choose: an edge that is missing

Delete the direct treatment–outcome edge and the three complete graphs become a fork, a chain and a v-structure — the covariate causing both, the covariate between them, and the covariate caused by both. Now the fits stop being free, because a two-edge world has five parameters against a covariance matrix’s six, and the sixth entry has to come out right on its own.

Counting how many of the three structural coefficients each causal order needs to reproduce each world’s covariance is the whole test. A fork’s data is reproduced by a common-cause order on 2 edges and by a mediator order on 2 edges, and needs 3 for a common-effect order. A chain’s data is the same: 2, 2 and 3. A v-structure’s data is reproduced on 2 edges only by a common-effect order and needs 3 from either of the others. Every one of those nine reparameterisations is exact — the largest gap across all nine is 2.22·10⁻¹⁶.

So the fork and the chain describe each other’s data at no extra cost and are not separable at all: they are Markov-equivalent, which is the name for two diagrams implying exactly the same conditional independences, and no test of independence can prefer one. The v-structure is, and the signature is legible in two numbers. A fork implies r(t,y)=0.3836r(t, y) = 0.3836 marginally and 0.0000 conditionally on the covariate; a chain implies 0.45860.4586 and 0.0000; a v-structure implies 0.0000 marginally and 0.3836-0.3836 conditionally. The two causes of a common effect are independent until their common effect is held fixed, and then they are not — which is the reverse of every other arrangement, and the reason this one shape has a signature while the other two do not.

Two edges is what an independence can see. The three ways to join a treatment, a covariate and an outcome with two edges, and what each of them implies. A fork and a chain both put the treatment and the outcome at a correlation of 0.3836 and 0.4586 and both make them exactly independent given the covariate. A common effect does the reverse: a marginal correlation of 0.0e+0 and a partial correlation of -0.3836. Asked to reproduce a fork's distribution, a common effect needs 3 edges rather than 2; asked to reproduce a common effect's, a fork and a chain need 3. So one of the three shapes carries a signature and the other two remain a matched pair — which is the most a conditional independence can do, and it is not enough to decide whether to adjust.
Fig. 3 How many edges each causal order needs to reproduce each sparse world’s covariance. Two orders describe a fork or a chain on two edges apiece; only one describes a v-structure on two.

What the rule finds, and what it costs

The rule that reads that signature is: fail to reject the marginal correlation, reject the conditional one. Run at two hundred rows over two thousand simulated datasets at a level of 0.05 for each of the two tests, it calls a v-structure a common effect 95.95% of the time against a closed form of 94.99%, with a standard error on the count of 0.0044. It calls a fork one 0.05% of the time and a chain one 0.00%.

The closed form is a product of two Fisher-transform power calculations, and reading it apart says where the 95% comes from. Against a v-structure the marginal test rejects 4.05% of the time — it is a true null and the level holds — and the conditional test rejects 100.00% against a power of 99.99%. Against a fork the marginal test rejects 99.95% of the time against a power of 99.99%, so the first half of the rule almost never fails to reject and the false call rate is bounded by how rarely it does.

That is a diagnostic almost too good to be interesting, and the reason it is not is the shape of the first half. The rule requires a failure to reject, and a failure to reject is not evidence that the null holds — the argument the essay about what a p-value does not say makes about effect sizes is exactly the argument here, one level up: the rule cannot tell an exact zero from a small non-zero, and a v-structure with a small real treatment effect is a small non-zero.

The signature needs a zero that is exactly zero. How often the rule that finds a common effect — fail to reject the treatment–outcome correlation, reject it given the covariate — calls a world a common effect, as that world is given a real treatment effect, at 200 rows and 1200 draws a point. At a true effect of nothing the rule is right 95.8% of the time. At 0.1 it still says common effect 69.8% of the time, and it only falls below half at 0.138. The line is the asymptotic rate from Fisher's transform and the points are counted; the worst disagreement between them is 1.46 standard errors. What the rule reads is a correlation of exactly zero, which is a property of a world rather than a thing a sample reports, and failing to reject at 200 rows is not evidence that it holds.
Fig. 4 The rate at which the rule calls each of the three sparse worlds a common effect, at two hundred rows, count against closed form. It is 95.95% on the v-structure and 0.05% and 0.00% on the other two.

The signature is an exact zero, and here is the bill

Give the v-structure a real direct effect from treatment to outcome and walk it up from nothing. The rule keeps calling the world a common effect: 95.8% of the time at an effect of zero, 89.5% at 0.05, 69.8% at 0.1, 43.6% at 0.15, 19.5% at 0.2 and 1.3% at 0.3. The closed forms are 95.0%, 89.1%, 70.9%, 43.9%, 19.5% and 1.1%, and the worst disagreement anywhere in that sweep is 1.46 standard errors.

The effect at which the rule is wrong half the time is 0.1377 at two hundred rows. A direct effect of that size is a marginal correlation of about 0.14 — a real, non-trivial dependence between treatment and outcome, and the rule reports the two as unconnected causes of the covariate half the time it sees it.

Raising the sample size moves it, and moving it is what says what kind of failure this is. At four hundred rows the half-point is 0.1007. The rule has not acquired a new ability; it has resolved a smaller effect, in the ordinary way any test does.

The two halves of the rule scale together, which is why the whole thing behaves like one test rather than like a conjunction. Against the fork’s marginal correlation of 0.3836, the marginal test has power 0.7915 at fifty rows, 0.9784 at a hundred, 0.9999 at two hundred and 1.0000 at four hundred; the conditional test against the v-structure’s partial correlation of the same size has 0.7829, 0.9773, 0.9999 and 1.0000. The conditional test is fractionally the weaker of the two at every sample size, and by exactly the amount conditioning on one variable costs — the Fisher transform’s variance is one over n3n - 3 for a marginal correlation and one over n4n - 4 for a partial one, which is a single degree of freedom and is the whole difference between the two columns.

The signature needs a zero that is exactly zero. How often the rule that finds a common effect — fail to reject the treatment–outcome correlation, reject it given the covariate — calls a world a common effect, as that world is given a real treatment effect, at 400 rows and 1200 draws a point. At a true effect of nothing the rule is right 94.0% of the time. At 0.1 it still says common effect 50.5% of the time, and it only falls below half at 0.101. The line is the asymptotic rate from Fisher's transform and the points are counted; the worst disagreement between them is 1.46 standard errors. What the rule reads is a correlation of exactly zero, which is a property of a world rather than a thing a sample reports, and failing to reject at 400 rows is not evidence that it holds.
Fig. 5 The same three rates at four hundred rows rather than two hundred. The marginal test’s power against a fork rises and the level against a v-structure does not move, because a level does not depend on the sample size.

A resolution rather than a blind spot

One sample size reads as a threshold and four read as a rate, and the difference is worth the extra three computations.

The check that the counted rates are not confirming their own arithmetic is the same one this ladder’s first essay makes. The closed forms are non-central power calculations on a population correlation; the counts are Fisher tests run on sample correlations computed from simulated rows, and the rows are drawn one variable at a time from the structural equations rather than from any matrix. Nothing on the counting side ever forms a population correlation, and nothing on the closed-form side ever draws a row.

Solving the closed form for the effect at which the rule is wrong half the time gives 0.1392, 0.0985, 0.0695 and 0.0491 at two hundred, four hundred, eight hundred and sixteen hundred rows. Multiplied by the square root of the sample size those are 1.968, 1.970, 1.965 and 1.962 — a spread of 0.008 across an eightfold change in the sample. The effect a collider test mistakes for a zero falls as 1.97/n1.97/\sqrt{n}, and the counted half-point of 0.1377 sits beside the derived 0.1392.

That constant is not a coincidence and it is not deep. It is the critical value of the marginal test, in the units the Fisher transform puts a correlation in, at the point where the test’s power against the true correlation is a half. The rule is wrong about exactly the effects a two-sided test at the 5% level cannot see, which is what a rule built on a failure to reject has to be wrong about.

So the honest description is that structure learning from independences has a resolution, not a blind spot. A blind spot would be a class of worlds it never recovers however much data arrives; a resolution shrinks. Both readings tell somebody not to trust a v-structure found at two hundred rows against an effect of a tenth, and only one of them tells them what a larger sample buys.

Where this does not hold

Two limits, and the first is the one that would change the field’s headline claim rather than qualify it.

The equivalence is an equivalence of second moments. A covariance matrix is the whole of a Gaussian joint law, so two structures that share one share everything and no test of any kind separates them. A non-Gaussian law carries information a covariance does not, and there are model families in which the direction of an arrow is recoverable from observational data alone — a linear structure with non-Gaussian noise is the standard example, and the recovery is real rather than a technicality. Nothing here measures how much data such a recovery takes or how it degrades as the noise approaches normality, and it should: the claim proved above is that Gaussian worlds are inseparable, and somebody who has read only the headline will over-apply it.

The sparse worlds have exactly one edge missing, in one place. The two-against-three edge counting above is a three-node result, and a real search runs over many variables, many triples and many tests at once. What a full constraint-based search costs on these worlds — how many tests, at what multiplicity, and with what error rate over a whole graph rather than one triple — is a question about error rates across a family, which is not measured here.

The third limit is the one this essay shares with everything around it: the rule assumes the covariate was recorded for every row, so nothing here says what a selection into the sample would do to the very independences the rule reads. That is being in the data as a condition, and it is the mechanism by which a v-structure’s signature can appear where no v-structure is.

What identification means, in this exact case

The word is used loosely enough that the arithmetic above is worth restating as a definition.

A quantity is identified if two models that agree on the distribution of the observables must agree on it. The effect here is not, and the counterexample is the three-row table this essay opened on: two models agreeing to 4.4·10⁻¹⁶ on the distribution and disagreeing by 0.3481 on the effect. Nothing about that is statistical. There is no estimator to improve, no standard error to shrink, and no sample size at which the disagreement narrows — which is precisely what the reversal that no amount of data settles said in words and this says in a matrix.

The practical residue is that the assumption has to be argued for outside the data, and that the argument is about which edges are absent. A complete graph is unfalsifiable and uninformative; every claim any procedure here makes is made by a hole. That is the same shape as the essay that refuses “explained” as a causal word — a fit statistic rises with any structure and so distinguishes none — and it is why what randomisation buys is worth its cost: randomising the treatment deletes every arrow into it by construction, which is a hole put there on purpose rather than one argued for afterwards.

What this essay does not settle is what to do when the structure is not known, which is the ordinary case. Two crude rules are available — adjust for everything measured, or adjust for nothing — and both are answers to a question about structure given by somebody who does not have one. The sweep that priced them finds that the first is worse than the second on 65.5% of four thousand random structures, and that neither is the best rule on more than a fifth of them. Two mechanisms with one observable relation is not a problem confined to causal diagrams either: two series with identical long-run behaviour is the same difficulty with time doing the work the covariate does here, and four datasets with one set of summaries is the version where the shared object is a summary rather than a matrix.

The same covariate, three ways round. Three worlds over a treatment, an outcome and a covariate, joined by the same three edges at the same three strengths — 0.90, 0.50 and 0.70 — differing only in which way the two edges touching the covariate point. In the first the covariate causes both and adjusting for it recovers the effect of 0.50 exactly. In the second the treatment causes the covariate, the effect is 1.13, and adjusting returns 0.50 — the direct edge alone, with the part that travels through the covariate deleted. In the third the treatment and the outcome both cause the covariate, the effect is 0.50, and adjusting returns -0.087. The regression that produces those three numbers is one formula, and nothing in the data says which panel it is being run in.
Fig. 6 The three worlds as diagrams, with the effect each holds and what the two regressions return. The matched triple in this essay is these three structures at edge strengths chosen to give them one covariance matrix.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Causal diagramCausal effectConditional independenceCovariance matrixFisher transformIdentificationMarkov equivalenceNull hypothesisObservational equivalencePartial correlationSample sizeStatistical powerStructural modelV structure