What conditioning on a variable does

The sample is a condition

Two independent standard normals, selected on their sum exceeding its median, read a correlation of exactly −1/(π − 1) = −0.4669 inside the sample. Nothing is measured badly and nothing is missing — and both halves of that split read it, in the same direction, while the population containing both reads zero.

Worth reading first: One arithmetic, three decisions.

Two quantities are independent in the population, by construction, with a correlation of exactly zero. Take everyone whose sum of the two exceeds its median — half the population, no measurement error, no missing values, no non-response, every unit real and every value recorded correctly — and inside that sample the two correlate at

ρ=1π1=0.4669.\rho = -\frac{1}{\pi - 1} = -0.4669 .

That number is exact. It is not a limit, an approximation or a simulated estimate; it is the value the arithmetic returns, and it appears because the variance of a normal truncated at its own mean is 12/π=0.3633801 - 2/\pi = 0.363380 and for no other reason.

The word for this is not bias in the ordinary sense, because nothing about the sample is wrong. The sample is a condition: being in it is a statement about both quantities at once, and a statement about two things at once is information about each given the other. The covariate this ladder has been deciding whether to adjust for is not a column in the dataset here. It is the rule that decided which rows there are.

The rule that decided who is in the data

Two causes, xx and yy, independent standard normals. A rule that keeps a unit when

w1x+w2y  >  θ,w_1 x + w_2 y \;>\; \theta ,

for weights w1,w2>0w_1, w_2 > 0 and a threshold θ\theta. At θ=0\theta = 0 and equal weights that keeps 50.0% of the population — half of everyone, which is as far from a fringe sample as a rule of this shape gets.

Read as a causal diagram it is the shortest possible collider. Both causes point into the selection indicator, the indicator is held fixed at “in the sample” for every row anybody analyses, and holding a common effect fixed makes its causes dependent. That is the same structure as the pre-treatment covariate that satisfies every rule of thumb, with one difference that matters more than it looks: there, the collider was a column, and somebody chose to put it in the model. Here it is not a column at all, so there is no decision to make and no way to leave it out. Every analysis of the sample is already conditioned on it.

The examples are ordinary rather than contrived, and the oldest of them gives the effect its name. Two unrelated conditions both raise the chance of being in hospital, so among inpatients they are negatively associated — which is Berkson’s paradox, stated about hospital records and true of every sample assembled the same way. Two unrelated qualities both raise the chance of admission, so among those admitted they trade off. Two unrelated properties both raise the chance that a case is recorded, so among recorded cases they appear to substitute for each other. In every one of them the association is real inside the sample, the sample is what anybody has, and the population value is zero.

The line is the sample, and the sample is the finding900 draws of two independent standard normal causes, with the 453 of them past a threshold of 0.00 marked and the 447 that fall short left pale. In the population the two are independent by construction. Inside the selected sample the correlation is -0.4669 in closed form and -0.5050 counted on these 453 rows, and the least-squares line through them has a slope of -0.545. The mechanism is visible in the picture rather than argued: the threshold removes one corner of the cloud, and a cloud with a corner missing is a cloud whose two coordinates carry information about each other.-202-202the first causethe second causeselected r = -0.5050population r = 0900 draws, 453 kept past 0.00closed form -0.4669
Fig. 1 Nine hundred draws of two independent causes, with those past the threshold marked and those short of it left pale. The slider moves the threshold; the counted correlation among the kept rows is −0.5050 at the median and −0.6980 at a threshold keeping about a ninth of them.

Rotating onto the direction the selection acts in

The arithmetic is exact because the selection acts along one direction and leaves the perpendicular one untouched.

Write the pair in a rotated basis: the component along (w1,w2)(w_1, w_2), and the component across it. Both are normal and they are independent, because a rotation of independent standard normals is independent standard normals. The rule constrains the first and says nothing about the second. So the selected distribution is a truncated normal in one coordinate and an untouched normal in the other, with the two still independent — the picture is a normal cloud with one side of it sliced off, which is exactly what the figure above shows.

Everything then follows from the truncated variance. With τ=θ/w\tau = \theta/\lVert w \rVert and m=φ(τ)/(1Φ(τ))m = \varphi(\tau)/(1 - \Phi(\tau)) the mean of the truncated component, its variance is

s2=1+τmm2,s^{2} = 1 + \tau m - m^{2} ,

and rotating back gives

ρ=w1w2(s21)(w12s2+w22)(w22s2+w12),\rho = \frac{w_1 w_2 (s^{2} - 1)}{\sqrt{(w_1^{2}s^{2} + w_2^{2})(w_2^{2}s^{2} + w_1^{2})}} ,

which at equal weights collapses to (s21)/(s2+1)(s^2 - 1)/(s^2 + 1).

The sign is settled by that expression before any number is put into it. A truncated normal always has a variance below one, so s21s^2 - 1 is always negative and the numerator is always negative for positive weights. Selecting on a sum of two causes always induces a negative association between them, at every threshold, and the direction is a property of the selection rather than of where the line happened to be put.

At the median, τ=0\tau = 0, so m=2/π=0.7979m = \sqrt{2/\pi} = 0.7979 and s2=12/π=0.3634s^2 = 1 - 2/\pi = \mathbf{0.3634}. Then

ρ=s21s2+1=2/π22/π=1π1=0.4669422069,\rho = \frac{s^{2}-1}{s^{2}+1} = \frac{-2/\pi}{2 - 2/\pi} = -\frac{1}{\pi - 1} = -0.4669422069 ,

with π\pi arriving from the normal density at zero and from nowhere else. To ten places that is −0.4669422069, and nothing in the derivation is approximate at any step of it.

Twenty-five thresholds, and the worst is 1.81 standard errors

The form above is calculus on a truncated normal. Beside it: sixty thousand draws at each of twenty-five thresholds from −2 to 2.6, keeping whatever falls past the line and computing the sample correlation of what is kept.

The largest disagreement between the Monte Carlo count and the closed form anywhere in that sweep is 0.0109 in correlation, which is 1.81 of the count’s own standard errors. Nothing on the counting side ever forms a truncated moment and nothing on the closed-form side ever draws a row, so the agreement is two routes to one number rather than one route checking itself — the discipline two routes to every number is about, and the reason a result this counter-intuitive is worth stating without hedging.

Across the sweep the induced correlation is negative at every threshold and strengthens at every step. At a threshold keeping 92.1% of the population it is −0.1433; at 76.0%, −0.2953; at half, −0.4669; at 36.2%, −0.5463; at 24.0%, −0.6169; at the conventional 5% point of the sum, keeping 12.2%, it is −0.6934; at 7.9%, −0.7288; and at 3.3% of the population it is −0.7787.

The last of those is the number to hold beside any analysis of a highly selected group. A sample of the top three per cent on a criterion made of two things carries a correlation of nearly −0.78 between those two things, from the selection alone, before anything about the subject has been considered.

The intermediate quantities move the way the mechanism says they should, which is a check worth making because it would catch a sign error the correlation itself would not. The truncated component’s mean mm rises monotonically from 0.1593 at the loosest threshold to 2.2310 at the tightest, and its variance s2s^2 falls from 0.7494 to 0.1244. The correlation is a function of s2s^2 alone at equal weights, so a table in which s2s^2 fell and the correlation did not strengthen would mean the rotation had been applied wrongly, and a table in which s2s^2 exceeded one anywhere would mean the truncation had been applied to the wrong tail. Neither happens at any of the twenty-five settings.

The way a check like this goes hollow is worth naming too, since a count and a form agreeing is only evidence when they could have disagreed. The counted side draws pairs, applies the rule as an inequality on the raw values, and computes a Pearson correlation on whatever survives; it never rotates, never evaluates a normal density, and never uses π\pi. The closed-form side never draws anything. A shared error would have to be an error in the definition of the rule itself, which is one line and is the same line the picture is drawn from.

Two independent causes, correlated by being in the sample. The correlation between two independent standard normal causes, measured inside the sample of everyone whose sum of the two exceeds a threshold. It is -0.1433 where the threshold keeps 92.1% of the population and -0.7787 where it keeps 3.3%. At the median it is exactly −1/(π − 1) = -0.4669, because the variance of a normal truncated at its own mean is 1 − 2/π. The line is that closed form and the points are counted from 60000 draws apiece; the worst disagreement is 1.81 standard errors. Nothing has been measured badly and nothing is missing: the correlation is a property of the sample, and the sample is a condition.
Fig. 2 The induced correlation against the threshold, closed form as a line and counts from sixty thousand draws apiece as points. It runs from −0.1433 where the rule keeps 92.1% of the population to −0.7787 where it keeps 3.3%.

Both halves show it, and the population containing both does not

The reading that would rescue the ordinary intuition is that the selected group is unrepresentative and the discarded group is where the truth is. It is not, and the check is cheap.

Split the population at the median and read the correlation in each half. Above the line: −0.4669 on 50.0%. Below the line: −0.4669 on 50.0%. The two halves are the same number by symmetry, and the population containing both is exactly 0.

At a threshold of 1 the halves are no longer symmetric and both still show it: −0.6169 on the 24.0% above and −0.2953 on the 76.0% below. The kept group shows more of it than the rejected one, because a stricter cut truncates harder — but both show it, in the same direction, and the union of two groups each correlated at under −0.29 is a population correlated at zero.

There is no half to retreat to. Two analysts studying the two halves would both report a negative association, agree with each other, disagree with the truth, and have no disagreement to investigate. That is a harder failure than a biased sample, because a biased sample has a comparison group somewhere.

That is also the sharpest available statement of why this is not what the word bias usually means. A biased measurement is one whose expected value differs from the quantity it estimates, and the repair is to correct the measurement. Here the measurement is exactly right about the population it was taken from: the selected half really does have that correlation, and an analyst who reported it as a fact about the selected half would be reporting a true fact. The error is entirely in the transfer — in taking a quantity that is a property of the conditioned population and reading it as a property of the unconditioned one. Which of those two populations a study means to speak about is not a statistical question, and it is the same question the reversal that no amount of data settles leaves open about which weighting a comparison is meant to use.

Both halves are correlated and the whole is not. The correlation between two independent causes, read inside each half of one split at a threshold of 0.00. The half above the line reads -0.4669 on 50.0% of the population and the half below reads -0.4669 on 50.0%; the two together read exactly zero. Neither half is measured badly and neither is a biased sample in the ordinary sense of the word — every unit in it is a real unit with real values. What has happened is that being in the half is a condition on both causes at once, so knowing one of them says something about the other. A study of the selected half and a study of the rejected half would both report an association, in the same direction, and disagree with the population that contains both.
Fig. 3 The correlation inside each half of a split at the median, and in the population containing both. Each half reads −0.4669 on 50.0% of the population; the two together read exactly zero.

It is largest when the two causes matter equally

Unequal weights weaken it, and the way they weaken it says what the mechanism is.

At w2=2w_2 = 2 against w1=1w_1 = 1 the induced correlation at the median is −0.3891 rather than −0.4669; at w2=3w_2 = 3 it is −0.3020. As one cause dominates the selection, the rule stops being a joint condition on two things and becomes very nearly a condition on one of them — and a condition on one variable alone says nothing about its relationship with any other. In the limit the rule truncates the first cause and leaves the second untouched and independent.

So the effect is largest at equal weights, which is the least contrived case rather than the most. A selection rule built from two causes of comparable importance is the ordinary shape of a selection rule, and it is the shape that does the most damage.

Both halves are correlated and the whole is not. The correlation between two independent causes, read inside each half of one split at a threshold of 1.00. The half above the line reads -0.6169 on 24.0% of the population and the half below reads -0.2953 on 76.0%; the two together read exactly zero. Neither half is measured badly and neither is a biased sample in the ordinary sense of the word — every unit in it is a real unit with real values. What has happened is that being in the half is a condition on both causes at once, so knowing one of them says something about the other. A study of the selected half and a study of the rejected half would both report an association, in the same direction, and disagree with the population that contains both.
Fig. 4 The same split at a threshold of 1 rather than at the median. The 24.0% above the line read −0.6169 and the 76.0% below read −0.2953; the population holding both still reads zero.

What a small selected sample looks like, and why the picture is noisy

The cloud in this essay draws nine hundred points rather than sixty thousand, and the counted correlations it reports are correspondingly less settled: −0.3234 at a threshold of −1 against a closed form of −0.2953, −0.5050 at the median against −0.4669, −0.6136 at 0.8 against −0.5898, and −0.6980 at 1.6 against −0.6886.

That gap is the point of the picture rather than a defect in it. A stringent threshold produces a strong correlation read on very few rows: of nine hundred draws, 665 survive a threshold of −1, 453 the median, 240 a threshold of 0.8 and 102 a threshold of 1.6. The most dramatic reading in the sequence is the one computed from a hundred and two observations, and the visible flattening of the cloud is what a hundred and two selected points look like.

An analyst meeting the last of those in real data has a strong association on a small sample and every reason to believe it, since the standard error of a correlation of −0.70 on a hundred rows is about 0.05 and the association is fourteen standard errors from zero. Every conventional check passes. The association is real inside the sample, the sample is what exists, and the population value is zero.

The line is the sample, and the sample is the finding. 900 draws of two independent standard normal causes, with the 102 of them past a threshold of 1.60 marked and the 798 that fall short left pale. In the population the two are independent by construction. Inside the selected sample the correlation is -0.6886 in closed form and -0.6980 counted on these 102 rows, and the least-squares line through them has a slope of -0.642. The mechanism is visible in the picture rather than argued: the threshold removes one corner of the cloud, and a cloud with a corner missing is a cloud whose two coordinates carry information about each other.
Fig. 5 The same nine hundred draws at a threshold keeping 102 of them. The correlation counted on those rows is −0.6980 against a closed form of −0.6886, and the shape the selection leaves is the whole of the mechanism.

What this is not

Two other things on this site are called selection and neither of them is this one, so the distinction is worth stating in full rather than assumed.

Selection on a statistic is the winner’s curse: a quantity is estimated many times, the largest estimate is reported, and it is too large because it was chosen for being large. That failure is about a sampling distribution — it disappears in a large enough sample of the chosen quantity and it is repaired by shrinking the reported value, which is what the estimate after the choice and the interval after the choice measure. Nothing here is repaired by more data. The correlation of −0.4669 is a population quantity of the selected population, and a million selected rows estimate it more precisely.

Selection among models is a third thing again, and shares only the word: a criterion is evaluated over a family of candidate fits and the winner’s apparent quality is inflated by the search. That is a failure of a criterion under multiplicity, and it is repaired by charging for the search.

And this is not regression to the mean. The effect that appears with no intervention at all is about one variable measured twice and an imperfect correlation between the measurements; it is predicted from the correlation and it does not require two causes or a joint condition. The resemblance is that both are consequences of arithmetic that get read as effects of a treatment, which is the family resemblance across everything in this ladder rather than a shared mechanism.

What it does to a treatment effect

The essay has so far measured an association between two causes, and the field is about estimating an effect, so the bridge is worth building rather than left implied.

Let the two quantities be a treatment and an outcome, independent because the treatment was randomised and does nothing. Select on a rule depending on both — units are recorded only if they were treated or they did well, or only if some score built from both exceeded a bar — and inside the recorded data the treatment and the outcome correlate at −0.4669 at the median cut. A randomised trial, analysed correctly, on a sample assembled by a rule that touched both arms and the outcome, reports a substantial harmful effect of a treatment that does nothing. Randomisation deletes the arrows into the treatment and this route does not use one: it goes forward from the treatment into the selection indicator and back out of it into the outcome.

That is why this failure sits outside everything the design essays protect against. What randomisation buys is a reference distribution for the assignment, and the assignment is not what went wrong. Adjusting for the right covariates does not help either, because the conditioning that caused the trouble is not in the model and cannot be taken out of it — it is the same argument as the pre-treatment covariate that satisfies every rule of thumb with the covariate moved out of the dataset entirely. The only repairs are to know the selection rule and model it, or to have a sample that was not selected on the outcome.

And there is no diagnostic. The data has the shape the analysis expects, the residuals behave, the standard errors are right for the population that was sampled, and the estimate is a consistent estimate of a real parameter of a real population. This is the same reason a description that survives censoring still needs everything a causal claim needs: internal consistency is not a check on what the number is about.

Where this does not hold

The causes are jointly normal and the rule is a linear threshold. That is what makes the truncation a truncated normal and the answer exact. A different joint law or a rule that is not a half-plane gives an association of the same sign — any selection that depends on both causes does — but not this number, and not necessarily a bound this clean.

The two causes are exactly independent to begin with. A real pair is not, and what the sample then reads is the population correlation plus this term, which means the induced association can equally well hide a real one as manufacture a false one. Nothing here measures that composition, and it is the case a reader is most likely to be in.

The selection rule is known. It has to be, for the closed form to exist, and in practice it is exactly what nobody knows: the units that are not in the data are the units nothing was recorded about. A sensitivity analysis in the weights and the threshold is the honest tool, and this essay’s sweep is what such an analysis would range over rather than a claim about any dataset.

What remains true across all three of those limits is the shape of the finding, and it is the one to carry: an association measured inside a sample is an association conditional on the rule that made the sample, and that rule is part of the analysis whether or not anybody wrote it down. The rest of this ladder decides which columns to put in a model; this one says that the row filter is a column too, and it is one that cannot be left out. That the two are not distinguishable from inside the data is the same statement one covariance matrix holding two effects makes about arrows, arriving here as a statement about rows.

The same covariate, three ways round. Three worlds over a treatment, an outcome and a covariate, joined by the same three edges at the same three strengths — 0.90, 0.50 and 0.70 — differing only in which way the two edges touching the covariate point. In the first the covariate causes both and adjusting for it recovers the effect of 0.50 exactly. In the second the treatment causes the covariate, the effect is 1.13, and adjusting returns 0.50 — the direct edge alone, with the part that travels through the covariate deleted. In the third the treatment and the outcome both cause the covariate, the effect is 0.50, and adjusting returns -0.087. The regression that produces those three numbers is one formula, and nothing in the data says which panel it is being run in.
Fig. 6 The three arrangements of a covariate beside a treatment and an outcome. The selection rule in this essay is the third of them, with the covariate replaced by an indicator that is fixed at “in the sample” for every row.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Berksons paradoxCausal diagramClosed formColliderCollider biasConditional distributionCorrelationIndependenceMonte CarloNormal distributionSample selectionSelection biasThresholdTruncated normal