The sample is a condition
Worth reading first: One arithmetic, three decisions.
Two quantities are independent in the population, by construction, with a correlation of exactly zero. Take everyone whose sum of the two exceeds its median — half the population, no measurement error, no missing values, no non-response, every unit real and every value recorded correctly — and inside that sample the two correlate at
That number is exact. It is not a limit, an approximation or a simulated estimate; it is the value the arithmetic returns, and it appears because the variance of a normal truncated at its own mean is and for no other reason.
The word for this is not bias in the ordinary sense, because nothing about the sample is wrong. The sample is a condition: being in it is a statement about both quantities at once, and a statement about two things at once is information about each given the other. The covariate this ladder has been deciding whether to adjust for is not a column in the dataset here. It is the rule that decided which rows there are.
The rule that decided who is in the data
Two causes, and , independent standard normals. A rule that keeps a unit when
for weights and a threshold . At and equal weights that keeps 50.0% of the population — half of everyone, which is as far from a fringe sample as a rule of this shape gets.
Read as a causal diagram it is the shortest possible collider. Both causes point into the selection indicator, the indicator is held fixed at “in the sample” for every row anybody analyses, and holding a common effect fixed makes its causes dependent. That is the same structure as the pre-treatment covariate that satisfies every rule of thumb, with one difference that matters more than it looks: there, the collider was a column, and somebody chose to put it in the model. Here it is not a column at all, so there is no decision to make and no way to leave it out. Every analysis of the sample is already conditioned on it.
The examples are ordinary rather than contrived, and the oldest of them gives the effect its name. Two unrelated conditions both raise the chance of being in hospital, so among inpatients they are negatively associated — which is Berkson’s paradox, stated about hospital records and true of every sample assembled the same way. Two unrelated qualities both raise the chance of admission, so among those admitted they trade off. Two unrelated properties both raise the chance that a case is recorded, so among recorded cases they appear to substitute for each other. In every one of them the association is real inside the sample, the sample is what anybody has, and the population value is zero.
Rotating onto the direction the selection acts in
The arithmetic is exact because the selection acts along one direction and leaves the perpendicular one untouched.
Write the pair in a rotated basis: the component along , and the component across it. Both are normal and they are independent, because a rotation of independent standard normals is independent standard normals. The rule constrains the first and says nothing about the second. So the selected distribution is a truncated normal in one coordinate and an untouched normal in the other, with the two still independent — the picture is a normal cloud with one side of it sliced off, which is exactly what the figure above shows.
Everything then follows from the truncated variance. With and the mean of the truncated component, its variance is
and rotating back gives
which at equal weights collapses to .
The sign is settled by that expression before any number is put into it. A truncated normal always has a variance below one, so is always negative and the numerator is always negative for positive weights. Selecting on a sum of two causes always induces a negative association between them, at every threshold, and the direction is a property of the selection rather than of where the line happened to be put.
At the median, , so and . Then
with arriving from the normal density at zero and from nowhere else. To ten places that is −0.4669422069, and nothing in the derivation is approximate at any step of it.
Twenty-five thresholds, and the worst is 1.81 standard errors
The form above is calculus on a truncated normal. Beside it: sixty thousand draws at each of twenty-five thresholds from −2 to 2.6, keeping whatever falls past the line and computing the sample correlation of what is kept.
The largest disagreement between the Monte Carlo count and the closed form anywhere in that sweep is 0.0109 in correlation, which is 1.81 of the count’s own standard errors. Nothing on the counting side ever forms a truncated moment and nothing on the closed-form side ever draws a row, so the agreement is two routes to one number rather than one route checking itself — the discipline two routes to every number is about, and the reason a result this counter-intuitive is worth stating without hedging.
Across the sweep the induced correlation is negative at every threshold and strengthens at every step. At a threshold keeping 92.1% of the population it is −0.1433; at 76.0%, −0.2953; at half, −0.4669; at 36.2%, −0.5463; at 24.0%, −0.6169; at the conventional 5% point of the sum, keeping 12.2%, it is −0.6934; at 7.9%, −0.7288; and at 3.3% of the population it is −0.7787.
The last of those is the number to hold beside any analysis of a highly selected group. A sample of the top three per cent on a criterion made of two things carries a correlation of nearly −0.78 between those two things, from the selection alone, before anything about the subject has been considered.
The intermediate quantities move the way the mechanism says they should, which is a check worth making because it would catch a sign error the correlation itself would not. The truncated component’s mean rises monotonically from 0.1593 at the loosest threshold to 2.2310 at the tightest, and its variance falls from 0.7494 to 0.1244. The correlation is a function of alone at equal weights, so a table in which fell and the correlation did not strengthen would mean the rotation had been applied wrongly, and a table in which exceeded one anywhere would mean the truncation had been applied to the wrong tail. Neither happens at any of the twenty-five settings.
The way a check like this goes hollow is worth naming too, since a count and a form agreeing is only evidence when they could have disagreed. The counted side draws pairs, applies the rule as an inequality on the raw values, and computes a Pearson correlation on whatever survives; it never rotates, never evaluates a normal density, and never uses . The closed-form side never draws anything. A shared error would have to be an error in the definition of the rule itself, which is one line and is the same line the picture is drawn from.
Both halves show it, and the population containing both does not
The reading that would rescue the ordinary intuition is that the selected group is unrepresentative and the discarded group is where the truth is. It is not, and the check is cheap.
Split the population at the median and read the correlation in each half. Above the line: −0.4669 on 50.0%. Below the line: −0.4669 on 50.0%. The two halves are the same number by symmetry, and the population containing both is exactly 0.
At a threshold of 1 the halves are no longer symmetric and both still show it: −0.6169 on the 24.0% above and −0.2953 on the 76.0% below. The kept group shows more of it than the rejected one, because a stricter cut truncates harder — but both show it, in the same direction, and the union of two groups each correlated at under −0.29 is a population correlated at zero.
There is no half to retreat to. Two analysts studying the two halves would both report a negative association, agree with each other, disagree with the truth, and have no disagreement to investigate. That is a harder failure than a biased sample, because a biased sample has a comparison group somewhere.
That is also the sharpest available statement of why this is not what the word bias usually means. A biased measurement is one whose expected value differs from the quantity it estimates, and the repair is to correct the measurement. Here the measurement is exactly right about the population it was taken from: the selected half really does have that correlation, and an analyst who reported it as a fact about the selected half would be reporting a true fact. The error is entirely in the transfer — in taking a quantity that is a property of the conditioned population and reading it as a property of the unconditioned one. Which of those two populations a study means to speak about is not a statistical question, and it is the same question the reversal that no amount of data settles leaves open about which weighting a comparison is meant to use.
It is largest when the two causes matter equally
Unequal weights weaken it, and the way they weaken it says what the mechanism is.
At against the induced correlation at the median is −0.3891 rather than −0.4669; at it is −0.3020. As one cause dominates the selection, the rule stops being a joint condition on two things and becomes very nearly a condition on one of them — and a condition on one variable alone says nothing about its relationship with any other. In the limit the rule truncates the first cause and leaves the second untouched and independent.
So the effect is largest at equal weights, which is the least contrived case rather than the most. A selection rule built from two causes of comparable importance is the ordinary shape of a selection rule, and it is the shape that does the most damage.
What a small selected sample looks like, and why the picture is noisy
The cloud in this essay draws nine hundred points rather than sixty thousand, and the counted correlations it reports are correspondingly less settled: −0.3234 at a threshold of −1 against a closed form of −0.2953, −0.5050 at the median against −0.4669, −0.6136 at 0.8 against −0.5898, and −0.6980 at 1.6 against −0.6886.
That gap is the point of the picture rather than a defect in it. A stringent threshold produces a strong correlation read on very few rows: of nine hundred draws, 665 survive a threshold of −1, 453 the median, 240 a threshold of 0.8 and 102 a threshold of 1.6. The most dramatic reading in the sequence is the one computed from a hundred and two observations, and the visible flattening of the cloud is what a hundred and two selected points look like.
An analyst meeting the last of those in real data has a strong association on a small sample and every reason to believe it, since the standard error of a correlation of −0.70 on a hundred rows is about 0.05 and the association is fourteen standard errors from zero. Every conventional check passes. The association is real inside the sample, the sample is what exists, and the population value is zero.
What this is not
Two other things on this site are called selection and neither of them is this one, so the distinction is worth stating in full rather than assumed.
Selection on a statistic is the winner’s curse: a quantity is estimated many times, the largest estimate is reported, and it is too large because it was chosen for being large. That failure is about a sampling distribution — it disappears in a large enough sample of the chosen quantity and it is repaired by shrinking the reported value, which is what the estimate after the choice and the interval after the choice measure. Nothing here is repaired by more data. The correlation of −0.4669 is a population quantity of the selected population, and a million selected rows estimate it more precisely.
Selection among models is a third thing again, and shares only the word: a criterion is evaluated over a family of candidate fits and the winner’s apparent quality is inflated by the search. That is a failure of a criterion under multiplicity, and it is repaired by charging for the search.
And this is not regression to the mean. The effect that appears with no intervention at all is about one variable measured twice and an imperfect correlation between the measurements; it is predicted from the correlation and it does not require two causes or a joint condition. The resemblance is that both are consequences of arithmetic that get read as effects of a treatment, which is the family resemblance across everything in this ladder rather than a shared mechanism.
What it does to a treatment effect
The essay has so far measured an association between two causes, and the field is about estimating an effect, so the bridge is worth building rather than left implied.
Let the two quantities be a treatment and an outcome, independent because the treatment was randomised and does nothing. Select on a rule depending on both — units are recorded only if they were treated or they did well, or only if some score built from both exceeded a bar — and inside the recorded data the treatment and the outcome correlate at −0.4669 at the median cut. A randomised trial, analysed correctly, on a sample assembled by a rule that touched both arms and the outcome, reports a substantial harmful effect of a treatment that does nothing. Randomisation deletes the arrows into the treatment and this route does not use one: it goes forward from the treatment into the selection indicator and back out of it into the outcome.
That is why this failure sits outside everything the design essays protect against. What randomisation buys is a reference distribution for the assignment, and the assignment is not what went wrong. Adjusting for the right covariates does not help either, because the conditioning that caused the trouble is not in the model and cannot be taken out of it — it is the same argument as the pre-treatment covariate that satisfies every rule of thumb with the covariate moved out of the dataset entirely. The only repairs are to know the selection rule and model it, or to have a sample that was not selected on the outcome.
And there is no diagnostic. The data has the shape the analysis expects, the residuals behave, the standard errors are right for the population that was sampled, and the estimate is a consistent estimate of a real parameter of a real population. This is the same reason a description that survives censoring still needs everything a causal claim needs: internal consistency is not a check on what the number is about.
Where this does not hold
The causes are jointly normal and the rule is a linear threshold. That is what makes the truncation a truncated normal and the answer exact. A different joint law or a rule that is not a half-plane gives an association of the same sign — any selection that depends on both causes does — but not this number, and not necessarily a bound this clean.
The two causes are exactly independent to begin with. A real pair is not, and what the sample then reads is the population correlation plus this term, which means the induced association can equally well hide a real one as manufacture a false one. Nothing here measures that composition, and it is the case a reader is most likely to be in.
The selection rule is known. It has to be, for the closed form to exist, and in practice it is exactly what nobody knows: the units that are not in the data are the units nothing was recorded about. A sensitivity analysis in the weights and the threshold is the honest tool, and this essay’s sweep is what such an analysis would range over rather than a claim about any dataset.
What remains true across all three of those limits is the shape of the finding, and it is the one to carry: an association measured inside a sample is an association conditional on the rule that made the sample, and that rule is part of the analysis whether or not anybody wrote it down. The rest of this ladder decides which columns to put in a model; this one says that the row filter is a column too, and it is one that cannot be left out. That the two are not distinguishable from inside the data is the same statement one covariance matrix holding two effects makes about arrows, arriving here as a statement about rows.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A lead that a heavy tail keeps — both name closed form, correlation, monte carlo, selection bias
- A threshold in the tail — both name closed form, correlation, normal distribution, threshold
- The arcsine that closes it, and the error that was overstated — both name closed form, correlation, monte carlo, threshold
- The effect a stopped trial reports — both name closed form, conditional distribution, monte carlo, selection bias
- The fourth moment that was missing — both name closed form, correlation, monte carlo, threshold
- Three mechanisms and one dataset — both name closed form, conditional distribution, monte carlo, selection bias
Named objects
A flat tag is an object no other essay names yet.
Berksons paradoxCausal diagramClosed formColliderCollider biasConditional distributionCorrelationIndependenceMonte CarloNormal distributionSample selectionSelection biasThresholdTruncated normal