The value that is not there

The assumption that identifies the mechanism

A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.

Worth reading first: Three mechanisms and one dataset.

Two worlds with identical recorded data — the same numbers to the last bit, with true slopes of 0.6 and 0.315452 — are the construction behind the conclusion that no statistic, test or likelihood ratio can say which world a dataset came from. The coefficient that would decide it — how strongly the outcome itself decides whether it is recorded — is not a quantity the data contain.

There is a standard model that reports that coefficient anyway, with a standard error and a test. A selection model writes the outcome as a normal regression, writes the chance of being recorded as a probit that may read the outcome, and fits both by maximum likelihood. It returns an estimate of the very number the twin construction says cannot be estimated. Both statements are correct, and the reconciliation is one assumption: the outcome is normal. Everything the model knows about the mechanism, it knows through that.

So this essay measures what the assumption buys and what it costs, on four kinds of study of 200 and 800 rows with 35% of outcomes missing. Where the outcome does decide and the residuals are normal, the model repairs what dropping the rows breaks: complete cases put the slope at 0.4318 and cover 1.5% at 800 rows, and the selection model puts it at 0.5795 and covers 90.5%. Where the missingness is at random and the residuals are skewed, the same model finds selection that is not there, puts the slope at 1.0319 against a truth of 0.6, and rejects missingness at random in 72.5% of studies.

A likelihood that can see the mechanism

The model has two equations. The outcome is y=α+βx+γz+σεy = \alpha + \beta x + \gamma z + \sigma\varepsilon with ε\varepsilon standard normal, and a row is recorded when c0+cxx+czz+cyy+u>0c_0 + c_x x + c_z z + c_y y + u > 0 with uu standard normal and independent of everything. The coefficient cyc_y is the mechanism: zero is missingness at random on the covariates, and anything else is missingness that depends on the value that went missing.

A recorded row contributes its outcome’s normal density times the chance of being recorded at that outcome. An unrecorded row has no outcome, so it contributes the chance of not being recorded with the outcome integrated out — and because the outcome is normal and the index is linear in it, that integral is closed:

P(unrecordedx,z)=Φ ⁣(c0+cxx+czz+cyμ1+cy2σ2),μ=α+βx+γz.P(\text{unrecorded} \mid x, z) = \Phi\!\left(-\frac{c_0 + c_x x + c_z z + c_y \mu}{\sqrt{1 + c_y^2 \sigma^2}}\right), \qquad \mu = \alpha + \beta x + \gamma z.

That line is where the identification happens. The unrecorded rows carry a count, and the count becomes information about cyc_y only because the law of the outcomes nobody saw is assumed to be the same normal law as the law of the ones that were seen, shifted by the covariates. Change the assumed law and the integral changes, and so does what the count says.

A peak where the recorded data have none. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 0.35, where the fitted slope is 0.839, and the values of cy within the 95% cut run from −0.13 to 0.75; the likelihood-ratio statistic against cy = 0 is 1.47. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.
Fig. 1 One study of 800 rows whose missingness is at random and whose residuals are normal: the log-likelihood with cyc_y held at each value and everything else maximised, against its maximum. A likelihood that left the outcome’s law free would be flat across the whole axis.

On one study of 800 rows drawn at random with normal residuals, the profile likelihood in cyc_y is not flat. It peaks at 0.349, falls by 48.52 units by cy=2c_y = -2, and the likelihood-ratio statistic against zero is 1.47, well inside the 5% cut. That curvature is not in the recorded rows. The twin construction shows there is none there to find, since a world with the unseen outcomes shifted produces the same recorded rows. The curvature is the normal law, drawn as a function of the one parameter it pins down.

The same study shows what rides on the peak. Held at cy=0.75c_y = -0.75 the fitted slope is 0.410; at cy=1.125c_y = 1.125 it is 1.081. The selection model’s profile carries the twin construction’s sensitivity line inside it — a slope for every value of the mechanism — and then uses the normal law to choose one point on it. At this study’s peak that point is a slope of 0.8395, a quarter of a unit from the truth, on a study where nothing is wrong with the model at all.

The twins, again

The model chooses a world, and says so with an interval. The true slope in the two worlds whose recorded data are identical, beside what complete cases and a normal selection model estimate from those data, averaged over 200 studies of 800 rows. Both estimates are the same numbers in both worlds, bit for bit. Complete cases read 0.6037 and the selection model 0.6151: the model has estimated the mechanism and concluded the missingness is at random, because the recorded rows look like a normal outcome truncated by a rule on the covariates. Its 95% interval covers the first world's slope of 0.60 on 85.5% of studies and the second world's 0.3155 on 30.5%. Nothing in the data preferred the first world; the assumption that the outcomes nobody saw were normal did.
Fig. 2 The true slope in the two worlds with identical recorded data, beside complete cases and the selection model, each of which returns the same number in both. The model’s interval covers the first world’s slope far more often than the second’s.

Fitted to the two worlds’ recorded data, the selection model returns the same parameters, the same likelihood and the same standard errors, bit for bit — equal as numbers, not approximately, since it reads nothing the two worlds do not share. So it cannot tell the worlds apart either. What it can do is pick one, and it does. Over 200 studies of 800 rows it reports a slope of 0.6151 and a mechanism with a median of 0.016, which is the first world: missingness at random, slope 0.6. Its 95% interval covers that world’s slope on 85.5% of studies and the second world’s 0.315452 on 30.5%.

The reason is not that the recorded data favour the first world. It is that the second world is outside the model. Its unrecorded outcomes are shifted by a whole unit, so its complete data are a normal regression on the recorded rows and a shifted one on the others, which no choice of α\alpha, β\beta, γ\gamma, σ\sigma and cyc_y can write down. A model that can only express normal complete data finds the one world of the two in which the complete data are normal, and reports it with an interval. Nothing in that interval is a statement about the second world, and a reader holding it has no way to know that.

Where the assumption is true

A peak at a mechanism that is there. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is on the outcome, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 1.45, where the fitted slope is 0.721, and the values of cy within the 95% cut run from 1.00 to 1.88; the likelihood-ratio statistic against cy = 0 is 13.59. The study was drawn with cy = 1.20. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.
Fig. 3 One study of 800 rows in which the outcome does decide whether it is recorded, with normal residuals: the same profile likelihood, peaked well away from zero.

Now draw the studies the model was written for: the outcome reads itself into the chance of being recorded with a coefficient of 1.2, and the residuals are normal. Dropping the incomplete rows is exactly wrong here — the field’s first table put the attenuation at 0.727448 of every coefficient — and complete cases report 0.4367 at 200 rows and 0.4318 at 800, covering 48.7% and then 1.5% as the interval narrows around the wrong number, the collapse dropping the incomplete rows measured as the sample grew.

The selection model reports 0.5458 at 200 rows and 0.5795 at 800, and its interval covers 80.0% and 90.5%. The mechanism it estimates has a median of 1.249 at 200 rows and 1.184 at 800, against the 1.2 the studies were drawn with, and at 800 rows nine studies in ten put it between 0.840 and 1.500. On the one study in the figure the profile peaks at 1.450 and the likelihood-ratio statistic against zero is 13.59.

That is the case for the model, and it is a real one. When the outcome does decide what is recorded and the outcome is normal, the model recovers most of a bias no amount of data removes from complete cases, and its test finds the dependence in 76.5% of studies of 800 rows.

The repair is not free at the smaller size, and the error column says how much of it survives. At 800 rows the selection model’s root mean squared error is 0.0770 against complete cases’ 0.1728, less than half. At 200 rows it is 0.1681 against 0.1843: the bias has been mostly removed and the spread that freeing the mechanism adds has taken nearly all of the gain back. A study of two hundred rows with a third of its outcomes missing is not a small study by the standards of most fields, and on it the model that is exactly right is barely more accurate than the one that is exactly wrong. What it is instead is centred. Complete cases average 1.91 of their own spreads below the truth at 200 rows, so nearly every study errs the same way; the selection model averages 0.34 of its spread below, so its errors fall on both sides. That is worth something to a reader combining studies, and it does not show in a single error figure.

What a skewed residual looks like to it

A peak built out of a skewed residual. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals skewed right; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 2.41, where the fitted slope is 1.089, and the values of cy within the 95% cut run from 1.38 to 3.88; the likelihood-ratio statistic against cy = 0 is 5.35. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.
Fig. 4 One study of 800 rows whose missingness is at random but whose residuals are skewed to the right: the profile likelihood peaks far from zero, with a second, lower peak near it.

Keep the missingness exactly as it was in the first world — a probit on the covariates, nothing to do with the outcome — and change only the shape of the residual, to a gamma with shape four, standardised: mean zero, variance one, skewness one. Complete cases do not care. Missingness on the regressors leaves the complete-case slope exactly right under any residual law, and they report 0.5971 at 800 rows and cover 96.5%.

The selection model cares a great deal. On the study in the figure its profile peaks at cy=c_y = 2.406 with a slope of 1.089, and the likelihood-ratio statistic against zero is 5.35. Over 200 studies of 800 rows the mechanism’s median is 1.895, the slope averages 1.0319, and the interval covers the true 0.6 on 14.0% of studies.

The reading that fits is the model doing what it was built to do with evidence it misreads. A recording rule whose chance rises with the outcome removes low outcomes more often than high ones, so the recorded residuals lose part of their lower tail and come out skewed to the right. A normal model shown recorded residuals that are skewed to the right has exactly one way to explain them: selection on the outcome, with cy>0c_y > 0. Having concluded that, it corrects the slope the way outcome selection requires, upward, because complete cases under that mechanism are attenuated. Every step is correct inference under the assumption, and the assumption is the only thing wrong.

The figure’s second peak is part of the same story. At cy=0.125c_y = -0.125 the profile has a local maximum 2.52 below the top — the missing-at-random reading, which is true and which the likelihood ranks second. That is not rare. At 200 rows, at least two of the nine starting points of the fit end on different peaks in 90.8% of skewed studies, and on 23.0% of them the best fit puts cyc_y past six, where the likelihood is still climbing towards a hard cut on the outcome.

Four studies side by side

Where the selection model repairs a slope, and where it breaks one. Complete-case least squares and a normal selection model, each as the mean slope and one standard deviation across 200 studies of 800 rows, on four kinds of study; the dashes mark the true slope, 0.60, and the right-hand column is how often each 95% interval covers it. When the missingness reads the outcome and the residuals are normal, complete cases read 0.4318 and the selection model 0.5795. When the missingness is at random and the residuals are skewed, complete cases read 0.5971 and the selection model 1.0319, covering 14.0%. With heavy-tailed residuals the model reads 0.5882. In the study it is built for — at random, normal residuals — it reads 0.6151 with a spread of 0.1310 against complete cases' 0.0580.
Fig. 5 Complete cases and the selection model on four kinds of study of 800 rows, each as a mean slope with one standard deviation either side, and the share of intervals covering the true 0.6.

At 800 rows the two estimators trade places cleanly. Complete cases are right on the three studies whose missingness is at random, whatever the residual, and wrong on the one where the outcome decides. The selection model is close to right on the two studies with normal residuals and wrong on the skewed one, with a root mean squared error of 0.4833 on a slope of 0.6, while complete cases on the same data have 0.0576.

The heavy-tailed residual, a tt on five degrees of freedom scaled to unit variance, sits between. It is symmetric, so there is no lower tail to explain away and no direction for the model to lean, and the slope averages 0.5882. But its tails are not normal either, and the interval covers 77.0%.

Even the study the model is exactly right about costs something. With missingness at random and normal residuals the selection model’s slope has a standard deviation of 0.1310 at 800 rows against complete cases’ 0.05802.26 times the spread, and five times the variance — because freeing cyc_y lets every recorded row’s outcome also inform the mechanism, and the slope and the mechanism move together along the profile. Its interval covers 85.5%, short of 95% on the one kind of study where every assumption it makes is true.

A test of the mechanism that tests the shape

A test of the mechanism that tests the shape. How often the likelihood-ratio test of cy = 0 rejects at 5%, on four kinds of study at 200 and 800 rows. Only the studies drawn with the outcome deciding what is recorded should be rejected, and there it rejects 36.5% and 76.5%. On studies whose missingness is at random with normal residuals it rejects 7.2% and 6.5%. On studies whose missingness is exactly as much at random, with residuals skewed right, it rejects 53.2% and 72.5%, and with heavy-tailed residuals 13.2% and 12.0%. Every one of those rejections is wrong about the mechanism and right about the shape.
Fig. 6 How often the likelihood-ratio test of cy=0c_y = 0 rejects at 5%, on the four kinds of study at 200 and 800 rows. Only the studies in which the outcome decides should be rejected.

The likelihood-ratio test of cy=0c_y = 0 is the diagnostic the twin construction says cannot exist, and the model supplies it. On studies at random with normal residuals it rejects 7.2% at 200 rows and 6.5% at 800, near its level. On studies where the outcome decides it rejects 36.5% and 76.5%. On studies at random with a skewed residual it rejects 53.2% and 72.5% — more often than on the studies it is meant to catch at 200 rows and nearly as often at 800, and every one of those rejections is wrong about the mechanism. On the heavy-tailed residual it rejects 13.2% and 12.0%, more than twice its level at both sizes and not falling as the sample grows — a symmetric departure from normality has no direction to lend the mechanism, but its tails still read as something the normal law did not expect, and some of that is spent on cyc_y.

So a rejection is evidence against a pair of claims, missingness at random and a normal residual, and the test cannot say which of the two failed. At 800 rows the skewed studies and the outcome-driven ones are rejected at 72.5% and 76.5%. A reader told only that cy=0c_y = 0 was rejected has almost no information about which of those worlds they are in, and the one they are most likely to conclude — the one the test is named for — is wrong in the first.

This is the same absence the twins showed, reached from the other side. There the data could not see the mechanism and every statistic said nothing. Here a statistic says something with confidence, and what it has seen is the residual’s shape.

Where the model’s estimate sits

The mechanism the model reports, study by study. The selection model's estimate of cy — how strongly the outcome itself decides whether it is recorded — across studies of 200 and 800 rows, from the tenth to the ninetieth percentile with the median marked; the vertical dashes are cy = 0, missingness at random, and the short marks are the value each study was drawn with. On studies drawn with cy = 1.2 and normal residuals the median estimate is 1.249 at 200 rows and 1.184 at 800. On studies drawn at random with residuals skewed right it is 2.160 and 1.895, and on studies drawn at random with normal residuals 0.036 and 0.016. The axis stops at 4: on at random, skewed right, 200 rows the ninetieth percentile lies beyond it, at 1.1e+8, where the fit's cy ran off with no interior peak.
Fig. 7 The estimated mechanism across studies of 200 and 800 rows, from the tenth to the ninetieth percentile with the median marked, with the value each kind of study was drawn with.

Across studies the estimated mechanism tells the same story as the test. At random with normal residuals, nine studies in ten of 800 rows put cyc_y between −0.325 and 0.409, around a median of 0.016. With a skewed residual the same range is 0.021 to 2.854: almost every study leans towards selection on the outcome, and the median at 1.895 is further from zero than the 1.184 the outcome-driven studies report for a mechanism that is really there.

Seen as a number on the sensitivity line, that is the whole account. A selection model does not estimate where on the line the data sit, because the data sit everywhere on it equally. It computes where a normal outcome would sit, and reports that position with a standard error that measures how precisely the normal law has been applied rather than how well it fits. The same move turns up wherever an analysis cannot see a shape and assumes one: a correction that needs the slope of a density is either estimated from the readings or assumed normal, and the two give different answers exactly where the population is not.

What a reader should take from it

A selection model is a sensitivity analysis with the parameter set by a distributional assumption. Its estimate of the mechanism is not a measurement of the mechanism. It is the point on a line nobody can see where a normal residual would put the analysis, and on this evidence the residual’s shape moves that point further than the mechanism itself does.

Its test of missingness at random is a joint test. A rejection says that missingness at random with a normal residual does not fit; on skewed residuals it says so in 72.5% of studies with the missingness at random throughout. Reporting it as evidence of informative missingness reports one of two things it cannot tell apart, whenever the residual’s shape has not been checked first.

It is the one repair in this field that does not assume the answer. Weighting by an estimated chance of being recorded, pooling imputations and matching the imputation model to the analysis are all correct under missingness at random and inherit that assumption whole. The selection model drops it, and has to put something in its place; what it puts there is a normal law, and a residual law is no easier to defend than a mechanism.

Where it is right it is worth having, and its interval still needs checking. With the outcome deciding and normal residuals it takes a slope complete cases put at 0.4318 back to 0.5795, and covers 90.5% where complete cases cover 1.5%. With nothing wrong at all it covers 85.5% at more than twice complete cases’ spread. A prior on the mechanism is the whole answer when the data say nothing; a normal law is a prior with the word model on it, and it deserves the same statement.

It is also not special to missing outcomes. A dropout nobody can see makes a survival curve converge on the same number from two different worlds, and any model that separated them would be separating them through something it assumed about the part of the population that left.

Still open, and what comes next

A variable that moves recording and not the outcome. Everything here identifies the mechanism through the normal law alone, because the recording rule and the outcome read the same covariates. A variable that shifts the chance of being recorded and has no place in the outcome equation identifies the model without leaning on the shape, and the question worth measuring is how strong it has to be before the skewed studies stop reporting selection. That is the next measurement, and it can be made on the same draws.

An interval for the mechanism that reads the profile. Every interval above is a Wald interval from the curvature at the peak, and the profiles are not quadratic: they have second peaks and, on a quarter of the skewed studies at 200 rows, no interior peak. An interval read off the likelihood-ratio cut would follow the curve instead, and on the study in the first figure it would not have helped: the values of cyc_y inside its 95% cut, −0.125 to 0.75, carry slopes from 0.650 to 0.980, and the true 0.6 is below all of them. Whether it rescues the 85.5% across studies is not measured.

A residual law estimated rather than assumed. A selection model whose residual law is left free is what the twin construction says cannot identify the mechanism at all, and fitting one would show the identification disappearing as the law is relaxed — the same line, reached from the model’s side rather than the data’s.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formComplete-caseConfidence intervalLikelihood ratio testMaximum likelihoodMissing at randomMissing not at randomModel misspecificationNon-identifiabilityObservation propensityProfile likelihoodSelection modelSensitivity analysisSensitivity parameter