The value that is not there

A variable that moves recording

A selection model fitted to a skewed outcome that is missing at random reports selection that is not there on three studies in four. Give it a variable that shifts who is recorded and has no place in the outcome — strong enough to carry 29% of the recording index's variance — and the false reports fall to 7.5%, the slope from 1.034 to 0.613 against a truth of 0.6. Leave the same variable out of the model and it does harm instead: false reports reach 100%, and where the outcome really does decide recording, the estimated pull falls from 1.18 to 0.78.

Worth reading first: Three mechanisms and one dataset.

The assumption that identifies the mechanism put a selection model to work on the quantity that two identical datasets had shown no statistic can see: how strongly an outcome decides whether it is recorded. The model estimates that pull by assuming the outcome is normal. Where the assumption holds and the outcome does decide, it repairs a slope that complete cases bias. Where the outcome is missing at random and merely skewed, it does the opposite. It reports selection that is not there on most studies and drags the slope far from the truth that complete cases had right.

That essay closed by naming the obvious repair. Every study it measured had the recording rule and the outcome read the same covariates, so the only thing separating “recorded because the outcome was high” from “recorded because the regressor was high, and the outcome happens to be skewed” was the normal shape assumed for the outcomes nobody saw. A variable that shifts the chance of being recorded and has no place in the outcome equation gives the model something else to identify the pull from. It is the missing-data version of an instrument, and the standard advice for selection models is to have one. The question is how strong it has to be before the skewed studies stop reporting selection, and that can be measured on the same kind of draws.

The studies and the variable

The world is the one this field has used throughout. An outcome depends on a regressor and a second covariate with slopes 0.6 and 0.4, 35% of the outcomes go unrecorded, and the chance of being recorded is a probit index. Two cases are run. In the first, recording depends on the regressor only, so the outcome is missing at random, but the residual is a gamma variable shifted and scaled to mean zero and variance one: right-skewed, and not the normal law the model assumes. In the second, the residual is normal and recording depends on the outcome itself with a pull of 1.2. That is the case the model is built for.

To each, a variable ww is added: standard normal, independent of everything, entering the recording index with coefficient cwc_w and the outcome equation not at all. The intercept is re-solved so the missing share stays at 35% exactly — the measured shares are 34.8% to 35.2% across every setting — so a difference between strengths is not a difference in how much data was lost. At strengths 0.5, 1 and 2 the variable carries 9.3%, 29.1% and 62.1% of the recording index’s variance on the skewed case. A strength of 1 is a variable about as important to recording as the regressor is in a typical study.

Every study of 800 rows is fitted twice. The selection model with ww added to its recording equation is what the advice recommends. The same model without ww is what a study that never measured the variable, or measured it and left it out, would fit. 120 studies are run at each setting.

On the skewed study, the variable takes the false report away

How often a skewed study reports selection that is not there, by the strength of a variable that moves recording only. 120 studies of 800 rows at each strength, missing at random on the regressor with a right-skewed residual, 35% unrecorded. The variable w shifts the chance of being recorded and has no place in the outcome; its share of the recording index's variance is 0.0%, 9.3%, 29.1%, 62.1% at strengths 0, 0.5, 1, 2. The selection model that includes w rejects missingness at random on 75.8%, 27.5%, 7.5%, 5.0% of studies; the same model fitted without w rejects on 78.3%, 85.0%, 100.0%, 100.0%.
Fig. 1 The share of skewed studies, missing at random, on which the selection model rejects missingness at random, by the strength of a variable that moves recording only — with that variable in the model and without it.

With no exclusion at all, the model rejects missingness at random on 75.8% of skewed studies, consistent with the 72.5% the earlier essay measured on its own draws. At a strength of 0.5 the model that uses ww rejects on 27.5%. At 1 it rejects on 7.5%, and at 2 on 5.0%, the test’s nominal size. A variable carrying under a tenth of the recording index’s variance removes most of the false reports, and one carrying under a third removes nearly all of them.

The mechanism is direct. Rows that differ only in ww are recorded at different rates and have the same outcome distribution, because ww is not in the outcome equation. If the outcome were pulling on recording, the recorded outcomes would shift as ww changes who gets recorded: where ww is high and more rows get through, more low outcomes would get through with them. On these studies nothing shifts, and the model can see that. The comparison is across levels of a variable whose effect on recording is estimated from the recording pattern alone, so it does not rest on the shape of the outcomes nobody saw.

The fitted slope of a skewed study missing at random, by the strength of a variable that moves recording only. 120 studies of 800 rows at each strength, missing at random on the regressor with a right-skewed residual, 35% unrecorded. The variable w shifts the chance of being recorded and has no place in the outcome; its share of the recording index's variance is 0.0%, 9.3%, 29.1%, 62.1% at strengths 0, 0.5, 1, 2. Mean slope with w in the model 1.0342, 0.7325, 0.6134, 0.6068; without it 1.0389, 1.0701, 1.1037, 0.9939; the true slope is 0.6.
Fig. 2 The mean fitted slope on the same skewed studies, with the exclusion variable in the model and without it, against the true slope of 0.6.

The slope follows. With no exclusion the selection model’s mean slope is 1.0342, far above the true 0.6 that complete cases get right in this case. With an exclusion of strength 0.5 it is 0.7325, at 1 0.6134, and at 2 0.6068. Coverage of the model’s own interval goes from 13.3% with no exclusion to 64.2%, 95.8% and 90.8%. The last of those is two standard errors of a 120-study count below 95%, so it is a hint rather than a finding. Root mean squared error falls from 0.4813 to 0.0693 at strength 1, roughly a sevenfold improvement.

What the model estimates as the outcome’s pull shows the same transition from the other side.

The estimated pull of the outcome on recording, middle eighty per cent of studies, with and without the exclusion. The selection model's estimate of how strongly the outcome decides whether it is recorded, whose true value here is 0. Median and tenth to ninetieth percentiles over 120 skewed studies at each strength: with the exclusion 2.02 (-0.02 to 2.75), 0.09 (-0.18 to 1.82), 0.03 (-0.19 to 0.21), -0.01 (-0.16 to 0.18); without it 2.03 (-0.02 to 2.68), 2.23 (1.24 to 3.19), 2.73 (2.11 to 3.54), 3.29 (2.65 to 4.40).
Fig. 3 The estimated pull of the outcome on recording, median and tenth to ninetieth percentile over the skewed studies at each strength, with the exclusion in the model and without it. The true pull is zero.

The model with ww puts the median pull at 2.02 with no exclusion, 0.09 at strength 0.5, 0.03 at 1 and −0.01 at 2. At strength 0.5 the middle eighty per cent of studies still runs from −0.18 to 1.82. A weak exclusion moves the typical study to the truth but leaves a long tail of studies where the normal shape is still doing the deciding. At strength 1 the same band is −0.19 to 0.21, and the exclusion has taken over.

Left out, the same variable does harm

The second curve in each figure is the one that matters for studies that never thought about an exclusion. The variable is there and moves recording, but the model does not include it. It does not merely fail to help. It makes the false report worse. The model without ww rejects missingness at random on 85.0% of skewed studies at strength 0.5 and on 100% at strengths 1 and 2. Its median estimated pull rises from 2.03 to 2.73 at strength 1 and 3.29 at 2, and its interval never covers the true slope once at either strength.

The reason is what the model has available to explain recording. Its recording equation says the chance of being recorded, given the covariates, is a probit whose only source of unexplained variation is the standard normal error and, through the pull, the outcome’s own residual. A variable that moves recording and is not in the equation adds variation the model has no other parameter for. The pull of the outcome is the one parameter that can absorb it, so it does, and every unmodelled reason for being recorded becomes evidence that the outcome decided.

The model also grows more confident as it grows more wrong. Its median reported standard error for the slope falls from 0.0757 with no such variable to 0.0658 at strength 1 and 0.0603 at strength 2, while its mean slope sits between 0.99 and 1.10 against a truth of 0.6. A reader shown only the fit would see a sharper estimate and a decisive test against missingness at random, both produced by a variable the analysis never mentions.

Real recording rules are full of such variables. A survey is answered or not partly because of the day the letter arrived, which interviewer called, whether a reminder went out. A clinic records a follow-up partly because of its opening hours and its staffing that month. None of these has any obvious place in the outcome equation, and a study that records them can put them in the recording equation at no cost. A study that does not record them has fitted a selection model whose estimated pull is, in part, a summary of everything about recording it chose not to model.

Where the outcome does decide

The case the selection model is built for, normal residuals and recording driven by the outcome, gives the variable a different job. There is real selection to estimate, and the question is how well.

Where the outcome does decide recording: the slope's error with an exclusion modelled and left out. 120 studies of 800 rows with normal residuals and recording driven by the outcome, true pull 1.2. Root mean squared error of the slope, mean slope, coverage and rejection of missingness at random: no exclusion exists: 0.0657, 0.5947, 90.8%, 90.0%; an exclusion, modelled: 0.0454, 0.5937, 95.8%, 100.0%; an exclusion, left out: 0.1191, 0.5439, 80.8%, 33.3%.
Fig. 4 On normal studies whose outcome decides recording, the root mean squared error of the fitted slope with no exclusion, with an exclusion of strength 1 in the model, and with the same exclusion left out.

With no exclusion, the selection model already does well here: mean slope 0.5947, root mean squared error 0.0657, coverage 90.8%, and it detects the selection on 90.0% of studies. With an exclusion of strength 1 in the model, the error falls to 0.0454, coverage rises to 95.8%, and every one of the 120 studies detects the selection. Its median estimated pull is 1.18 against a true 1.2. The exclusion is now a second source of identification beside the normal shape, and the two agree, so the estimate sharpens.

Left out, the same variable roughly doubles the error, to 0.1191. The mean slope falls to 0.5439, coverage to 80.8%, and selection is detected on only 33.3% of studies. The estimated pull falls to a median of 0.78. That is close to what the omission predicts: an independent standard normal term of strength 1 left out of a probit index scales every coefficient in it by 1/1+cw21/\sqrt{1+c_w^2}, which takes 1.2 to 0.85. The model under-estimates the selection, under-corrects for it, and the slope lands between the complete-case answer and the truth.

So the two cases agree about the omitted variable from opposite directions. On the skewed study where the outcome is innocent, leaving it out inflates the pull. On the normal study where the outcome is guilty, leaving it out deflates it. Either way the selection model’s verdict on the mechanism depends on a variable it was not given, which is the same lesson an imputation model missing a variable the analysis needed taught about filling in: what is left out of the model of the missingness shows up in the estimate of something else.

The exclusion’s strength is the one thing the data can see

There is an asymmetry between the two coefficients in the recording equation that is worth stating plainly. The outcome’s pull on recording is the quantity three mechanisms and one dataset set out to distinguish, and the twin worlds showed the recorded data cannot see it. The exclusion’s pull on recording is a different kind of quantity. Whether a row was recorded is observed for every row, and ww is observed for every row, so how strongly ww moves recording is estimated from complete data. Across the studies here, the model’s median estimate of it is 0.558 at a true 0.5, 1.023 at 1 and 2.048 at 2.

That makes the exclusion’s strength checkable in the way an instrument’s first stage is checkable, and the analogy carries its warning with it. A weak instrument leaves an estimate that drifts back towards the confounded answer it was meant to repair, and the drift is set by how much the instrument moves the treatment. A weak exclusion does the same here. At strength 0.5 the typical study is repaired, but more than a quarter still report selection and the slope’s mean sits at 0.73, a third of the way back towards the shape-identified 1.03. A study can see that its exclusion is weak before it sees anything about the outcome’s pull, and it should treat the selection model’s verdict accordingly.

The checkable part is only the strength, not the exclusion itself. A recording equation can show that ww moves recording; nothing can show that ww does not also move the outcome. That is the next section’s subject, and it is the reason the strength is a necessary condition rather than a sufficient one.

What the exclusion assumes in its turn

The exclusion takes the decision away from the normal law and hands it to a different assumption: that ww has no place in the outcome equation. That assumption is as untestable as the normal law was. The recorded data cannot tell a variable that moves only recording from one that also moves the outcome slightly, because the outcome’s dependence on ww among recorded rows is exactly what selection would also produce. The model has traded a shape assumption for an exclusion assumption, and the twin worlds still exist; they are now worlds that differ in whether ww touches the outcome.

The trade is still worth making, for two reasons visible in the measurements. The exclusion assumption is about a named variable, so it can be argued from subject knowledge in a way a residual’s shape seldom can: a survey’s interviewer, a clinic’s opening hours, a reminder letter’s timing. And the variable is cheap to check partially. If the selection model with ww and the one without it disagree as sharply as they do above, the disagreement is itself a signal that one of the two identifying assumptions is doing all the work. This is the role an instrument plays for a treatment: it does not make an untestable assumption go away, it moves it somewhere it can be defended.

Complete-case least squares needs none of this when the outcome is missing at random, as dropping the incomplete rows showed: on these skewed studies it is unbiased, with or without ww, because ww is independent of the outcome given the covariates. The selection model earns its place only where the outcome decides, and there an exclusion, if one exists, should be in it.

What a selection-model analysis should report

Whether an exclusion variable was used, and why it is believed to move recording only. Without one, the estimated pull is identified by the normal law alone, and on a skewed outcome that is enough to report selection on three studies in four that have none.

The variable’s strength in the recording equation. Here, a variable carrying a tenth of the recording index’s variance removes most false reports and one carrying under a third removes nearly all. A weak exclusion leaves a long tail of studies still decided by the shape.

The fit without the exclusion beside the fit with it. If they agree, the normal shape and the exclusion are telling the same story. If they disagree — a pull of 2.7 against 0.03, as on the skewed studies — at least one of the two identifying assumptions is false, and the report should say which one it is relying on.

Every variable known to move recording, in the recording equation. Leaving one out is not neutral. It inflates the estimated pull where the outcome is innocent and deflates it where the outcome is guilty.

What is claimed and what is not

Counted. Every rate, slope, coverage and pull on 120 studies of 800 rows at each setting, with both fits made on every study so the difference between them is the variable and nothing else. The missing share is solved to 35% at every strength and measured at 34.8% to 35.2%.

Checked. The selection likelihood’s gradient in the new coefficient agrees with its central difference, as the other eight already did.

Not claimed. One skewed law, one outcome-driven strength and one sample size. The exclusion here is standard normal and independent of the covariates. An exclusion correlated with the regressor carries less independent information about recording and would need to be stronger to do the same work. At 120 studies a setting a rate is good to about four points, which is why the 90.8% at strength 2 is described as a hint.

Still open: an exclusion that is slightly wrong

The measurement above takes the exclusion at its word: ww never touches the outcome. A real candidate — an interviewer, a reminder, a clinic’s distance — will touch it a little. A variable that moves recording strongly and the outcome weakly is still informative, but the model now attributes the outcome’s dependence on ww to selection. How large that leak can be before the exclusion’s repair of the skewed studies turns into a manufactured pull of its own, and whether the leak and the skew can cancel or compound, is a measurement on the same draws. It is the measurement that would decide whether an exclusion is a repair or a second assumption of the same kind, and it has not been made.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Complete-caseLikelihood ratio testMaximum likelihoodMissing at randomMissing not at randomMissingness mechanismModel misspecificationMonte CarloNon-identifiabilityObservation propensitySelection model