The mechanism the data cannot see
Worth reading first: Three mechanisms and one dataset.
Two worlds are constructed. They produce the same covariates, the same pattern of which rows are recorded, and the same recorded outcomes — not close, the same numbers, bit for bit. In one the true slope is 0.6. In the other it is 0.315452.
Every gate the earlier arguments in this field walk through — is the estimator unbiased, does the interval cover, is the imputation model congenial, were enough draws taken — assumes the answer to one question that comes first. Which of those two worlds is this? The data does not answer it, and this is not a statement about power. There is no test, no diagnostic, no likelihood ratio and no sample size, because the two worlds imply the same distribution for everything that can be observed.
What a dataset with gaps actually contains
A study with missing outcomes tells a reader exactly three things, and it is worth listing them because the fourth is what everything above depends on.
It gives the joint law of the covariates, which are recorded on every row. It gives the chance of being recorded given those covariates, which is estimable to any precision the sample size allows, since the indicator is observed. And it gives the law of the outcome given the covariates among the rows where the outcome was recorded, which is what a complete-case regression fits.
What it does not give, at all, is the law of the outcome given the covariates among the rows where it was not recorded. There is nothing there to read. Every estimand anybody wants is a mixture of that fourth law with the third, weighted by the second, so every answer is a function of something the data is silent about. Missing at random is the assumption that the fourth law equals the third. It is an assumption, and it is not a testable one: it is a statement about cells that are empty.
The usual instinct at this point is to look for a diagnostic anyway — to model the chance of being recorded as a function of the outcome and see whether the coefficient is nonzero. That model cannot be fitted. Every row that has an outcome to read is a row where the indicator is one, so the likelihood is separated by construction and there is no estimate to inspect. The diagnostic a reader reaches for is not weak; it does not exist.
It is not that the recorded data are uninformative. They are highly informative — about themselves. Under a rule that reads a covariate the mean of the recorded outcomes is wrong by 0.315192 and under one that reads the outcome it is wrong by 0.617443, and both of those are exactly computable from the recording rule and the recorded rows. What is not computable is which of the two rules produced the table in hand, and every quantity worth reporting depends on that. A subject still event-free when a study ends is the contrasting case and the contrast is instructive: censoring leaves a bound behind, which is information, and an estimator that reads the bound recovers the truth. An unrecorded outcome leaves nothing.
Two worlds, drawn from one stream
The construction fixes all three of the things the data contains and varies the fourth.
Both worlds have the same covariates, the same residuals and the same recording rule — a probit index in with intercept 0.60189 and coefficient 1.2, deleting 35% of the outcomes. In the first world the unrecorded outcomes follow the same law as the recorded ones. In the second they are shifted by . Since the shift applies only to cells that are never read, every recorded number is the same number in both.
The truths differ by a stated amount. Writing for the variance of the index’s covariate part, 1.44 here,
which is exact rather than approximate, because applying to the covariance of the covariates with the index returns the index coefficients themselves. At that puts the second world’s slope at 0.315452, a gap of −0.284548, which is 47.4247% of the first world’s.
Every statistic is the same number
“No test can distinguish them” is a claim about every test rather than about the ones somebody thought of, so it is asserted as equality of numbers rather than as a p-value.
Across forty studies, every recorded outcome is the same number in both worlds, and so is every statistic computed from the recorded data: the fitted slope, the fitted second coefficient, its standard error, the residual sum of squares, the mean and standard deviation of the recorded outcome, the coefficients of a fitted probit for the chance of being recorded and its log-likelihood, the difference in covariate means between recorded and unrecorded rows, and the regression of the recorded outcome on the fitted chance of being recorded. The largest gap across all of them is 0 — not a tolerance, the digit.
That includes the two statistics a reader would actually propose. Comparing the covariates of the recorded rows against the unrecorded ones is a real diagnostic and it detects a real thing; it just detects the wrong thing, since both worlds have identical covariates and identical indicators, and what differs is the outcome in the rows where it was not written down. And regressing the recorded outcome on its own fitted chance of being recorded — the nearest available proxy for “does the outcome predict its own absence” — returns the same coefficient in both, for the same reason.
The complete-case estimate is 0.59955 in both — exactly right in one world, wrong by nearly half in the other, and identical in each. This is the failure whose symptom is absence in its cleanest form: nothing is wrong with what is there, and a resampling method’s inability to see past its own sample is the same limitation reached from a different direction. There the data is a poor guide to a region it barely covers; here it is no guide at all to a region it does not cover.
The shape the fitted line keeps
There is a reason the outcome-driven case is so hard to catch by inspection, and it is a fact about the arithmetic rather than about diagnostics being underused.
When the recording rule reads the outcome and nothing else, every coefficient is multiplied by the same factor — 0.727448 at the field’s own settings, on both the slope on and the slope on , because the bias vector is parallel to the coefficient vector. The fitted relation is scaled rather than tilted. Signs are kept, the ratio between two coefficients is kept, the residuals are well behaved and the fit is a correct fit to the rows in front of it.
So a reader comparing the relative importance of two predictors reaches exactly the right conclusion, and a reader reading a level off the same output is out by more than a quarter. There is no internal inconsistency to notice, because a scaled linear model is a linear model. That is what makes the absence of a diagnostic more than an inconvenience: the failure does not merely evade the tests that exist, it produces output that passes them.
The answer is a line
If a number cannot be given, the honest report is the function that would give it.
| shift in the unseen outcomes | the truth, closed | counted | what the data returns |
|---|---|---|---|
| −1 | 0.88455 | 0.88350 | 0.59955 |
| −0.5 | 0.74227 | 0.74100 | 0.59955 |
| −0.25 | 0.67114 | 0.66974 | 0.59955 |
| 0 | 0.60000 | 0.59849 | 0.59955 |
| +0.25 | 0.52886 | 0.52724 | 0.59955 |
| +0.5 | 0.45773 | 0.45599 | 0.59955 |
| +1 | 0.31545 | 0.31348 | 0.59955 |
The truth moves at −0.284548 per unit of shift, and the counted line has a slope of −0.285011 over two thousand studies at each point. Across the swept range the true slope runs from 0.88455 to 0.31545, a span of 0.56910 — very nearly the value of the slope itself in the world where the shift is zero.
The last column is the point of the table. It does not move. Every row of it is computed from the same recorded cells, so the estimate is not merely similar across the seven worlds; it is the same number seven times.
The slope of that line is not estimated. It is a function of the missingness model and the missing fraction and of nothing else — times the index coefficient — so it can be computed before any data arrive. What cannot be computed, ever, is where on the line to stand.
No sample size touches this
It is worth separating this from the failure that looks like it.
Under a mechanism that reads the outcome, the complete-case interval covers 70.92%, 50.42%, 21.17% and 2.42% at 100, 200, 400 and 800 rows. That is bad and it is a different kind of bad. The bias there is a fixed −0.163531, so a reader who knew the mechanism could correct for it, bound it, or model it — the quantity is estimable given the assumption, and more data estimates it better.
Here there is no quantity. The shift is not a parameter that a larger study pins down more precisely; it is a parameter about which the likelihood is flat, in every sample of every size. Doubling the study halves the width of the interval around 0.59955 and does not move 0.59955, and the interval then excludes the truth with more confidence, in exactly the way an interval narrowing around a fixed bias does — but the reason is one level further back. The estimate is not converging to the wrong number because of a mechanism; it is converging to a number that is right in one of the worlds consistent with what was seen.
What could have made this wrong
The equality could be an artefact of a shared random stream. The two worlds are drawn from one generator in an order that consumes the same numbers in the same places, so it is fair to ask whether the identity is a property of the construction rather than of the statistics. It is a property of both, and the first is the point: the construction exists to make the two datasets the same dataset, which is what a claim of non-identifiability requires. The check that matters is that the statistics are equal as numbers, which is asserted bitwise on forty studies rather than to a tolerance, and a tolerance would have been the weaker claim.
The gap could be small enough to be uninteresting. It is 47.4247% of the slope. At a shift of one standard deviation of the residual — which is not an extreme assumption about people who did not answer, and is exactly the kind of thing a subject-matter argument might propose — the estimand changes by nearly half.
It could be a property of one mechanism. This is the real limit and it is a family rather than a setting. The construction is exact because the index is normal and the truncation identity is closed, which makes the sensitivity relation a straight line with a computable slope. Under a logistic mechanism, a hard threshold rule, or a mechanism depending on the outcome non-monotonically, the relation is presumably still a curve of some kind and nothing here says it is a line or that its slope has a closed form. The non-identifiability is general; the tidy arithmetic is not.
It could be an artefact of the analysis being least squares. Partly. A different analysis has a different sensitivity slope, and the span would change. What does not change is that the recorded data are identical, so whatever quantity is computed, it is computed from cells the two worlds share.
A flat likelihood, and what a prior would do with it
The construction says something specific about every route to an answer, not only about the frequentist ones, and it is worth stating because it is where the reflex repairs stop.
The recorded data have the same likelihood under both worlds. So a likelihood ratio is 1, a maximised likelihood is identical, and any criterion that reads the likelihood — a fit statistic, an information criterion, a Bayes factor — returns the same value in both. There is nothing for a model comparison to compare.
An analysis with a prior on the shift is not exempt either, and the arithmetic is unusually simple: the likelihood contributes nothing about the shift, so the posterior for the shift is the prior, and the posterior for the estimand is that prior pushed through the straight line above. Whatever interval comes out is a restatement of what was assumed, scaled by 0.284548. That is not an argument against putting a prior on it — what a prior is worth is measurable when the data have something to say, and here they have nothing, so the prior is the whole answer and should be presented as one. The failure mode is a prior chosen for convenience and then reported as a result.
What a sensitivity parameter is for
The table of four mechanisms this field opened on is, read against all this, a table of four assumptions rather than four findings. Three of its rows describe worlds where dropping the incomplete rows is exactly right and one describes a world where it costs a quarter of the coefficient, and nothing in a dataset distinguishes them.
So the shift is not a nuisance to be minimised. It is the place where an argument that is being made anyway becomes visible. An analysis that reports 0.59955 with an interval has assumed a shift of zero and has not said so; an analysis that reports the line has made the same computation and declared its terms. The second is not more uncertain than the first — they are the same arithmetic — it is more honest about what the arithmetic rests on, which is the distinction between a number and the context that makes it interpretable applied to an assumption rather than to a sample size.
That reframes every repair this field has priced. Weighting by a chance of being recorded estimated from the same data, pooling twenty imputations, matching the imputation model to the analysis — each is correct under missingness at random and each inherits the assumption whole. The four estimators above differ by a third of a unit on a mean and by a thousandth on a slope, and all four are answers to a question that begins “supposing the unrecorded outcomes follow the same law as the recorded ones”. None of them tests that supposition, and a method that repairs a bias under an assumption is not evidence for the assumption. Pooling several imputations prices the uncertainty of the model that filled the gaps and prices nothing about whether the gaps were filled from the right population; matching the model that fills to the model that analyses is a second condition on top of that one, and both are conditions on top of this.
What is left over
Two things this argument opens and does not close.
The family. The twin construction and the closed slope are properties of a normal index. Whether the sensitivity relation remains a line under a logistic rule or a threshold rule is not measured here, and the number −0.284548 is quoted as a property of one model of missingness rather than of missingness. A reader taking a sensitivity slope from this argument into a different mechanism is taking the idea and not the number.
The other direction of the shift. The sweep is symmetric in and the two halves are not equally plausible in any given study. A shift of −1 puts the true slope at 0.88455, half again what the recorded data returns, and it corresponds to unrecorded outcomes running below the recorded ones at the same covariates. Which sign is credible is a subject-matter question with a different answer in every study, and the arithmetic is indifferent: the line is straight through zero and its two arms are the same length. An analysis that reported only the arm it expected would be back to reporting a point with extra steps.
The range. Nothing here says how far along the line to look. The shift is swept over ±1 because that is a legible range on a residual of standard deviation 1, and there is no principled bound in the arithmetic — a bound would have to come from the subject the data are about, which is exactly the kind of external argument this construction exists to force into the open. Reporting a line without a defensible range is a smaller failure than reporting a point without one, and it is not no failure.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- One imputation is not an observation — both name closed form, complete-case, confidence interval
- One minus Kaplan–Meier is not a risk — both name closed form, estimand, non-identifiability
- Two intervals for one return level — both name closed form, confidence interval, maximum likelihood
- A coverage table with its own error — both name closed form, confidence interval
- A flat point with more than one direction — both name closed form, confidence interval
- A simulation that stops when it looks settled — both name closed form, confidence interval
Named objects
A flat tag is an object no other essay names yet.
Closed formComplete-caseConfidence intervalEstimandMaximum likelihoodMissing at randomMissing not at randomMissingness mechanismNon-identifiabilityObservation propensityObserved data distributionSelection modelSensitivity analysisSensitivity parameter