An imputation model the analysis does not contain
Worth reading first: Three mechanisms and one dataset.
Fill the missing outcomes from a model that knows only , then analyse the completed data by regressing on and . The coefficient on converges to 0.26 against a true 0.4 — a bias of −0.14, which is exactly the true coefficient times the missing fraction. The coefficient on , the one the imputation model did carry, converges to 0.642 against a true 0.6, a bias of +0.042, which is exactly .
Both are closed forms out of one two-by-two solve, and the counts land on them: −0.14046 against −0.14 and +0.04020 against +0.042, on standard errors of 0.00235.
The second number is the one worth carrying. An argument that checked only the coefficient the imputer omitted would report uncongeniality as a local defect — a term left out of a model comes out short, which is unsurprising. What actually happens is that the shortfall is redistributed: the coefficient the imputation model did contain absorbs it, in the opposite direction, and comes out of the analysis wrong by a tenth of itself while looking entirely healthy.
Two models, and they are two choices
Multiple imputation is usually described as one procedure, and it contains two independent specifications that nothing forces to agree.
The imputation model says what the missing values are drawn from. The analysis model says what is being estimated from the completed data. In every count of the pooled interval reaching its promise at five draws the two were the same specification, which is the case where the arithmetic has the easiest job and is also the case least likely to arise. In practice the imputer and the analyst are often different people, the imputation is often done once for a dataset that several analyses will use, and the analysis model is often chosen after the completed data arrive.
The mechanism throughout here reads nothing at all — thirty-five per cent of the outcomes go missing completely at random — so dropping the incomplete rows would be correct, and every failure below belongs to the pairing of models rather than to what went missing. That is the same discipline as holding the missing fraction fixed across mechanisms: change one thing and only one.
An imputation done once and an analysis chosen later
The reason this is a failure mode rather than a curiosity is that the two models are usually separated in time and often in authorship. A completed dataset is an artefact: somebody fills the gaps, the filled version is circulated, and analyses are run on it for years afterwards by people who did not choose the imputation model and cannot see it in the file. The completed table has no column recording what its filled values were drawn from.
Every one of those later analyses inherits an attenuation towards whatever the imputer omitted, and none of them can detect it. A completed dataset looks like a dataset: it has no gaps, its residuals behave, and its coefficients have intervals. Four datasets that share a summary are the reminder that a table of numbers does not announce what it cannot support, and a filled table is a stronger version of the same problem, because the numbers were manufactured to be unremarkable.
The practical residue is that an imputation model is a commitment made on behalf of analyses that do not exist yet, and the only safe form of that commitment is to include everything any of them might fit. Which is the standard advice, and is also what makes the reverse case — an imputer richer than the analyst — worth measuring rather than assuming.
The two-by-two solve
Impute from alone and the draws carry the marginal relation between and , which is , correctly. So in the completed population
unchanged. The covariance with is what suffers. The recorded rows carry the full , and the filled rows carry only what inherits through its correlation with , so the completed column reads
short by . Solving the two normal equations with one covariance intact and one short gives, exactly,
Neither is an approximation and neither depends on the sample size. The first says the omitted term is attenuated by precisely the share of rows that were filled — the imputation contributed nothing about beyond what already implied, so a third of the dataset votes for a coefficient of zero. The structure is not destroyed and it is not preserved either; it is diluted, in the way a balancing rule handed both main effects removes exactly none of a pure interaction — what a procedure does not carry, it does not carry, and the arithmetic is indifferent to whether that was intended. The second is the least squares fit doing its job: with ’s coefficient pulled down, the part of the outcome that was explaining has to go somewhere, and it goes into the correlated regressor.
The damage does not stay where the model was short
Coverage separates the two coefficients sharply and not in the way the biases alone would predict. The interval for covers 72.00%. The interval for covers 94.73% ± 0.58 — five hundredths short of its promise, on a bias of 0.042, which is nearly a fifth of the coefficient’s own standard error at this sample size and is therefore invisible in any one study.
That asymmetry is the practically dangerous part. A reader who suspected the imputation model was missing something would inspect the coefficient it was missing, find its interval clearly misbehaving, and could reasonably conclude that the defect is confined there. It is not: the other coefficient is wrong by seven per cent of itself, with an interval that keeps its nominal rate because the bias is small relative to the width rather than because the estimate is right. At four times the sample size the same bias would break that interval too, exactly as a fixed bias breaks a narrowing interval under a mechanism that reads the outcome.
Matching the two models repairs both: the congenial pairing covers 95.80% ± 0.52 on and 95.07% on , with biases of −0.00279 and +0.00120. The machinery is fine. The pairing was not.
The closed forms are checked against the counts on the same fifteen hundred studies rather than argued: −0.14046 against −0.14 and +0.04020 against +0.042, each within two of its own standard errors of 0.00235, and each about a sixtieth of the quantity it is a bias in. Two routes that share no arithmetic — one a solve in the completed population, one an average over simulated studies — is the discipline every number here is held to, and it is what says the compensating movement in is a property of the pairing rather than a feature of one seed.
Rubin’s total variance notices something. The ratio of to the actual sampling variance of the pooled estimate is 1.0769 under the matching pairing and 1.1735 under the short one — the rules widen the interval when the imputations disagree more, which is what they are built to do. It is not enough, and it could not be: the estimate has moved, and a variance is not a location.
The same factor, arriving from a second direction
The factor has now appeared twice in this field, from two different acts, and the coincidence is worth naming because it is not one.
Filling every gap with the observed mean multiplies the fitted slope by , because the filled rows contribute nothing to the cross-product. Filling every gap from a model that omits multiplies ’s coefficient by , because the filled rows contribute nothing about to the cross-product. The second is the first, restricted to one direction of the covariate space: mean imputation is the special case in which the imputation model omits everything, and its attenuation applies to every coefficient at once.
That gives the reading a shape a practitioner can use. An imputation model attenuates, towards zero, exactly the structure it does not contain, by exactly the share it filled. A variable left out of the imputer is not neutral in the analysis; it is actively shrunk, by an amount that is knowable in advance from the missing fraction alone.
There is a tidier way to see why the interval for survives and the interval for does not, and it is arithmetic rather than intuition. Both coefficients have standard errors of the same order at this sample size, and the two biases are 0.042 and 0.140 — so one is a fifth of an error and the other is most of one. Coverage is a function of the ratio of bias to width, and nothing else, which is why a coefficient can be seven per cent wrong with a perfectly healthy interval. It also means the healthy interval is a fact about this sample size and not about this pairing.
The case that is supposed to be unsafe
Turn the pairing round. Let the imputation model know and the analysis fit only the marginal slope of on — so the imputer conditions on more than the analyst does.
The received statement about this case is that Rubin’s variance estimator is biased upward when the imputer is richer than the analyst, so the pooled interval over-covers and the procedure pays for its extra information in width. It is the standard warning attached to the standard advice that an imputation model should be as rich as possible.
It is not there.
At the world’s own settings the estimate is unbiased for the analysis’s own estimand — the marginal slope, 0.72, not the partial 0.6 — with a bias of −0.00164 ± 0.00236, and the interval covers 95.07% ± 0.56 against the matching pairing’s 95.80%. If anything it covers slightly less, not more, and the ratio of Rubin’s total variance to the estimate’s own sits at 1.0530 against the matching pairing’s 1.0769 — below it rather than above, which is the opposite of the direction the claim names.
A single setting is not a measurement of a claim, so the reading is taken again at four strengths of the auxiliary, from a covariate explaining none of the outcome’s variance to one explaining 89.12% of it:
| share of the outcome’s variance explains | estimand | bias | coverage | ÷ the estimate’s own |
|---|---|---|---|---|
| 0.0000 | 0.6000 | −0.00192 | 95.00% ± 0.63 | 1.0167 ± 0.0415 |
| 0.1271 | 0.7200 | −0.00121 | 94.83% ± 0.64 | 1.0164 ± 0.0415 |
| 0.5672 | 0.9600 | +0.00021 | 95.08% ± 0.62 | 0.9815 ± 0.0401 |
| 0.8912 | 1.5000 | +0.00340 | 93.92% ± 0.69 | 0.9360 ± 0.0382 |
The drift across the sweep is −0.0807. The claim predicts a rise; the measurement gives a fall, and the counted coverage stays within 1.08 points of ninety-five throughout. So the honest form of the result is no rise — which is what a claim about an effect that is not there is entitled to say, and no more.
Where the received claim comes from
The claim being refused is not a folk belief; it has a derivation behind it, and being clear about what the derivation says is part of refusing it honestly.
The general treatment of uncongeniality asks what happens when the imputer’s model and the analyst’s procedure imply different answers to the same question, and shows that Rubin’s variance estimator is not guaranteed to be consistent for the repeated-sampling variance of the pooled estimate under either direction of mismatch. The direction where the imputer knows more is then usually described as the benign one, on the grounds that the failure is conservative — the interval is too wide rather than too narrow — which is the version that reaches most descriptions of the method as a warning about paying in width for information.
What is measured above is the second half of that statement, in one world, at four strengths. The first half is not contradicted by anything here: the ratio of to the estimate’s own variance is not 1 at any row of the sweep, which is itself a failure of consistency, and at the strongest setting it is 0.9360 — Rubin’s variance below the truth rather than above it. So the measurement is not “uncongeniality does not matter”. It is that the sign of the mismatch, in the direction everybody calls safe, is the opposite of the one that is quoted, and that the interval nonetheless covers because the departure is small compared with what a t interval at these degrees of freedom is doing anyway.
What could have made the null reading wrong
A measurement that fails to find something is easier to get wrong than one that finds something, so the ways this could be an artefact are worth working through.
The sweep could be too coarse to resolve a small rise. Each ratio carries a standard error of about 0.04, which is what fifteen hundred to twelve hundred studies buy for a ratio of variances. A rise of a couple of per cent would be invisible. A rise large enough to matter for anybody’s interval — five per cent of a variance, two and a half of a width — would be at the edge of resolution at one setting and would show as a trend across four, and the trend that is there runs the other way and is twice the size of a single reading’s error.
The auxiliary could be too weak to trigger it. It is not: the strongest setting has explaining 89.12% of the outcome’s variance, which is far past anything a real auxiliary would do, and it is the setting where the ratio is furthest below one.
The estimand could be doing the work. This is the qualification that matters most. The analysis here fits the marginal slope, so its estimand moves with — 0.6000, 0.7200, 0.9600, 1.5000 across the sweep — and each row is compared against its own target rather than a common one. A sweep that had held the estimand fixed would have been measuring two things at once. What it does mean is that the four rows are four different analyses, and the claim being refused is a claim about a relationship rather than about a level.
The world could be too well behaved. It is a linear-Gaussian world with a correctly specified imputation model and a correctly specified analysis; the auxiliary is fully recorded on every row, and both the imputer and the analyst are right about everything they model. One explanation of the absence is available on that basis and is not tested here: when the auxiliary is complete, the extra information tightens both the between-imputation variance and the actual sampling variance of the pooled estimate, and a ratio of two quantities that fall together need not move. Whether the received claim requires an incomplete auxiliary, a misspecified analysis, or a nonlinear estimand is exactly what this measurement cannot say, and saying so is not the same as saying the claim is wrong.
What is settled, and what the pairing does not touch
None of this is about how many imputations were taken. Twenty is used throughout, which is past the point where the pooled interval is still improving — it reaches its promise at five and is flat by ten — so the failures above cannot be repaired by drawing more. That is worth stating because “take more imputations” is the reflex answer to an under-covering pooled interval, and here it is the wrong one: the interval for would cover 72% at two hundred draws as readily as at twenty, since the estimate it is centred on has moved.
The same point in reverse is why matching the models is not merely good practice. A trial whose analysis does not know how the allocation was made fails in the identical shape — a procedure that is individually correct at both ends and wrong as a pair — and so does including a covariate that the causal structure says should be left out. In each case the arithmetic is not in question and the composition is.
Two limits, stated rather than hedged. Everything above holds the missing fraction at 0.35 and lets the models disagree; nothing sweeps the disagreement against the fraction, though the closed forms say what would happen, since both biases are linear in and vanish with it. And everything above is an imputation model that omits or adds a linear term. An imputation model that gets a functional form wrong, or that omits an interaction the analysis fits, is a different failure with no closed form here, and it is the one most likely to arise from an imputer built for a dataset rather than for an analysis. What the closed forms do establish is that congeniality is not a matter of degree: the attenuation is exactly , whatever the term was worth.
One further thing is untouched and is worth naming because the whole field turns on it. Everything here holds the mechanism at completely random, so the imputation model is fitted to rows that are a plain sample of the population and its only defect is the terms it omits. Under a mechanism that reads the outcome itself the imputation model is fitted to a population that is not the one it extrapolates into, and a correctly specified imputer is then correctly specified for the wrong law. Congeniality and the mechanism are separate conditions, both required, and neither visible in a completed table.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A weight fitted to balance — both name closed form, estimand, model misspecification
- The assumption that identifies the mechanism — both name closed form, confidence interval, model misspecification
- The fit that takes the memory out — both name attenuation, closed form, nuisance parameter
- The moments a balance is told — both name closed form, estimand, model misspecification
- The outcomes a trial could have stopped with — both name closed form, conditional distribution, confidence interval
- What a wrong model estimates — both name closed form, estimand, model misspecification
Named objects
A flat tag is an object no other essay names yet.
AttenuationBetween imputation varianceClosed formConditional distributionConfidence intervalCongenialityEstimandImputation modelMissing completely at randomModel misspecificationMultiple imputationNuisance parameterParameter uncertaintyRubin's rules