The value that is not there

An imputation model the analysis does not contain

A model that fills the gaps without a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing share, 0.4 to 0.26, and moves the one it did carry by exactly γρf, 0.6 to 0.642. The reverse case is supposed to inflate the interval, and at four strengths of the extra knowledge it does not.

Worth reading first: Three mechanisms and one dataset.

Fill the missing outcomes from a model that knows only xx, then analyse the completed data by regressing on xx and zz. The coefficient on zz converges to 0.26 against a true 0.4 — a bias of −0.14, which is exactly the true coefficient times the missing fraction. The coefficient on xx, the one the imputation model did carry, converges to 0.642 against a true 0.6, a bias of +0.042, which is exactly γρf\gamma\rho f.

The damage does not stay in the term that was left out. Where each coefficient lands when the model that fills the missing outcomes and the model that analyses them disagree, over 1500 studies of 200 rows at 35.0% missing and 20 imputations. An imputer that omits a covariate the analysis fits attenuates that covariate's coefficient by exactly the missing fraction — -0.1405 counted against a closed -0.1400 — and pushes the coefficient it did impute on the other way by exactly the product of the omitted coefficient, the covariates' correlation and the missing fraction: 0.0402 counted against 0.0420. Both closed forms come out of the same two-by-two solve. Matching models leave both alone, and so does an imputer that knows more than the analysis.
Fig. 1 Where each coefficient lands when the model that fills and the model that analyses disagree, over fifteen hundred studies at twenty imputations each, with the closed form printed beside each count.

Both are closed forms out of one two-by-two solve, and the counts land on them: −0.14046 against −0.14 and +0.04020 against +0.042, on standard errors of 0.00235.

The second number is the one worth carrying. An argument that checked only the coefficient the imputer omitted would report uncongeniality as a local defect — a term left out of a model comes out short, which is unsurprising. What actually happens is that the shortfall is redistributed: the coefficient the imputation model did contain absorbs it, in the opposite direction, and comes out of the analysis wrong by a tenth of itself while looking entirely healthy.

Two models, and they are two choices

Multiple imputation is usually described as one procedure, and it contains two independent specifications that nothing forces to agree.

The imputation model says what the missing values are drawn from. The analysis model says what is being estimated from the completed data. In every count of the pooled interval reaching its promise at five draws the two were the same specification, which is the case where the arithmetic has the easiest job and is also the case least likely to arise. In practice the imputer and the analyst are often different people, the imputation is often done once for a dataset that several analyses will use, and the analysis model is often chosen after the completed data arrive.

The mechanism throughout here reads nothing at all — thirty-five per cent of the outcomes go missing completely at random — so dropping the incomplete rows would be correct, and every failure below belongs to the pairing of models rather than to what went missing. That is the same discipline as holding the missing fraction fixed across mechanisms: change one thing and only one.

An imputation done once and an analysis chosen later

The reason this is a failure mode rather than a curiosity is that the two models are usually separated in time and often in authorship. A completed dataset is an artefact: somebody fills the gaps, the filled version is circulated, and analyses are run on it for years afterwards by people who did not choose the imputation model and cannot see it in the file. The completed table has no column recording what its filled values were drawn from.

Every one of those later analyses inherits an attenuation towards whatever the imputer omitted, and none of them can detect it. A completed dataset looks like a dataset: it has no gaps, its residuals behave, and its coefficients have intervals. Four datasets that share a summary are the reminder that a table of numbers does not announce what it cannot support, and a filled table is a stronger version of the same problem, because the numbers were manufactured to be unremarkable.

The practical residue is that an imputation model is a commitment made on behalf of analyses that do not exist yet, and the only safe form of that commitment is to include everything any of them might fit. Which is the standard advice, and is also what makes the reverse case — an imputer richer than the analyst — worth measuring rather than assuming.

The two-by-two solve

Impute yy from xx alone and the draws carry the marginal relation between xx and yy, which is β+γρ\beta + \gamma\rho, correctly. So in the completed population

cov(x,yfilled)=cov(x,y),\operatorname{cov}(x, y_{\text{filled}}) = \operatorname{cov}(x, y),

unchanged. The covariance with zz is what suffers. The recorded rows carry the full cov(z,y)=ρβ+γ\operatorname{cov}(z, y) = \rho\beta + \gamma, and the filled rows carry only what zz inherits through its correlation with xx, so the completed column reads

cov(z,yfilled)=ρβ+γκ,κ=1f(1ρ2)=0.6815,\operatorname{cov}(z, y_{\text{filled}}) = \rho\beta + \gamma\kappa, \qquad \kappa = 1 - f(1 - \rho^2) = \mathbf{0.6815},

short by γf(1ρ2)\gamma f(1-\rho^2). Solving the two normal equations with one covariance intact and one short gives, exactly,

γ^γ(1f)=0.26,β^β+γρf=0.642.\hat\gamma \to \gamma(1-f) = 0.26, \qquad \hat\beta \to \beta + \gamma\rho f = 0.642.

Neither is an approximation and neither depends on the sample size. The first says the omitted term is attenuated by precisely the share of rows that were filled — the imputation contributed nothing about zz beyond what xx already implied, so a third of the dataset votes for a coefficient of zero. The structure is not destroyed and it is not preserved either; it is diluted, in the way a balancing rule handed both main effects removes exactly none of a pure interaction — what a procedure does not carry, it does not carry, and the arithmetic is indifferent to whether that was intended. The second is the least squares fit doing its job: with zz’s coefficient pulled down, the part of the outcome that zz was explaining has to go somewhere, and it goes into the correlated regressor.

The damage does not stay where the model was short

Where an uncongenial pairing loses its promise. What each pooled 95% interval covers, over 1500 studies of 200 rows at 35.0% missing and 20 imputations apiece. When the imputation model and the analysis model agree, both coefficients cover their promise — 95.80% and 95.07%. When the imputer leaves out a covariate the analysis fits, the interval for the coefficient it left out covers 72.00% and the interval for the one it kept covers 94.73% — the damage is largest where the model was short and does not stay there. When the imputer knows a covariate the analysis omits, the interval covers 95.07%, which is its promise.
Fig. 2 What each pooled interval covers under three pairings of imputation and analysis model. The interval for the coefficient the imputer omitted covers 72.00%; the interval for the one it kept covers 94.73%.

Coverage separates the two coefficients sharply and not in the way the biases alone would predict. The interval for γ\gamma covers 72.00%. The interval for β\beta covers 94.73% ± 0.58 — five hundredths short of its promise, on a bias of 0.042, which is nearly a fifth of the coefficient’s own standard error at this sample size and is therefore invisible in any one study.

That asymmetry is the practically dangerous part. A reader who suspected the imputation model was missing something would inspect the coefficient it was missing, find its interval clearly misbehaving, and could reasonably conclude that the defect is confined there. It is not: the other coefficient is wrong by seven per cent of itself, with an interval that keeps its nominal rate because the bias is small relative to the width rather than because the estimate is right. At four times the sample size the same bias would break that interval too, exactly as a fixed bias breaks a narrowing interval under a mechanism that reads the outcome.

Matching the two models repairs both: the congenial pairing covers 95.80% ± 0.52 on β\beta and 95.07% on γ\gamma, with biases of −0.00279 and +0.00120. The machinery is fine. The pairing was not.

The closed forms are checked against the counts on the same fifteen hundred studies rather than argued: −0.14046 against −0.14 and +0.04020 against +0.042, each within two of its own standard errors of 0.00235, and each about a sixtieth of the quantity it is a bias in. Two routes that share no arithmetic — one a solve in the completed population, one an average over simulated studies — is the discipline every number here is held to, and it is what says the compensating movement in β\beta is a property of the pairing rather than a feature of one seed.

Rubin’s total variance notices something. The ratio of TT to the actual sampling variance of the pooled estimate is 1.0769 under the matching pairing and 1.1735 under the short one — the rules widen the interval when the imputations disagree more, which is what they are built to do. It is not enough, and it could not be: the estimate has moved, and a variance is not a location.

The same factor, arriving from a second direction

One act, two different factors. What survives mean imputation, as a share of the true quantity, against the share of outcomes filled — lines from the closed forms, dots counted over 1200 studies of 200 rows each. Replacing a fraction of a column by its own observed mean multiplies that column's variance by exactly one minus the fraction and multiplies its correlation with the regressor by the square root of the same number, because the correlation carries the filled column's standard deviation in its denominator and the covariance does not. At 35.0% filled the variance keeps 0.6500 of itself, the slope the same 0.6500, and the correlation 0.8062. The counts at that fraction are 0.5986 and 0.7764 at the nearest swept point.
Fig. 3 What survives filling a share of a column with one number, against the share filled. The variance and the slope keep 1 − f of themselves and the correlation keeps its square root.

The factor 1f1 - f has now appeared twice in this field, from two different acts, and the coincidence is worth naming because it is not one.

Filling every gap with the observed mean multiplies the fitted slope by 1f1 - f, because the filled rows contribute nothing to the cross-product. Filling every gap from a model that omits zz multiplies zz’s coefficient by 1f1 - f, because the filled rows contribute nothing about zz to the cross-product. The second is the first, restricted to one direction of the covariate space: mean imputation is the special case in which the imputation model omits everything, and its attenuation applies to every coefficient at once.

That gives the reading a shape a practitioner can use. An imputation model attenuates, towards zero, exactly the structure it does not contain, by exactly the share it filled. A variable left out of the imputer is not neutral in the analysis; it is actively shrunk, by an amount that is knowable in advance from the missing fraction alone.

There is a tidier way to see why the interval for β\beta survives and the interval for γ\gamma does not, and it is arithmetic rather than intuition. Both coefficients have standard errors of the same order at this sample size, and the two biases are 0.042 and 0.140 — so one is a fifth of an error and the other is most of one. Coverage is a function of the ratio of bias to width, and nothing else, which is why a coefficient can be seven per cent wrong with a perfectly healthy interval. It also means the healthy interval is a fact about this sample size and not about this pairing.

The case that is supposed to be unsafe

Turn the pairing round. Let the imputation model know zz and the analysis fit only the marginal slope of yy on xx — so the imputer conditions on more than the analyst does.

The received statement about this case is that Rubin’s variance estimator is biased upward when the imputer is richer than the analyst, so the pooled interval over-covers and the procedure pays for its extra information in width. It is the standard warning attached to the standard advice that an imputation model should be as rich as possible.

It is not there.

The safe case is safe, which was not the guess. Rubin's total variance over the actual sampling variance of the pooled estimate, when the imputation model conditions on a covariate the analysis leaves out, at four strengths of that covariate — counted over 1200 studies of 200 rows apiece at 35.0% missing, with each reading's own standard error drawn as a bar. The received statement is that an imputer richer than the analyst inflates this ratio and so over-covers. It reads 1.0167, 1.0164, 0.9815, 0.9360 as the extra covariate goes from explaining none of the outcome's variance to 89.1% of it, and the counted coverage stays between 93.92% and 95.08%. The drift across the sweep is -0.0807, and it runs the wrong way for the claim.
Fig. 4 Rubin’s total variance over the actual sampling variance of the pooled estimate, at four strengths of a covariate the imputer knows and the analysis omits. The received claim says this should rise; the drift across the sweep is −0.0807.

At the world’s own settings the estimate is unbiased for the analysis’s own estimand — the marginal slope, 0.72, not the partial 0.6 — with a bias of −0.00164 ± 0.00236, and the interval covers 95.07% ± 0.56 against the matching pairing’s 95.80%. If anything it covers slightly less, not more, and the ratio of Rubin’s total variance to the estimate’s own sits at 1.0530 against the matching pairing’s 1.0769 — below it rather than above, which is the opposite of the direction the claim names.

A single setting is not a measurement of a claim, so the reading is taken again at four strengths of the auxiliary, from a covariate explaining none of the outcome’s variance to one explaining 89.12% of it:

share of the outcome’s variance zz explains estimand bias coverage TT ÷ the estimate’s own
0.0000 0.6000 −0.00192 95.00% ± 0.63 1.0167 ± 0.0415
0.1271 0.7200 −0.00121 94.83% ± 0.64 1.0164 ± 0.0415
0.5672 0.9600 +0.00021 95.08% ± 0.62 0.9815 ± 0.0401
0.8912 1.5000 +0.00340 93.92% ± 0.69 0.9360 ± 0.0382

The drift across the sweep is −0.0807. The claim predicts a rise; the measurement gives a fall, and the counted coverage stays within 1.08 points of ninety-five throughout. So the honest form of the result is no rise — which is what a claim about an effect that is not there is entitled to say, and no more.

Where the received claim comes from

The claim being refused is not a folk belief; it has a derivation behind it, and being clear about what the derivation says is part of refusing it honestly.

The general treatment of uncongeniality asks what happens when the imputer’s model and the analyst’s procedure imply different answers to the same question, and shows that Rubin’s variance estimator is not guaranteed to be consistent for the repeated-sampling variance of the pooled estimate under either direction of mismatch. The direction where the imputer knows more is then usually described as the benign one, on the grounds that the failure is conservative — the interval is too wide rather than too narrow — which is the version that reaches most descriptions of the method as a warning about paying in width for information.

What is measured above is the second half of that statement, in one world, at four strengths. The first half is not contradicted by anything here: the ratio of TT to the estimate’s own variance is not 1 at any row of the sweep, which is itself a failure of consistency, and at the strongest setting it is 0.9360 — Rubin’s variance below the truth rather than above it. So the measurement is not “uncongeniality does not matter”. It is that the sign of the mismatch, in the direction everybody calls safe, is the opposite of the one that is quoted, and that the interval nonetheless covers because the departure is small compared with what a t interval at these degrees of freedom is doing anyway.

What could have made the null reading wrong

A measurement that fails to find something is easier to get wrong than one that finds something, so the ways this could be an artefact are worth working through.

The sweep could be too coarse to resolve a small rise. Each ratio carries a standard error of about 0.04, which is what fifteen hundred to twelve hundred studies buy for a ratio of variances. A rise of a couple of per cent would be invisible. A rise large enough to matter for anybody’s interval — five per cent of a variance, two and a half of a width — would be at the edge of resolution at one setting and would show as a trend across four, and the trend that is there runs the other way and is twice the size of a single reading’s error.

The auxiliary could be too weak to trigger it. It is not: the strongest setting has zz explaining 89.12% of the outcome’s variance, which is far past anything a real auxiliary would do, and it is the setting where the ratio is furthest below one.

The estimand could be doing the work. This is the qualification that matters most. The analysis here fits the marginal slope, so its estimand moves with γ\gamma — 0.6000, 0.7200, 0.9600, 1.5000 across the sweep — and each row is compared against its own target rather than a common one. A sweep that had held the estimand fixed would have been measuring two things at once. What it does mean is that the four rows are four different analyses, and the claim being refused is a claim about a relationship rather than about a level.

The world could be too well behaved. It is a linear-Gaussian world with a correctly specified imputation model and a correctly specified analysis; the auxiliary is fully recorded on every row, and both the imputer and the analyst are right about everything they model. One explanation of the absence is available on that basis and is not tested here: when the auxiliary is complete, the extra information tightens both the between-imputation variance and the actual sampling variance of the pooled estimate, and a ratio of two quantities that fall together need not move. Whether the received claim requires an incomplete auxiliary, a misspecified analysis, or a nonlinear estimand is exactly what this measurement cannot say, and saying so is not the same as saying the claim is wrong.

What is settled, and what the pairing does not touch

The extra 1/m, and the correction nobody quotes. What a pooled 95% interval covers against the number of imputations, counted over 2000 studies of 200 rows at 35.0% of outcomes missing. Rubin's rules — total variance W̄ + (1 + 1/m)B, read against a t distribution on (m − 1)(1 + W̄/((1 + 1/m)B))² degrees of freedom — cover 94.10% at two imputations and reach their promise by 5, at 95.25%. Dropping the (1 + 1/m) factor takes two imputations to 93.10%; using a normal quantile instead of the degrees-of-freedom correction takes it to 92.55%; dropping both takes it to 91.45%. The median degrees of freedom at two imputations is 12.95, which is why the second correction is the larger.
Fig. 5 The pooled interval against the number of imputations under a matching pair of models. Every count here uses twenty, which is far past the point where the number of draws is what limits anything.

None of this is about how many imputations were taken. Twenty is used throughout, which is past the point where the pooled interval is still improving — it reaches its promise at five and is flat by ten — so the failures above cannot be repaired by drawing more. That is worth stating because “take more imputations” is the reflex answer to an under-covering pooled interval, and here it is the wrong one: the interval for γ\gamma would cover 72% at two hundred draws as readily as at twenty, since the estimate it is centred on has moved.

The same point in reverse is why matching the models is not merely good practice. A trial whose analysis does not know how the allocation was made fails in the identical shape — a procedure that is individually correct at both ends and wrong as a pair — and so does including a covariate that the causal structure says should be left out. In each case the arithmetic is not in question and the composition is.

The estimand decides whether a repair is needed. The bias of four estimators of the outcome's population mean, when 35.0% of outcomes are missing at a rate that depends on the regressor, over 1500 studies of 200 rows. The mean of the complete cases is wrong by 0.3143 against a closed prediction of 0.3152, on the same rows whose fitted slope is exactly right. Weighting each complete case by the reciprocal of its chance of being observed removes most of it: 0.0183 ± 0.0051 with the true chances and 0.0229 ± 0.0044 with fitted ones — small, and several standard errors from zero, which is what consistent rather than unbiased looks like at this size. Pooling 20 imputations leaves -0.0009 ± 0.0030.
Fig. 6 Four estimators of the outcome’s population mean under a mechanism that reads the regressor. Twenty imputations leave −0.00089 where the complete cases are wrong by 0.31427.

Two limits, stated rather than hedged. Everything above holds the missing fraction at 0.35 and lets the models disagree; nothing sweeps the disagreement against the fraction, though the closed forms say what would happen, since both biases are linear in ff and vanish with it. And everything above is an imputation model that omits or adds a linear term. An imputation model that gets a functional form wrong, or that omits an interaction the analysis fits, is a different failure with no closed form here, and it is the one most likely to arise from an imputer built for a dataset rather than for an analysis. What the closed forms do establish is that congeniality is not a matter of degree: the attenuation is exactly 1f1 - f, whatever the term was worth.

One further thing is untouched and is worth naming because the whole field turns on it. Everything here holds the mechanism at completely random, so the imputation model is fitted to rows that are a plain sample of the population and its only defect is the terms it omits. Under a mechanism that reads the outcome itself the imputation model is fitted to a population that is not the one it extrapolates into, and a correctly specified imputer is then correctly specified for the wrong law. Congeniality and the mechanism are separate conditions, both required, and neither visible in a completed table.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AttenuationBetween imputation varianceClosed formConditional distributionConfidence intervalCongenialityEstimandImputation modelMissing completely at randomModel misspecificationMultiple imputationNuisance parameterParameter uncertaintyRubin's rules