Concept

Non-identifiability — where it appears

The condition where two different states of the world imply exactly the same distribution for everything that can be observed. No amount of data separates them, so the choice between them is made by assumption, and the size of the gap between their answers is what a sensitivity analysis reports.

Named by 7 essays across 4 fields — each of them below, with the objects they name alongside it.

The first stage an instrument needs is set by the violation nobody can see. The error each estimator converges on when the instrument has a direct effect of 0.05 on the outcome — a path the exclusion restriction asserts is zero and no sample can check. The instrument's error is δ/π exactly, so it is the reciprocal of the very quantity that made the method work: 1.0000 at a first stage of 0.05 and 0.0833 at 0.60. Least squares carries the confounding instead, at 0.3440 at a first stage of 0.30. The two cross at π = 0.1389, and the crossing is exactly δ times 2.7778 — the first stage an instrument needs is proportional to the violation it is assumed not to have, and below that line the method being corrected is the better estimator.

The assumption nothing tests

An instrument buys a causal effect with an assumption no sample can check, and the price is set by the same quantity that made the method work. The first stage it needs is 2.7778 times the violation it is assumed not to have, so a direct effect of 0.05 demands a first stage of 0.1389 and least squares wins below it.

instrument · Exclusion
Three mechanisms leave the slope alone; one does not. The bias of the complete-case slope under each of four missingness rules, counted over 4000 studies of 200 rows at 35.0% missing, with the closed form printed beside each count. Missingness that depends on nothing, on the regressor, or on the second covariate leaves the slope exactly where it was — the closed forms are zero to machine precision and the counts are -0.0005, -0.0005 and -0.0011 against standard errors of about 0.0018. Missingness that depends on the outcome moves it by -0.1635, which is 27.3% of the slope being estimated. The same share of rows is lost in every case.

Three mechanisms and one dataset

Four rules for which outcomes go missing, each calibrated to lose the same 35% of the rows and each leaning on what it reads with the same coefficient. Three leave the fitted slope exactly where it was, and the one that reads the outcome moves it by 0.163531.

missing · Missingness
Two worlds, one Kaplan–Meier curve, two truths. World A gives each subject a frailty with mean one and variance 1, and multiplies both its event hazard (0.35) and its dropout hazard (0.5) by it, so the subjects likeliest to leave are the ones likeliest to fail. World B has independent event and dropout times whose hazards are world A's crude hazards. Kaplan–Meier over 1000 studies of 400 gives the same curve from both — 0.7750 and 0.7763 at t = 1; 0.6627 and 0.6641 at t = 2; 0.5924 and 0.5932 at t = 3; 0.5047 and 0.5042 at t = 5 — and that curve is world B's truth, 0.5052 at t = 5. World A's truth is 0.3636 there. The dashed lines are the two bounds that assume nothing, from every dropout failing on leaving (0.1905 at t = 5) to none ever failing (0.6667).

A dropout the data cannot see

Two worlds leave the same record to the last detail a study can write down — the same times, the same share ending in the event, the same share leaving first — and a log-rank test between them rejects at its own 5% level at every sample size from a hundred to sixteen hundred. Kaplan–Meier converges on 0.5052 at t = 5 from both. The truth is 0.5052 in one and 0.3636 in the other, and what is left to argue about is where between two bounds to stand.

survival · Censoring
The risk of one cause, estimated two ways. Two causes of an ending event with constant hazards 0.2 (the one of interest) and 0.3 (the competitor), random dropout at 0.1 and follow-up to 6. The lower line is the cumulative incidence, (0.2/0.5)(1 − e^(−0.5t)), the chance of actually having had this event by t; the dots on it are the Aalen–Johansen estimate over 2000 studies of 300, 0.3670 at t = 5 against 0.3672. The upper line is 1 − e^(−0.2t), and the dots on it are one minus Kaplan–Meier with the competing event treated as censoring: 0.6318 at t = 5 against 0.6321. The second is larger by a factor of 1.722 at t = 5, and it is not an error of estimation. It estimates, correctly, the risk in a population where the competing cause does not exist.

One minus Kaplan–Meier is not a risk

With two ways for observation to end, one minus Kaplan–Meier for one cause reads 0.6318 at t = 5 where the chance of actually having had that event is 0.3670. Added across the two causes, the complements reach 1.4088 — more than the whole cohort. Nothing is estimated badly: the complement estimates, correctly, the risk in a world where the other cause does not exist.

survival · Censoring
An estimate reported as a function of an assumption. What the slope really is, against a shift in the outcomes nobody saw — line from the closed form, dots counted over 2000 studies of 200 rows at 35.0% missing. Every point on this line produces exactly the same observed data, and the complete-case estimate is the flat line at 0.5996 regardless. The truth moves at -0.2845 per unit of shift, which is a function of the missingness model and the missing fraction and of nothing that can be estimated: across the swept range the true slope runs from 0.8845 to 0.3155, a span of 0.5691 against a value of 0.60 in the world where the shift is zero. Reporting the line is the honest form of the answer.

The mechanism the data cannot see

Two worlds produce identical covariates, identical patterns of what is recorded and identical recorded outcomes, to the last bit. Their true slopes are 0.6 and 0.315452, and the truth moves at 0.284548 per unit of an assumption nothing in the data can inform.

missing · Missingness
A peak where the recorded data have none. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 0.35, where the fitted slope is 0.839, and the values of cy within the 95% cut run from −0.13 to 0.75; the likelihood-ratio statistic against cy = 0 is 1.47. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.

The assumption that identifies the mechanism

A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.

missing · Missingness
One of these converges and the other does not. Two measurements on the same fits, against the sample length, for a system with 2 genuine relations. The distance from the fitted plane to the true plane falls from 0.1438 at 100 observations to 0.0075 at 1600 — halving with each doubling, which is the 1/n rate this field's estimates converge at. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and is flat in between. The plane is an estimate; the relation inside it is not.

A space is not a relation

The fitted plane approaches the true one at rate 1/n — 0.1438 at a hundred observations and 0.0075 at sixteen hundred. The angle between the leading fitted relation and the leading generating one reads 29.6° and 29.0° at those same lengths, and never moves.

systems · Rank

Named alongside it

The objects these essays reach for when they reach for this one.

Closed formEstimandComplete-caseConfidence intervalMissing at randomMissing not at randomSelection modelSensitivity analysisSensitivity parameterCause specific hazardCensoringCompeting risks

All concepts