Concept

Sensitivity analysis — where it appears

Recomputing a result under a range of values for an assumption the data cannot check, to show how far the answer moves. It does not test the assumption; it reports what each value of it implies, and the range tried has to be argued for from how the data arose.

Named by 4 essays across 3 fields — each of them below, with the objects they name alongside it.

Two worlds, one Kaplan–Meier curve, two truths. World A gives each subject a frailty with mean one and variance 1, and multiplies both its event hazard (0.35) and its dropout hazard (0.5) by it, so the subjects likeliest to leave are the ones likeliest to fail. World B has independent event and dropout times whose hazards are world A's crude hazards. Kaplan–Meier over 1000 studies of 400 gives the same curve from both — 0.7750 and 0.7763 at t = 1; 0.6627 and 0.6641 at t = 2; 0.5924 and 0.5932 at t = 3; 0.5047 and 0.5042 at t = 5 — and that curve is world B's truth, 0.5052 at t = 5. World A's truth is 0.3636 there. The dashed lines are the two bounds that assume nothing, from every dropout failing on leaving (0.1905 at t = 5) to none ever failing (0.6667).

A dropout the data cannot see

Two worlds leave the same record to the last detail a study can write down — the same times, the same share ending in the event, the same share leaving first — and a log-rank test between them rejects at its own 5% level at every sample size from a hundred to sixteen hundred. Kaplan–Meier converges on 0.5052 at t = 5 from both. The truth is 0.5052 in one and 0.3636 in the other, and what is left to argue about is where between two bounds to stand.

survival · Censoring
An estimate reported as a function of an assumption. What the slope really is, against a shift in the outcomes nobody saw — line from the closed form, dots counted over 2000 studies of 200 rows at 35.0% missing. Every point on this line produces exactly the same observed data, and the complete-case estimate is the flat line at 0.5996 regardless. The truth moves at -0.2845 per unit of shift, which is a function of the missingness model and the missing fraction and of nothing that can be estimated: across the swept range the true slope runs from 0.8845 to 0.3155, a span of 0.5691 against a value of 0.60 in the world where the shift is zero. Reporting the line is the honest form of the answer.

The mechanism the data cannot see

Two worlds produce identical covariates, identical patterns of what is recorded and identical recorded outcomes, to the last bit. Their true slopes are 0.6 and 0.315452, and the truth moves at 0.284548 per unit of an assumption nothing in the data can inform.

missing · Missingness
A peak where the recorded data have none. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 0.35, where the fitted slope is 0.839, and the values of cy within the 95% cut run from −0.13 to 0.75; the likelihood-ratio statistic against cy = 0 is 1.47. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.

The assumption that identifies the mechanism

A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.

missing · Missingness
A proxy removes less than its reliability, always. The share of the confounding bias removed by adjusting for a proxy, against how well the proxy measures the confounder. The diagonal is the answer a reader would guess — a covariate that is 80% signal removes 80% of the problem. The curve is what the arithmetic gives: the reliability, times one minus the squared correlation between the treatment and the confounder, divided by one minus the product of those two. That squared correlation is 0.4475. A reliability of 0.8 removes 68.85% and one of 0.6 removes 45.32%. The two agree only at the ends, and the gap is widest where most applied covariates sit.

Adjusting for a shadow

A covariate that is 80% signal removes 68.85% of the confounding, not 80% — the share is λ(1 − ρ²)/(1 − λρ²) and it is below the reliability everywhere. The residual bias is 0.1084 against an effect of 0.5, and at 25,600 rows it is 17.6 standard errors wide.

collider · Conditioning

Named alongside it

The objects these essays reach for when they reach for this one.

Closed formNon-identifiabilitySensitivity parameterComplete-caseConfidence intervalEstimandMaximum likelihoodMissing at randomMissing not at randomObservation propensitySelection modelAdjustment set

All concepts