Field

When the data stops early

A subject still event-free when a study ends is not missing and not observed — it is known to exceed something, which is a third state most tools have no slot for. The estimator that gives it that slot recovers the curve, and everything read off the curve afterwards has a condition attached: the interval printed around its end stops covering where the curve is read, a dropout that carries information leaves a record identical to one that does not, one minus the curve is not a risk once something else can end observation first, and a hazard ratio is an average whose weights the length of follow-up chose.
The same study read three ways. At time 2 the truth is 0.497. Kaplan–Meier gives 0.532; dropping the censored subjects gives 0.180; treating the censoring time as the event time gives 0.392. Both naive readings understate survival, because the subjects they mishandle are the ones doing well.

The data that stops early

A subject still event-free when a study ends is not missing and not observed. It is known to exceed something, which is a third state most tools have no slot for — and the two obvious ways of forcing it into one are wrong by 31 and 13 percentage points.

Kaplan–Meier from 120 subjects, 59 of them censored. The step curve is the estimate, the smooth curve is the truth it is trying to recover. 59 of 120 subjects were still event-free when observation stopped; they are not dropped, and they are not counted as events — they leave the risk set at the time they were last seen.

The curve that survives censoring

Kaplan–Meier recovers the true survival curve to within a fraction of a point at every censoring level from 37% to 71%, where dropping the censored subjects is off by 28 and then by 43. The estimator is a running product and the reason it works is in its denominator.

One cohort of 40, and two intervals around the end of its curve. A single simulated study of 40 subjects with exponential survival at rate 0.35, dropout at rate 0.15 and follow-up to 6 — the first seed from 8811 upward whose plain band reaches below −0.05, chosen to show the failure rather than its frequency. The step curve is Kaplan–Meier and the smooth curve the truth. The plain band, the estimate plus and minus 1.96 Greenwood standard errors, first dips below zero at t = 3.78 and reaches −0.052; early on it also rises to 1.023, above one. At t = 5 the estimate is 0.069 with 1 subject still under observation, the plain interval runs from −0.052 to 0.191 and the log-log interval from 0.006 to 0.251, against a truth of 0.174. The log-log band is built on a scale that cannot leave [0, 1], and it bends away from the edge rather than through it.

The interval at the end of the curve

The interval most software prints around a survival curve covers 89.7% at five years, where 3.3 of forty subjects are still being watched and where the curve is actually read. The same variance carried on a log–log scale covers 94.8% there — and the failure was never the width.

Two worlds, one Kaplan–Meier curve, two truths. World A gives each subject a frailty with mean one and variance 1, and multiplies both its event hazard (0.35) and its dropout hazard (0.5) by it, so the subjects likeliest to leave are the ones likeliest to fail. World B has independent event and dropout times whose hazards are world A's crude hazards. Kaplan–Meier over 1000 studies of 400 gives the same curve from both — 0.7750 and 0.7763 at t = 1; 0.6627 and 0.6641 at t = 2; 0.5924 and 0.5932 at t = 3; 0.5047 and 0.5042 at t = 5 — and that curve is world B's truth, 0.5052 at t = 5. World A's truth is 0.3636 there. The dashed lines are the two bounds that assume nothing, from every dropout failing on leaving (0.1905 at t = 5) to none ever failing (0.6667).

A dropout the data cannot see

Two worlds leave the same record to the last detail a study can write down — the same times, the same share ending in the event, the same share leaving first — and a log-rank test between them rejects at its own 5% level at every sample size from a hundred to sixteen hundred. Kaplan–Meier converges on 0.5052 at t = 5 from both. The truth is 0.5052 in one and 0.3636 in the other, and what is left to argue about is where between two bounds to stand.

The risk of one cause, estimated two ways. Two causes of an ending event with constant hazards 0.2 (the one of interest) and 0.3 (the competitor), random dropout at 0.1 and follow-up to 6. The lower line is the cumulative incidence, (0.2/0.5)(1 − e^(−0.5t)), the chance of actually having had this event by t; the dots on it are the Aalen–Johansen estimate over 2000 studies of 300, 0.3670 at t = 5 against 0.3672. The upper line is 1 − e^(−0.2t), and the dots on it are one minus Kaplan–Meier with the competing event treated as censoring: 0.6318 at t = 5 against 0.6321. The second is larger by a factor of 1.722 at t = 5, and it is not an error of estimation. It estimates, correctly, the risk in a population where the competing cause does not exist.

One minus Kaplan–Meier is not a risk

With two ways for observation to end, one minus Kaplan–Meier for one cause reads 0.6318 at t = 5 where the chance of actually having had that event is 0.3670. Added across the two causes, the complements reach 1.4088 — more than the whole cohort. Nothing is estimated badly: the complement estimates, correctly, the risk in a world where the other cause does not exist.

One treatment, a different hazard ratio at every follow-up. The hazard ratio a Cox model converges to, found as the root of its expected score by numerical integration, as the trial runs longer; dropout at 0.1 throughout. The proportional treatment reads 0.5 at every τ. The waning treatment reads 0.5000 at τ = 1, 0.6362 at τ = 3 and 0.7890 at τ = 8 — the same two arms, the same effect in the same first year, and a number that drifts towards one as later, effect-free events are added to the average. The dots are the mean of 400 Cox fits with 400 subjects an arm: 0.5014 at τ = 1, 0.7020 at τ = 2, 0.7635 at τ = 3, 0.8034 at τ = 5, 0.8208 at τ = 8. The crossing treatment reads 0.3429 at τ = 1, exactly 1 at τ = 3 by construction, and 1.0611 at τ = 8: beneficial, null or harmful according to when the trial stopped.

The hazard ratio the follow-up chose

A treatment that halves the hazard for one year and then does nothing has a Cox hazard ratio of 0.5000 if the trial stops at one year, 0.7617 at three and 0.8194 at eight. Nothing about the treatment differs between those numbers. When hazards are not proportional the hazard ratio is an average, and the length of follow-up and the dropout rate choose its weights.

The difference in restricted mean survival at every horizon, in three worlds. Treatment minus control, in closed form, with dropout irrelevant to the truth. The proportional treatment's difference grows to 0.4766 at τ = 3 and the waning treatment's to 0.2675. The crossing treatment's rises to 0.1776 at τ = 2, near where the two survival curves cross, and falls back to 0.1366 at τ = 3. The ticks along the bottom are the eleven horizons, from 0.5 to 3 in quarters, at which a trial below reads its differences.

A horizon chosen after looking

A difference in restricted mean survival read at whichever of eleven horizons looks most convincing rejects 11.24% of trials in which the treatment does nothing, against 4.70% at a horizon fixed in advance. The correlation of the differences across horizons is closed, and the Gaussian process it defines prices the choice at a critical value of 2.317 — which brings the counted size back to 4.99% and keeps 96.92% of the power that a horizon nobody could have known to fix would have had.

All essays