When the data stops early

A dropout the data cannot see

Two worlds leave the same record to the last detail a study can write down — the same times, the same share ending in the event, the same share leaving first — and a log-rank test between them rejects at its own 5% level at every sample size from a hundred to sixteen hundred. Kaplan–Meier converges on 0.5052 at t = 5 from both. The truth is 0.5052 in one and 0.3636 in the other, and what is left to argue about is where between two bounds to stand.

Worth reading first: The data that stops early.

Every argument about a survival curve so far has rested on one sentence: a subject who stops being observed carries no information about when the event would have happened. It is the assumption that lets Kaplan–Meier keep a censored subject in the risk set for as long as they were watched and then drop them without prejudice, and it is the only assumption the estimator makes.

In a clinic it is frequently false in a direction everybody can name. The patients who stop coming are not a random sample of the patients. Some stop because they are too ill to travel, and those are the ones closest to the event. When that is how dropout works, the subjects who leave the risk set take their high hazard with them, the ones who remain look healthier than the cohort was, and the curve drawn from them is too high.

It is often said that nothing in the data can detect this. That is a strong claim, and it is usually offered as a warning rather than demonstrated. Here it is built.

Two worlds, one Kaplan–Meier curve, two truths. World A gives each subject a frailty with mean one and variance 1, and multiplies both its event hazard (0.35) and its dropout hazard (0.5) by it, so the subjects likeliest to leave are the ones likeliest to fail. World B has independent event and dropout times whose hazards are world A's crude hazards. Kaplan–Meier over 1000 studies of 400 gives the same curve from both — 0.7750 and 0.7763 at t = 1; 0.6627 and 0.6641 at t = 2; 0.5924 and 0.5932 at t = 3; 0.5047 and 0.5042 at t = 5 — and that curve is world B's truth, 0.5052 at t = 5. World A's truth is 0.3636 there. The dashed lines are the two bounds that assume nothing, from every dropout failing on leaving (0.1905 at t = 5) to none ever failing (0.6667).
Fig. 1 Two worlds and the curve Kaplan–Meier draws from each. The estimates from both sit on one line, which is the truth in one world and not in the other; the dashed lines are the two bounds that assume nothing.

Two worlds built to leave one record

World A is a shared frailty. Every subject carries an unobserved number ZZ, drawn from a gamma distribution with mean one and variance one, and both of their hazards are multiplied by it: the event at 0.35Z0.35Z, dropout at 0.5Z0.5Z, and the study ends at six. A frail subject is likelier to fail and likelier to leave, which is exactly the clinic’s story written as arithmetic. The true survival curve, averaging over the frailty, is (1+0.35t)1(1 + 0.35t)^{-1}, and at t = 5 it is 0.3636.

World B has no frailty and no dependence at all. Event time and dropout time are independent, and each has a hazard that falls with time,

hevent(t)=0.351+0.85t,hdropout(t)=0.51+0.85t.h_{\text{event}}(t) = \frac{0.35}{1 + 0.85t}, \qquad h_{\text{dropout}}(t) = \frac{0.5}{1 + 0.85t}.

Its true survival curve is (1+0.85t)0.35/0.85(1 + 0.85t)^{-0.35/0.85}, and at t = 5 it is 0.5052.

Those two hazards were not chosen for looking plausible. They are world A’s crude hazards — the rate at which subjects still under observation in world A have the event, or leave, at each time. In world A that rate falls with time for a reason: the frail are removed first, so the survivors’ average frailty at time t is 1/(1+0.85t)1/(1 + 0.85t). World B takes the falling rate and makes it a property of every individual instead of a property of who is left.

Kaplan–Meier, run on four hundred subjects from each world a thousand times, reads 0.5047 from world A and 0.5042 from world B at t = 5, each with a standard error of about 0.001. At t = 2 it reads 0.6627 and 0.6641, against world B’s truth of 0.6643 and world A’s of 0.5882. It converges on world B’s curve from both worlds’ data, because it is a product of crude hazards and the crude hazards are the same.

Why the curve follows the survivors

It is worth seeing why that has to happen rather than taking it from a simulation. Each factor in Kaplan–Meier’s product is one minus the share of the subjects still being watched who have the event at that moment, and over many subjects that share, per unit of time, is the crude hazard: in world A, 0.35 times the average frailty of whoever is still under observation.

Two different selections have acted on that group by t = 2, and only one of them acts on the cohort the true curve describes. Having had no event already removes the frail preferentially, so among everybody still event-free the average frailty has fallen to 1/(1+0.7)=1/(1 + 0.7) = 0.588, and the hazard the true curve is built from is 0.35 × 0.588, 0.2059. But the subjects Kaplan–Meier can see have also not dropped out, and dropout removed the frail as well. Among them the average frailty is 1/(1+1.7)=1/(1 + 1.7) = 0.370, and the crude hazard is 0.1296 — about five-eighths of the true one.

So the estimator is doing exactly what it was built to do, faithfully, to a group that the process of being observed has already selected. A running product of the survivors’ hazards is the survival curve of a population with the survivors’ hazards, which is world B. Independence of censoring is precisely the assumption that the second selection removes nobody in particular, and under a frailty it removes the people most likely to fail.

Everything a record contains

A survival dataset writes down two things for each subject: a time, and which of two things ended observation. So the whole of what any analysis can learn is the joint law of that pair — the chance of still being observed at each time, and how the endings divide between the event and dropout.

Everything the record can show, and it is the same in both worlds. The three things a survival dataset records by each time: the share still under observation, the share whose observation ended with the event, and the share who left first. The lines are the closed forms, and they are one set of lines, because the two worlds were built to share them: still observed (1 + 0.85t)^(−1), and the event and dropout shares 0.35/0.85 and 0.5/0.85 of the remainder. The solid dots are counted from world A and the open dots from world B, 400,000 subjects each; at t = 5 the three shares read 0.1897, 0.3334, 0.4769 in world A and 0.1900, 0.3334, 0.4766 in world B, against 0.1905, 0.3333, 0.4762 closed. The worlds differ only in what happens to people after they stop being recorded.
Fig. 2 The share still under observation, the share whose observation ended in the event and the share who left first, by each time: one set of closed curves, with dots counted from world A and open dots from world B.

In both worlds the chance of still being observed at t is (1+0.85t)1(1 + 0.85t)^{-1}, and the endings divide between the event and dropout in the fixed proportion 0.35 to 0.5. That is shown two ways that share no arithmetic. The closed forms are derived from world A’s frailty; independently, world B’s own hazards are integrated against its own chance of still being observed, numerically, and the crude incidences that come out match world A’s to within 2.0×10132.0\times10^{-13}, the error of the quadrature. And counted over 400,000 subjects in each world, the three shares at t = 5 read 0.1897, 0.3334 and 0.4769 in world A, 0.1900, 0.3334 and 0.4766 in world B, against 0.1905, 0.3333 and 0.4762 closed.

This is Tsiatis’s theorem from 1975, made concrete: for any joint model of dependent event and censoring times there is a model with independent times whose recorded data have exactly the same law, and its event-time curve is the one Kaplan–Meier estimates. The construction is always available. So every survival dataset, however it arose, is consistent with Kaplan–Meier being right, and consistent with it being wrong by any amount the bounds below allow. The crude hazards are identified; the event time’s own distribution is not.

It is the same statement as the one competing risks rests on, and for the same reason. Dropout here is a competing reason for observation to end, and what the data determine is each ending’s cause-specific hazard — never the joint law of the latent times behind them.

A test with nothing to find

“No test can distinguish them” is a claim about every test, so it is checked in the form a skeptic would choose: take the most natural two-sample test for survival data, give it a sample from each world, and count how often it says they differ.

A test that cannot tell the twins apart, and can tell a real difference. A two-sample log-rank test on the recorded data, run 600 times at each size. Between world A and its independent twin it rejects 5.2%, 5.8%, 5.0%, 5.2%, 4.8% of the time at 100, 200, 400, 800, 1600 subjects a side — its own 5% level, with a standard error of 0.9 points, at every size. Between world A and a control world with the same event times but dropout that ignores the frailty, whose recorded data genuinely differ, the same test rejects 17.7%, 32.3%, 61.2%, 87.5%, 99.5%. The test is not weak; the difference between the twins is simply not in the record.
Fig. 3 How often a log-rank test rejects at 5% between world A and its twin, and between world A and a control world whose dropout ignores the frailty, as each sample grows from a hundred to sixteen hundred.

Between the twins, the log-rank test rejects 5.2%, 5.8%, 5.0%, 5.2% and 4.8% of the time at 100, 200, 400, 800 and 1,600 subjects a side, over six hundred tests at each size. Those are its own nominal level, inside a standard error of about 0.9 points, and they do not move as the samples grow. A test that is exactly at its size under a true null is behaving perfectly, and here the null — same recorded law — is true.

The control is what stops that from being read as a weak test. Take world A’s event times and give its subjects dropout at a flat 0.5, ignoring the frailty. The recorded data now genuinely differ, because the frail no longer leave faster, and the same test rejects 17.7%, 32.3%, 61.2%, 87.5% and 99.5%. The test has power against every difference the record can carry. The difference between world A and world B is not one of them.

A bound is left behind, and that changes what can be said

This is recognisably the situation a missing value leaves, and the contrast is worth drawing carefully, because it is where survival data are better off.

An unrecorded outcome leaves nothing. The shift that separates the worlds in that construction is a parameter about which the data are silent in every direction, and the truth moves along a straight line with no end.

A censored subject leaves a bound. The two readings that need no assumption at all — every dropout failed the instant they left, or none of them ever failed — bracket the truth in every world consistent with the record. In this one, at t = 5, those bounds are 0.1905 and 0.6667. The unknown is confined to an interval 0.4762 wide, which is wide, but it has ends. Independent censoring puts the curve 66.1% of the way up that interval. World A’s truth sits 36.4% of the way up.

The interval is not a fixed width, and how it grows says where the assumption does its work. At t = 1 the bounds are 0.5405 and 0.8108; at t = 2, 0.3704 and 0.7407; at t = 5, 0.1905 and 0.6667. The widths — 0.2703, 0.3704 and 0.4762 — are exactly the share of the cohort that had dropped out by each time, which is what the bounds are: every subject who left, placed all above the curve or all below it. The assumption is asked to do nothing at the start of follow-up and more with every subject who leaves, so it is doing the most at the end of the curve, which is where survival is read.

So the question an analysis actually faces is not “is Kaplan–Meier right”, which has no answer, but “where inside a known interval does an assumption place the answer, and how far does a different assumption move it”. That question has an answer for every assumption, and it can be drawn.

A dial between the two bounds

The assumption Kaplan–Meier makes can be written as a number. Suppose a subject who drops out has, from the moment of leaving, an event hazard δ times the crude hazard of the subjects still being watched. Then δ = 1 is independent censoring; δ = 0 is a dropout who never fails; δ = ∞ is a dropout who fails on leaving.

One dataset, read at six assumptions about the people who left. The survival curve implied by the recorded law when a dropout's event hazard after leaving is δ times what it would have been had it stayed. Each curve uses only quantities the record contains — the share still observed, the crude dropout density and the crude event hazard — so every curve is equally consistent with the data. At δ = 0 the curve is the upper bound, 0.6667 at t = 5; at δ = 1 it is Kaplan–Meier's limit, 0.5052; at δ = ∞ it is the lower bound, 0.1905. In between, δ = 0.5, 2 and 5 give 0.5759, 0.4063 and 0.2780. World A's truth, the frailty world, is the separate curve, and it crosses no single one of these: it sits at δ = 2.25 at t = 1 and δ = 2.64 at t = 5.
Fig. 4 The survival curve the recorded data imply at six settings of δ, the dropouts’ hazard after leaving as a multiple of staying, with world A’s truth as the separate curve.

Written that way, the implied survival at t uses only what the record contains:

Sδ(t)=P(observed past t)+0tfdropout(s)eδ[Λ(t)Λ(s)]ds,S_\delta(t) = P(\text{observed past } t) + \int_0^t f_{\text{dropout}}(s)\, e^{-\delta\,[\Lambda(t) - \Lambda(s)]}\,ds,

with fdropoutf_{\text{dropout}} the crude density of leaving and Λ\Lambda the crude cumulative hazard of the event. Every curve in the figure is therefore exactly as consistent with the data as every other, and at its ends the dial reproduces the bounds: at t = 5 it reads 0.6667 at δ = 0, 0.5052 at δ = 1, and 0.1905 at δ = ∞. Between them, δ = 0.5, 2 and 5 give 0.5759, 0.4063 and 0.2780.

The same quantity can be read off a single dataset without knowing any closed form — the share still observed at t, plus each dropout’s contribution discounted by the Nelson–Aalen estimate of the hazard between their leaving and t. Over two thousand studies of four hundred from world A, that plug-in reads 0.6668, 0.5760, 0.5056, 0.4073, 0.2800 and 0.1901 at the six settings, against the closed values above. At δ = 1 it is a second estimator of Kaplan–Meier’s limit built by different arithmetic, and it agrees with Kaplan–Meier’s own 0.5051. The one visible departure is at δ = 5, where the plug-in is high by 0.0020, about three of its standard errors: an exponential of an estimated hazard is biased upward in a finite sample, and the bias grows with how hard δ leans on the estimate.

Where on the dial a frailty sits

The frailty world is not a setting of the dial by construction, so it is fair to ask where it lands.

Survival at t = 5 across every assumption about dropouts. The survival at t = 5 that the recorded law implies, as δ runs from nought to infinity, drawn against δ/(1 + δ) so both ends fit on one axis. It falls from the upper bound 0.6667 to the lower 0.1905 — the whole width the two worst-case bounds leave, 0.4762. Independent censoring is the single point δ = 1, at 0.5052. The dots are the dial read from 2000 studies of 400 drawn from world A, by a plug-in using the Nelson–Aalen hazard: 0: 0.6668, 0.5: 0.5760, 1: 0.5056, 2: 0.4073, 5: 0.2800, ∞: 0.1901, against 0.6667, 0.5759, 0.5052, 0.4063, 0.2780, 0.1905 closed. World A's truth, 0.3636, sits at δ = 2.637. A sensitivity range of one half to two, which reads as generous, runs from 0.5759 to 0.4063 and does not reach it.
Fig. 5 Survival at t = 5 implied by every δ from nought to infinity, drawn at δ/(1 + δ), with the plug-in from world A’s data as dots and world A’s truth as the horizontal line.

World A’s truth at t = 5, 0.3636, is reached at δ = 2.637. A subject who leaves in world A has, in effect, a hazard over two and a half times that of a subject still being watched — which is not an exotic assumption. It is what a frailty of variance one does: at the moment of leaving, a dropout’s expected frailty is twice that of the subjects remaining, and it stays high because what made them leave is still true of them.

A sensitivity analysis that swings δ from one half to two would be described in most reports as generous. It covers survival at t = 5 from 0.5759 down to 0.4063, and it does not reach this world’s truth. The range was wide in the units of δ and narrow in the units that matter, because the dial is steep near one and flattens only far above it.

A frailty is not one setting of the dial

There is a second thing the sweep hides, and it is the more important one for anybody designing a sensitivity analysis.

The assumption a frailty makes is not one number. For shared-frailty worlds with frailty variance 1, ½, ¼ and ⅛, the value of δ at which the dial reads each world's true survival, at every half-unit of time. With variance one it is 2.246 at t = 1, 2.395 at t = 2 and 2.637 at t = 5; with variance ⅛ it is 1.148, 1.170 and 1.232. Every one of these worlds is dependent censoring in the direction clinicians worry about — the sick leave — and every one sits above δ = 1. But none of them is a constant δ: the dial reading that reproduces a world drifts with the time read. A sensitivity analysis that fixes δ is a family of worlds, and a plausible mechanism need not be a member of it.
Fig. 6 For frailty worlds of variance one, a half, a quarter and an eighth, the δ that reproduces each world’s true survival, as a function of the time the curve is read at.

The δ that reproduces world A is 2.246 when the curve is read at t = 1, 2.395 at t = 2, 2.499 at t = 3 and 2.637 at t = 5. It drifts. A world with a real, fixed, easily described mechanism of dependence is not a member of the constant-δ family at all; it passes through the family, touching a different member at each time. With a smaller frailty variance the dependence is weaker and everything moves towards one — at t = 5 the gap between Kaplan–Meier’s limit and the truth is 0.1416, 0.1068, 0.0693 and 0.0403 for variances of one, a half, a quarter and an eighth, and δ* is 2.637, 1.853, 1.446 and 1.232 — but at every variance the reading still depends on when the curve is read.

The consequence is specific. A sensitivity parameter is a family of worlds chosen for being easy to compute, and a plausible mechanism need not belong to it. Reporting “robust for δ between 0.5 and 2” is a statement about that family. It is not a statement about dependent censoring.

What is proved here, and what is only constructed

Some of this is general and some belongs to one world, and they need keeping apart.

The non-identifiability is general. Tsiatis’s result holds for any joint distribution of event and censoring times, and nothing here depends on the gamma, the constant hazards or the numbers chosen; the twin world is one instance of a construction that is always available. The bounds are general too: they need no model and hold in every world consistent with the record.

The size of the gap is not general. A frailty of variance one acting multiplicatively on both hazards, with dropout faster than the event, is a strong dependence chosen to make the effect visible, and 0.1416 at t = 5 is a property of that choice. The direction is not guaranteed either. When the subjects who leave are the ones doing well — patients who recover and stop attending, participants who move away for a better job — δ is below one, the curve Kaplan–Meier draws is too low, and nothing in the record distinguishes that world from this one or from the independent twin. On this record, leavers with half the hazard of those who stay put survival at t = 5 at 0.5759, 0.0707 above the curve Kaplan–Meier draws; the error is the same kind in the other direction.

And the dial is one family among many. It fixes the dropouts’ hazard after leaving as a constant multiple of the remaining subjects’ crude hazard, which is simple and legible and demonstrably not flexible enough to contain a frailty. Other families — a multiple that decays with time since leaving, a dependence that runs through a latent class — would place world A differently. What stays true across all of them is that the data choose the curve’s crude hazards and never the assumption.

What an analysis can report instead of one curve

A curve with an interval around it, reported as though independent censoring were known, has made an assumption and not said so. Three things can be said instead, and none of them needs data the study did not collect.

The two bounds. They require nothing, and they say at once how much the assumption is being asked to do. At t = 5 in this world they are nearly half a unit apart. In a study with little dropout they may be close enough to settle the question with no model at all.

The dial across a range, with the range argued for. A reader who can see survival at t = 5 against δ can see how far any assumption they find credible moves the answer. The range has to come from somewhere other than the data — from why subjects in this study were known to leave — and that argument belongs in the report beside the curve.

Which direction is credible. The dial is not symmetric in plausibility. In a trial where leaving was driven by illness, δ below one is hard to defend and the upper part of the interval can be set aside; in a trial where leaving was driven by recovery, the reverse. That is an argument about the study, and it is the only kind of argument that can narrow an interval the record has left open.

Every one of those is a more honest report than a single Kaplan–Meier curve, and they are the same arithmetic: the curve is the dial read at one, and a prior placed on δ would simply be this figure averaged over the prior, because the likelihood carries no information about δ at all.

Where this goes next

The dependence here ran through something nobody recorded. The case that should come next on censoring is the one where it runs through something that was recorded: dropout that depends on a measured covariate — a symptom score, a visit history — which makes the censoring independent given that covariate, and so identifiable. Weighting each subject still under observation by the inverse of their estimated chance of still being observed, given the covariate, removes the bias the way a score that balances removes confounding, and it has the same failure modes: a chance of remaining that is estimated from the same data, extreme weights where few similar subjects stay, and a model for dropout that can be wrong. It is a distinct argument from this one — this essay concerns what no measurement can reach, and the next would ask how much of the dependence a measurement does reach, and what the weights cost in variance when it does.

There is also the question this one deliberately set aside: the curve at the end of follow-up is where the dial is widest and where any interval is least trustworthy, so the two uncertainties compound exactly where survival is usually read.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Cause specific hazardCensoringCompeting risksDependent censoringEstimandFrailtyKaplan–MeierLog-rank testNelson–AalenNon-identifiabilityObservational equivalencePartial identificationSensitivity analysisSensitivity parameterSurvival curve