A dropout the data cannot see
Worth reading first: The data that stops early.
Every argument about a survival curve so far has rested on one sentence: a subject who stops being observed carries no information about when the event would have happened. It is the assumption that lets Kaplan–Meier keep a censored subject in the risk set for as long as they were watched and then drop them without prejudice, and it is the only assumption the estimator makes.
In a clinic it is frequently false in a direction everybody can name. The patients who stop coming are not a random sample of the patients. Some stop because they are too ill to travel, and those are the ones closest to the event. When that is how dropout works, the subjects who leave the risk set take their high hazard with them, the ones who remain look healthier than the cohort was, and the curve drawn from them is too high.
It is often said that nothing in the data can detect this. That is a strong claim, and it is usually offered as a warning rather than demonstrated. Here it is built.
Two worlds built to leave one record
World A is a shared frailty. Every subject carries an unobserved number , drawn from a gamma distribution with mean one and variance one, and both of their hazards are multiplied by it: the event at , dropout at , and the study ends at six. A frail subject is likelier to fail and likelier to leave, which is exactly the clinic’s story written as arithmetic. The true survival curve, averaging over the frailty, is , and at t = 5 it is 0.3636.
World B has no frailty and no dependence at all. Event time and dropout time are independent, and each has a hazard that falls with time,
Its true survival curve is , and at t = 5 it is 0.5052.
Those two hazards were not chosen for looking plausible. They are world A’s crude hazards — the rate at which subjects still under observation in world A have the event, or leave, at each time. In world A that rate falls with time for a reason: the frail are removed first, so the survivors’ average frailty at time t is . World B takes the falling rate and makes it a property of every individual instead of a property of who is left.
Kaplan–Meier, run on four hundred subjects from each world a thousand times, reads 0.5047 from world A and 0.5042 from world B at t = 5, each with a standard error of about 0.001. At t = 2 it reads 0.6627 and 0.6641, against world B’s truth of 0.6643 and world A’s of 0.5882. It converges on world B’s curve from both worlds’ data, because it is a product of crude hazards and the crude hazards are the same.
Why the curve follows the survivors
It is worth seeing why that has to happen rather than taking it from a simulation. Each factor in Kaplan–Meier’s product is one minus the share of the subjects still being watched who have the event at that moment, and over many subjects that share, per unit of time, is the crude hazard: in world A, 0.35 times the average frailty of whoever is still under observation.
Two different selections have acted on that group by t = 2, and only one of them acts on the cohort the true curve describes. Having had no event already removes the frail preferentially, so among everybody still event-free the average frailty has fallen to 0.588, and the hazard the true curve is built from is 0.35 × 0.588, 0.2059. But the subjects Kaplan–Meier can see have also not dropped out, and dropout removed the frail as well. Among them the average frailty is 0.370, and the crude hazard is 0.1296 — about five-eighths of the true one.
So the estimator is doing exactly what it was built to do, faithfully, to a group that the process of being observed has already selected. A running product of the survivors’ hazards is the survival curve of a population with the survivors’ hazards, which is world B. Independence of censoring is precisely the assumption that the second selection removes nobody in particular, and under a frailty it removes the people most likely to fail.
Everything a record contains
A survival dataset writes down two things for each subject: a time, and which of two things ended observation. So the whole of what any analysis can learn is the joint law of that pair — the chance of still being observed at each time, and how the endings divide between the event and dropout.
In both worlds the chance of still being observed at t is , and the endings divide between the event and dropout in the fixed proportion 0.35 to 0.5. That is shown two ways that share no arithmetic. The closed forms are derived from world A’s frailty; independently, world B’s own hazards are integrated against its own chance of still being observed, numerically, and the crude incidences that come out match world A’s to within , the error of the quadrature. And counted over 400,000 subjects in each world, the three shares at t = 5 read 0.1897, 0.3334 and 0.4769 in world A, 0.1900, 0.3334 and 0.4766 in world B, against 0.1905, 0.3333 and 0.4762 closed.
This is Tsiatis’s theorem from 1975, made concrete: for any joint model of dependent event and censoring times there is a model with independent times whose recorded data have exactly the same law, and its event-time curve is the one Kaplan–Meier estimates. The construction is always available. So every survival dataset, however it arose, is consistent with Kaplan–Meier being right, and consistent with it being wrong by any amount the bounds below allow. The crude hazards are identified; the event time’s own distribution is not.
It is the same statement as the one competing risks rests on, and for the same reason. Dropout here is a competing reason for observation to end, and what the data determine is each ending’s cause-specific hazard — never the joint law of the latent times behind them.
A test with nothing to find
“No test can distinguish them” is a claim about every test, so it is checked in the form a skeptic would choose: take the most natural two-sample test for survival data, give it a sample from each world, and count how often it says they differ.
Between the twins, the log-rank test rejects 5.2%, 5.8%, 5.0%, 5.2% and 4.8% of the time at 100, 200, 400, 800 and 1,600 subjects a side, over six hundred tests at each size. Those are its own nominal level, inside a standard error of about 0.9 points, and they do not move as the samples grow. A test that is exactly at its size under a true null is behaving perfectly, and here the null — same recorded law — is true.
The control is what stops that from being read as a weak test. Take world A’s event times and give its subjects dropout at a flat 0.5, ignoring the frailty. The recorded data now genuinely differ, because the frail no longer leave faster, and the same test rejects 17.7%, 32.3%, 61.2%, 87.5% and 99.5%. The test has power against every difference the record can carry. The difference between world A and world B is not one of them.
A bound is left behind, and that changes what can be said
This is recognisably the situation a missing value leaves, and the contrast is worth drawing carefully, because it is where survival data are better off.
An unrecorded outcome leaves nothing. The shift that separates the worlds in that construction is a parameter about which the data are silent in every direction, and the truth moves along a straight line with no end.
A censored subject leaves a bound. The two readings that need no assumption at all — every dropout failed the instant they left, or none of them ever failed — bracket the truth in every world consistent with the record. In this one, at t = 5, those bounds are 0.1905 and 0.6667. The unknown is confined to an interval 0.4762 wide, which is wide, but it has ends. Independent censoring puts the curve 66.1% of the way up that interval. World A’s truth sits 36.4% of the way up.
The interval is not a fixed width, and how it grows says where the assumption does its work. At t = 1 the bounds are 0.5405 and 0.8108; at t = 2, 0.3704 and 0.7407; at t = 5, 0.1905 and 0.6667. The widths — 0.2703, 0.3704 and 0.4762 — are exactly the share of the cohort that had dropped out by each time, which is what the bounds are: every subject who left, placed all above the curve or all below it. The assumption is asked to do nothing at the start of follow-up and more with every subject who leaves, so it is doing the most at the end of the curve, which is where survival is read.
So the question an analysis actually faces is not “is Kaplan–Meier right”, which has no answer, but “where inside a known interval does an assumption place the answer, and how far does a different assumption move it”. That question has an answer for every assumption, and it can be drawn.
A dial between the two bounds
The assumption Kaplan–Meier makes can be written as a number. Suppose a subject who drops out has, from the moment of leaving, an event hazard δ times the crude hazard of the subjects still being watched. Then δ = 1 is independent censoring; δ = 0 is a dropout who never fails; δ = ∞ is a dropout who fails on leaving.
Written that way, the implied survival at t uses only what the record contains:
with the crude density of leaving and the crude cumulative hazard of the event. Every curve in the figure is therefore exactly as consistent with the data as every other, and at its ends the dial reproduces the bounds: at t = 5 it reads 0.6667 at δ = 0, 0.5052 at δ = 1, and 0.1905 at δ = ∞. Between them, δ = 0.5, 2 and 5 give 0.5759, 0.4063 and 0.2780.
The same quantity can be read off a single dataset without knowing any closed form — the share still observed at t, plus each dropout’s contribution discounted by the Nelson–Aalen estimate of the hazard between their leaving and t. Over two thousand studies of four hundred from world A, that plug-in reads 0.6668, 0.5760, 0.5056, 0.4073, 0.2800 and 0.1901 at the six settings, against the closed values above. At δ = 1 it is a second estimator of Kaplan–Meier’s limit built by different arithmetic, and it agrees with Kaplan–Meier’s own 0.5051. The one visible departure is at δ = 5, where the plug-in is high by 0.0020, about three of its standard errors: an exponential of an estimated hazard is biased upward in a finite sample, and the bias grows with how hard δ leans on the estimate.
Where on the dial a frailty sits
The frailty world is not a setting of the dial by construction, so it is fair to ask where it lands.
World A’s truth at t = 5, 0.3636, is reached at δ = 2.637. A subject who leaves in world A has, in effect, a hazard over two and a half times that of a subject still being watched — which is not an exotic assumption. It is what a frailty of variance one does: at the moment of leaving, a dropout’s expected frailty is twice that of the subjects remaining, and it stays high because what made them leave is still true of them.
A sensitivity analysis that swings δ from one half to two would be described in most reports as generous. It covers survival at t = 5 from 0.5759 down to 0.4063, and it does not reach this world’s truth. The range was wide in the units of δ and narrow in the units that matter, because the dial is steep near one and flattens only far above it.
A frailty is not one setting of the dial
There is a second thing the sweep hides, and it is the more important one for anybody designing a sensitivity analysis.
The δ that reproduces world A is 2.246 when the curve is read at t = 1, 2.395 at t = 2, 2.499 at t = 3 and 2.637 at t = 5. It drifts. A world with a real, fixed, easily described mechanism of dependence is not a member of the constant-δ family at all; it passes through the family, touching a different member at each time. With a smaller frailty variance the dependence is weaker and everything moves towards one — at t = 5 the gap between Kaplan–Meier’s limit and the truth is 0.1416, 0.1068, 0.0693 and 0.0403 for variances of one, a half, a quarter and an eighth, and δ* is 2.637, 1.853, 1.446 and 1.232 — but at every variance the reading still depends on when the curve is read.
The consequence is specific. A sensitivity parameter is a family of worlds chosen for being easy to compute, and a plausible mechanism need not belong to it. Reporting “robust for δ between 0.5 and 2” is a statement about that family. It is not a statement about dependent censoring.
What is proved here, and what is only constructed
Some of this is general and some belongs to one world, and they need keeping apart.
The non-identifiability is general. Tsiatis’s result holds for any joint distribution of event and censoring times, and nothing here depends on the gamma, the constant hazards or the numbers chosen; the twin world is one instance of a construction that is always available. The bounds are general too: they need no model and hold in every world consistent with the record.
The size of the gap is not general. A frailty of variance one acting multiplicatively on both hazards, with dropout faster than the event, is a strong dependence chosen to make the effect visible, and 0.1416 at t = 5 is a property of that choice. The direction is not guaranteed either. When the subjects who leave are the ones doing well — patients who recover and stop attending, participants who move away for a better job — δ is below one, the curve Kaplan–Meier draws is too low, and nothing in the record distinguishes that world from this one or from the independent twin. On this record, leavers with half the hazard of those who stay put survival at t = 5 at 0.5759, 0.0707 above the curve Kaplan–Meier draws; the error is the same kind in the other direction.
And the dial is one family among many. It fixes the dropouts’ hazard after leaving as a constant multiple of the remaining subjects’ crude hazard, which is simple and legible and demonstrably not flexible enough to contain a frailty. Other families — a multiple that decays with time since leaving, a dependence that runs through a latent class — would place world A differently. What stays true across all of them is that the data choose the curve’s crude hazards and never the assumption.
What an analysis can report instead of one curve
A curve with an interval around it, reported as though independent censoring were known, has made an assumption and not said so. Three things can be said instead, and none of them needs data the study did not collect.
The two bounds. They require nothing, and they say at once how much the assumption is being asked to do. At t = 5 in this world they are nearly half a unit apart. In a study with little dropout they may be close enough to settle the question with no model at all.
The dial across a range, with the range argued for. A reader who can see survival at t = 5 against δ can see how far any assumption they find credible moves the answer. The range has to come from somewhere other than the data — from why subjects in this study were known to leave — and that argument belongs in the report beside the curve.
Which direction is credible. The dial is not symmetric in plausibility. In a trial where leaving was driven by illness, δ below one is hard to defend and the upper part of the interval can be set aside; in a trial where leaving was driven by recovery, the reverse. That is an argument about the study, and it is the only kind of argument that can narrow an interval the record has left open.
Every one of those is a more honest report than a single Kaplan–Meier curve, and they are the same arithmetic: the curve is the dial read at one, and a prior placed on δ would simply be this figure averaged over the prior, because the likelihood carries no information about δ at all.
Where this goes next
The dependence here ran through something nobody recorded. The case that should come next on censoring is the one where it runs through something that was recorded: dropout that depends on a measured covariate — a symptom score, a visit history — which makes the censoring independent given that covariate, and so identifiable. Weighting each subject still under observation by the inverse of their estimated chance of still being observed, given the covariate, removes the bias the way a score that balances removes confounding, and it has the same failure modes: a chance of remaining that is estimated from the same data, extreme weights where few similar subjects stay, and a model for dropout that can be wrong. It is a distinct argument from this one — this essay concerns what no measurement can reach, and the next would ask how much of the dependence a measurement does reach, and what the weights cost in variance when it does.
There is also the question this one deliberately set aside: the curve at the end of follow-up is where the dial is widest and where any interval is least trustworthy, so the two uncertainties compound exactly where survival is usually read.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The assumption nothing tests — both name non-identifiability, observational equivalence, partial identification
- Three mechanisms and one dataset — both name estimand, non-identifiability
Named objects
A flat tag is an object no other essay names yet.
Cause specific hazardCensoringCompeting risksDependent censoringEstimandFrailtyKaplan–MeierLog-rank testNelson–AalenNon-identifiabilityObservational equivalencePartial identificationSensitivity analysisSensitivity parameterSurvival curve