When the data stops early

The data that stops early

A subject still event-free when a study ends is not missing and not observed. It is known to exceed something, which is a third state most tools have no slot for — and the two obvious ways of forcing it into one are wrong by 31 and 13 percentage points.

A study follows a hundred and twenty subjects for three years. At the end, some have had the event and some have not. The ones who have not are the problem, and the problem is that their data is neither present nor absent.

The same study read three waysAt time 2 the truth is 0.497. Kaplan–Meier gives 0.532; dropping the censored subjects gives 0.180; treating the censoring time as the event time gives 0.392. Both naive readings understate survival, because the subjects they mishandle are the ones doing well.00.2500.5000.75010123timesurvivingthe truthcensored droppedcensoring as event120 subjects, one datasetthe naive readings fail in the same direction
Fig. 1 The same study read three ways against the truth it is trying to recover. Two of the three readings are badly wrong and both fail in the same direction.

A third state

For a subject who has not had the event, the survival time is not unknown in the ordinary sense. Something specific is known about it: it is greater than three years.

That is real information. It is not a missing value, because a missing value tells nothing, and this tells something bounded and useful. It is also not an observation, because the quantity of interest was not observed.

Statistical software mostly offers two slots — a number or a blank — and the third state has to be encoded as a pair: the time last seen, and a flag saying whether that time was an event or the end of observation. Every method in this field is built on that pair, and the two failures below are what happens when the pair is collapsed back into one number.

The first wrong reading: drop them

The instinct when a tool will not accept a row is to remove the row. Analyse the subjects who had the event, ignore the rest.

Measured against a known truth, across fifteen hundred simulated studies: the true survival at two years is 0.497, and the complete-case reading gives 0.189.

Off by thirty-one percentage points, and off in a direction that is not random. The subjects being dropped are exactly the ones who survived longest — that is what being censored means — so removing them removes the best outcomes from the sample. What is left is the subset who had the event early, and their survival times are, by construction, short.

The clearest way to see the absurdity: a study where nobody has the event still produces an answer under this reading, computed from an empty set. And a study where the treatment works so well that most subjects survive the follow-up period will report worse survival than one where the treatment works less well, because the successful subjects are the ones excluded.

This is selection on the outcome, and it is the same structure as the winner’s curse: the selection rule is correlated with the quantity being estimated, so the surviving sample is systematically unrepresentative and no amount of data fixes it.

The second wrong reading: pretend it happened

The other instinct is to keep the row and treat the censoring time as the event time. The subject was last seen at three years, so record an event at three years.

This gives 0.370 against the true 0.497 — off by thirteen points, better than dropping them and still badly wrong.

The mechanism is straightforward. Every censored subject is recorded as having failed at the moment observation stopped, when in fact they had not failed at all and might have gone on for years. The estimate is a lower bound on survival dressed up as an estimate.

It is worth noticing that both naive readings fail in the same direction. Both understate survival, and both do so because censored subjects are the ones doing well — dropping them removes good outcomes, and converting them to events converts good outcomes into bad ones.

That the two obvious errors agree with each other is a trap. Two analysts using different wrong methods will reach similar conclusions and take the agreement as reassurance.

Kaplan–Meier from 120 subjects, 59 of them censored. The step curve is the estimate, the smooth curve is the truth it is trying to recover. 59 of 120 subjects were still event-free when observation stopped; they are not dropped, and they are not counted as events — they leave the risk set at the time they were last seen.
Fig. 2 What the information actually supports: the estimate that uses the censored subjects for as long as they were observed.

The two bounds nobody can argue with

Before any assumption about the censoring is made, there are two readings that need none, and they bracket every honest answer.

Treat every censored subject as having had the event the instant they left. That is the smallest survival curve consistent with the data.

Treat every censored subject as surviving for ever. That is the largest.

The truth is between them, always, with no assumption whatever — and the gap between them is the whole of what the censoring costs. On a study with heavy censoring that gap is wide enough to contain opposite conclusions, and on one with little censoring it is narrow enough to settle the question by itself.

What the Kaplan–Meier construction does is choose a point inside that interval, and it earns the choice by assuming the censoring carries no information about the event time. That is a real assumption and it is the only one, which is why the two bounds are worth computing beside the curve: they say how much the assumption is being asked to do.

What the censored subject actually contributes

The resolution is to use each subject for the period they were observed, and stop counting them afterwards.

A subject censored at eighteen months establishes that the event did not happen in the first eighteen months. That is a genuine contribution to every estimate about the first eighteen months, and no contribution at all to estimates about later times.

So the subject stays in the denominator — the risk set — until the moment they leave, and then they are gone. They are neither excluded from the analysis nor counted as an event; they are counted as present, and then absent.

That is the whole idea, and the estimator built on it recovers the truth to within a fraction of a point where the naive readings are off by thirteen and thirty-one.

The risk set, 60 subjects. The denominator Kaplan–Meier divides by at each event. It falls both when an event happens and when a subject is censored, which is exactly how a censored subject contributes: it was at risk until it left, and it is not counted afterwards.
Fig. 3 The denominator, falling as subjects have events and as subjects leave. Both reduce it, for different reasons.

Why the estimate has to be a curve

A structural point that follows from the above and explains the shape of the whole field.

With censoring, the amount of information available is different at different times. Everybody contributes to the estimate of survival at one month. Very few contribute to the estimate at three years, because most have either had the event or left.

So the precision of the estimate degrades as time increases, and a single summary number cannot represent that. The right output is a curve with the uncertainty widening to the right, and any reduction of it to one number — median survival, five-year survival, mean time to event — discards the structure that made the analysis honest.

The mean is the worst of those reductions and is worth a warning on its own. A mean survival time cannot be estimated at all if the study ends while subjects are still alive, because the mean depends on the tail of the distribution and the tail is precisely what was not observed. Software will compute one anyway, by assuming the curve drops to zero at the last observation, and that assumption is a fabrication about the part of the data that does not exist.

Medians are safer, because a median only needs the curve to reach 0.5 — which it may do well inside the follow-up period. When it does not, no median is estimable either, and the honest report says so.

What has to be true about the censoring

Everything here rests on an assumption that deserves to be stated as clearly as the failures it prevents.

The censoring must be independent of the event: subjects who leave the study must not be more or less likely to have the event than those who stay, given what has been observed.

That holds by construction for administrative censoring — the study ended, and the calendar does not know anything about the subjects. It is far less safe for dropout. A subject who leaves because they became too unwell to attend is censored for a reason connected to the outcome, and every method in this field will mis-estimate as a result.

The direction is predictable. If sicker subjects drop out, the remaining risk set is healthier than the cohort, and survival is overstated — the opposite direction from the naive readings, so the two errors can hide each other.

And it is not checkable from the data. Both the independent and dependent stories produce the same observed pattern of times and flags; distinguishing them requires knowing why people left, which is why survival studies collect reasons for withdrawal and why analyses report them.

That is the same shape as coverage under a misspecified model: the assumption that matters most is the one the data cannot verify, and the discipline is to state it rather than to test it.

Bias against the amount of censoring, truth 0.497. Kaplan–Meier tracks the truth from 37% censoring to 78%. Dropping the censored subjects gets steadily worse, from 0.220 to 0.037 against a truth of 0.497.
Fig. 4 How the naive error grows with how much censoring there is, and how the proper estimator does not.

The general form of the mistake

Censoring is a specific case of a pattern this site keeps meeting, and naming the pattern makes it transferable.

The data is incomplete in a way that carries information, and the incompleteness is related to the quantity being estimated. Handling it by removing the awkward cases selects on the outcome; handling it by imputing a convenient value fabricates observations. The correct treatment is to use exactly what is known and no more.

The same structure appears in the bootstrap for a maximum, where the resample can never exceed the largest observed value and the interval covers 0% of the time. It appears in detection limits, where a measurement below an instrument’s threshold is recorded as zero or as the threshold and is neither. It appears in survey non-response, where the people who do not answer differ from those who do in ways related to the question.

In every case the fix has the same character. Encode what is actually known — a bound rather than a value — and use a method that accepts bounds. The failures come from tools that only accept numbers, and from the habit of giving them one.

The same study read three ways. At time 2 the truth is 0.741. Kaplan–Meier gives 0.730; dropping the censored subjects gives 0.243; treating the censoring time as the event time gives 0.542. Both naive readings understate survival, because the subjects they mishandle are the ones doing well.
Fig. 5 The same three readings at a lower event rate, where more subjects are censored and the naive errors are larger.

Why the naive readings are still common

Three reasons, and none of them is that people do not know better.

The tool asks for a number. A spreadsheet column, a regression that takes a response vector, a machine learning library expecting a target: all of them want one value per row. Encoding a bound requires two columns and a method that reads both, and where the method is not to hand the pressure to collapse the pair is considerable.

The censored subjects look like a nuisance rather than data. They are the rows with no event, and in a study about events they read as the uninformative ones. The intuition is exactly backwards — at high censoring they are most of the information — but it is a natural intuition and nothing in the data display corrects it.

And the failure is invisible in the output. Both naive readings produce a plausible survival estimate, a plausible curve and a plausible comparison between groups. Nothing about 0.189 announces that the truth is 0.497. There is no error message, no warning, and no diagnostic in the standard output that flags the problem.

The third is the one that makes this dangerous rather than merely wrong, and it is the same property that makes the bootstrap for a maximum dangerous: a method producing a confident, well-formed, entirely wrong answer with no distress signal.

The only way to find out is to run the method on a problem whose answer is known and count. That is what the figures here do, and it is the reason this site simulates the data rather than using a real dataset — a real dataset cannot show that an estimator is biased, because the truth is not available to compare against.

How much censoring is too much

A question with a more encouraging answer than the failures above suggest.

Sweeping the amount of censoring from 37% of subjects to 71%, the proper estimator stays within a fraction of a point of the truth at every level. The complete-case reading gets worse at every step — 0.221, 0.186, 0.133 and 0.072 against a truth of 0.497 — degrading steadily as censoring rises.

So the answer is that heavy censoring is not by itself a problem for a correct analysis. A study where 71% of subjects never had the event still supports an accurate estimate of survival at two years, because the censored subjects were observed for part of the period and that part is used.

What heavy censoring does affect is precision, not accuracy. Fewer events means a wider interval, and at some point the curve becomes too uncertain at late times to say anything. That shows up honestly in the standard error, which is what a well-behaved estimator does.

The practical rule: heavy censoring is a reason to check the interval width at the times of interest, and not a reason to distrust the estimate. It is a reason to distrust the naive readings enormously, since their error scales directly with it.

What this field will establish

Two essays, and the division is simple.

This one has been about what censoring is — a bound rather than a value, information rather than missingness — and what the two obvious mishandlings cost: thirteen points for treating censoring as an event, thirty-one for dropping the censored subjects.

The next is about the estimator that uses the information properly, why it takes the form of a running product, how its uncertainty is computed, and what it can and cannot support.

Both rest on the same measurement discipline the rest of the site uses. The truth is known because the data is generated from a stated rule, the estimators are run against it thousands of times, and the bias of each is counted rather than reasoned about.

Kaplan–Meier from 300 subjects, 139 of them censored. The step curve is the estimate, the smooth curve is the truth it is trying to recover. 139 of 300 subjects were still event-free when observation stopped; they are not dropped, and they are not counted as events — they leave the risk set at the time they were last seen.
Fig. 6 And the estimate at a larger cohort, where the curve tightens onto the truth and the censoring marks thin out along it.

A note on the word “survival”

The vocabulary comes from medicine and the methods are not about death, which is worth saying because the terminology puts people off a technique they need.

The structure applies wherever the quantity of interest is a time until something happens and observation stops before it happens for some units. That covers: time until a component fails, until a customer cancels, until a released prisoner reoffends, until a machine part is replaced, until a patent is cited, until a process completes.

In every one of those, some units have not had the event when the data is pulled, and the same three readings are available with the same consequences. A churn analysis that drops customers who have not yet cancelled is making the thirty-one-point error, and it will report a customer lifetime far shorter than the truth — while looking entirely reasonable, because the average tenure of customers who cancelled is a plausible-sounding quantity.

That particular case is common enough to be worth stating on its own: the average lifetime of the customers who left is not the average customer lifetime, and the gap between them is the whole subject of this field.

The methods are also older and more general than their medical name suggests, having been developed independently in engineering as reliability analysis and in economics as duration analysis. Three literatures, three vocabularies, one estimator.

The one thing to check in a published analysis

A reader without the data can still ask one question, and it separates the careful analyses from the rest.

How many subjects were censored, and were they used?

A paper reporting an event-time analysis should state how many units had the event and how many did not. If the second number is large and the method described is a mean or a comparison of averages, the analysis is almost certainly making one of the two errors here — and the size of the error scales with that number.

If the method named is Kaplan–Meier, a proportional hazards model, or anything that takes a status indicator, the censored subjects are being used properly.

The question is answerable from the methods section alone, it takes a moment, and the difference between the two answers is thirty percentage points on the measurement in this essay.

The encoding, once, plainly

The whole field reduces to a data-representation decision, and it is worth writing out because it is the part that gets skipped.

Every unit gets two values: the last time it was observed, and whether that observation was the event or the end of watching. Not one value. Not a number with blanks for the ones that did not happen.

Given that pair, every method here works and the estimate is accurate at any level of censoring. Given a single number, no method can recover what was lost, because the information distinguishing “failed at three years” from “still fine at three years” has already been discarded.

The failures in this essay are not analysis errors. They are consequences of a representation choice made earlier, usually by whoever set up the spreadsheet, and usually without a decision being consciously taken.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CensoringComplete-case analysisRisk setSelection biasSurvival analysis