One minus Kaplan–Meier is not a risk
Worth reading first: The data that stops early.
A report of a long study says the five-year risk of the event it studied was sixty-three per cent. The number came from a Kaplan–Meier curve for that event, subtracted from one, with every subject who had some other ending event — a death from an unrelated cause, a transplant that removes the risk, a switch to a treatment that ends follow-up for this outcome — recorded as censored at that time. It is one of the most common constructions in applied survival analysis, and in the world measured below the share of the cohort that actually had the event by five years was thirty-seven per cent.
Two ways out, and one way to be counted
The world is as simple as a world with competing risks can be. Each subject faces two causes of an ending event with constant hazards, 0.2 for the event of interest and 0.3 for the competitor, and random dropout at 0.1, with follow-up to six. Three hundred subjects a study, two thousand studies.
The quantity a patient, a planner or a payer means by “risk” is the cumulative incidence: the share of the cohort that has had this particular event by time t. With constant hazards it has a closed form,
because the chance of any ending event by t is and a fixed share of those endings are this cause. At t = 5 it is 0.3672. The Aalen–Johansen estimator, which is built to estimate exactly that, reads 0.3670 averaged over the two thousand studies, with a standard error of 0.0007.
One minus Kaplan–Meier, with the competitor censored, converges to something else: , which at t = 5 is 0.6321. It reads 0.6318. The ratio between the two numbers is 1.722, and it is not an estimation error. The estimator has converged, precisely, on a well-defined quantity that is not the risk.
A risk in a world nobody lives in
What describes is the chance of having the event by t if nobody could have the competing event first. Censoring a subject at their competing event tells Kaplan–Meier that they were still capable of having the event of interest afterwards, and that their chance of doing so was the same as that of the subjects still being watched. The estimator then multiplies forward as though the competitor had been removed from the world and its victims returned to the risk set.
That is a meaningful number for some questions. It is the quantity behind “what would the risk of this disease be if the other cause were eliminated”, which is a real question in aetiology. But it answers that question only under two assumptions, and neither can be checked. The competing cause has to be removable without changing anything else about the subjects’ hazards; and the latent time to this event for a subject who died of the other cause has to be independent of that death. The second is exactly the non-identifiability a dropout the data cannot see turns on: every competing-risks dataset is consistent with independent latent times and with dependence of any strength, and the record cannot say which.
The cumulative incidence needs neither assumption. It describes the world the cohort was actually in, with the competitor present, and it is a function of the recorded data alone. The first number asks a counterfactual question and hides the fact; the second answers a factual one.
A time that does not exist is not missing
The instruction to censor the competing event borrows its authority from ordinary censoring, and the borrowing does not survive a close look. A subject censored because the study ended was still event-free when observation stopped: the event time exists, lies beyond the censoring time, and would have been seen with longer follow-up. Its value is unrecorded, and Kaplan–Meier is a correct way of using the partial information that remains.
A subject who died of the competing cause has no later time for the event of interest to be unrecorded at. There is no longer follow-up that would reveal it. Treating that subject as censored treats a quantity that does not exist as a quantity that is merely unseen, and then asks the estimator to fill it in from the subjects who are still alive — exactly the move an imputation makes for a value that is absent, applied to a value that was never going to be there.
The distinction matters because the language of missing data then misleads. Whether censoring is “at random” is a question about a value that exists and was not recorded, the question the three missingness mechanisms are built to ask. For a competing death that question has no object. The latent time to the event of interest after a competing death is a modelling device, and every conclusion that depends on its distribution depends on something no cohort, however large and however completely followed, contains.
That is also why no amount of care in follow-up repairs the complement. A trial that loses no subject at all to administrative censoring and records every competing death with its exact date still produces, from one minus Kaplan–Meier, the risk in a population where those deaths did not happen. Complete data make the estimate precise. They do not make the population real.
More than everybody
The clearest sign that the complement is not a risk is arithmetic that a risk cannot do.
Every subject has at most one ending event, so the share who had cause one plus the share who had cause two cannot exceed the share who had either, and that cannot exceed one. The cumulative incidences obey this: at t = 5 they add to 0.9179. The two complements, each computed with the other cause censored, add to 1.4088 over the two thousand studies, against 1.4090 closed. They pass one at t = 2.812. A table reporting the “risk” of each cause this way reports that more than a hundred and forty per cent of the cohort had an event.
The reason is double counting, and it is visible in the construction. A subject who died of the competitor at t = 2 is counted, by the first complement, as still able to have the event of interest; and by the second, as having had the competitor. Both complements claim the same person’s future.
Every subject in exactly one place
The cumulative incidence partitions the cohort instead of multiplying it forward.
At t = 5 the three shares are 0.3672, 0.5507 and 0.0821. The dashed complement, at 0.6321, runs up through the competitor’s band and into the event-free one: of the subjects it counts as having had the event of interest, a large share have already had the other event and cannot have this one.
The Aalen–Johansen estimator is the empirical form of that partition. At each event time it adds, to the running incidence of the cause that occurred,
the chance of being event-free of everything just before that time, times the share of those at risk who had this cause then. Kaplan–Meier for cause one uses the same ratio and multiplies it by the wrong thing: its own complement, which ignores that the other cause has been emptying the cohort. The difference between the two constructions is one factor, and it is the whole of the error.
Because the partition is built into the estimator, it holds on every dataset rather than on average. The event-free Kaplan–Meier estimate plus the two Aalen–Johansen incidences equals one at every step of every one of the two thousand studies, with a largest gap of . That is the second route here: the estimator’s unbiasedness is a count, and its coherence is an identity.
The same censoring is right for the hazard
None of this means censoring the competing event was a mistake in itself. For one quantity it is exactly right, and the distinction is the most useful thing to carry from this essay.
The cause-specific hazard is the rate at which subjects still at risk have this event. For a subject still at risk, a competing event is just another reason to stop being at risk, so treating it as censoring is the correct bookkeeping, and the Nelson–Aalen estimate of the cumulative hazard reads 0.9990 at t = 5 against a true of 1.0000. Every cause-specific hazard in a competing-risks analysis is estimated this way, and properly.
The mistake is only in the next step: turning the hazard into a risk by . Without a competitor, a cumulative hazard and a risk are transforms of each other. With one, they are not. At t = 5 the cumulative hazard is 1.0000, and minus the logarithm of one minus the actual risk is 0.4575. The hazard describes a rate among the subjects left; the risk describes the cohort; and the competitor is what makes those two groups different.
So “censor the competitor” is a correct instruction with a scope. It is correct for estimating a cause-specific hazard, for testing whether a cause-specific hazard differs between arms, and for fitting a model to one. It is incorrect for any statement that begins “the chance of having the event by”.
Preventing the other cause raises this one’s risk
The most practical consequence is also the least intuitive, and it is where the two numbers do not merely differ in size but move in different directions.
Hold the event of interest’s hazard at 0.2 and change only the competitor’s. At a competing hazard of 0.6 the cumulative incidence at t = 5 is 0.2454; at 0.3, 0.3672; at 0.15, 0.4721; at 0.075, 0.5434; and with no competitor, 0.6321. Subjects no longer removed by the other cause are still present to have this one, so the real risk of this event rises as the other cause is prevented. The Aalen–Johansen estimates, eight hundred studies at each setting, sit on the line.
One minus Kaplan–Meier is flat at 0.6321 across the whole sweep, counted at 0.6348, 0.6307, 0.6307, 0.6319 and 0.6325. It cannot see the change, because the change is in who is left rather than in the rate among those left. A treatment that halved the competing hazard from 0.3 to 0.15 would raise the actual five-year risk of the event of interest by 0.1050 and move the commonly reported number by nothing.
This is not a curiosity. A trial of a drug that reduces deaths from one cause will, in a population where several causes compete, raise the cumulative incidence of the others. Read through the cumulative incidence, that increase is real, and it is caused by success. Read through the complement of Kaplan–Meier it does not exist, and a report using the complement for the primary outcome and the cumulative incidence for a safety outcome would be comparing an effect measured in one world against a harm measured in another.
How large the overstatement is
The factor by which the complement exceeds the real risk has a closed form, and its shape settles when the distinction can safely be ignored.
It starts at one for every competitor, since before any events there is nothing to double-count, and grows with time and with the competitor’s hazard. At t = 5 it is 1.163 for a competing hazard of 0.075, 1.339 for 0.15, 1.722 for 0.3 and 2.576 for 0.6. A weak competitor and an early reading make it small; a strong competitor and a late reading make the complement more than twice the risk. Neither factor involves the sample size, so no study is large enough to shrink it: this is a difference of estimands, and precision only makes the wrong one more exact.
What is proved here, and what is only counted
The closed forms are exact for constant hazards, and the specific factors — 1.722 at t = 5, the crossing at 2.812 — belong to this world. The direction does not, and it needs no assumption at all. With the cause-specific cumulative hazard, the complement of Kaplan–Meier converges to and the cumulative incidence is . The chance of being free of every event, , can never exceed , so the complement is at least the incidence for any hazards whatever, with equality only when the competitor never acts. Independence of the latent times is not needed for the inequality. It is needed only to interpret the complement as the risk with the competitor removed — and without it the complement is a number with no world behind it.
The estimators’ behaviour is counted: two thousand studies of three hundred, standard errors under a thousandth, and every estimate within four standard errors of its closed form at every time checked. The partition identity is exact on every dataset. What is not examined here is the variance of the Aalen–Johansen estimator or the coverage of its interval, which has its own delta-method standard error and would face the same end-of-curve problem the interval at the end of a Kaplan–Meier curve runs into.
Which number answers which question
Three quantities are estimable from a competing-risks dataset, and each is the right answer to one question.
The cumulative incidence answers how many subjects in this population, with every cause acting, will have this event by a given time. It is the number for prognosis, planning and cost. An estimator that converges cleanly on the wrong estimand is a recurring shape in statistics, and here the right estimand is almost always this one.
The cause-specific hazard answers how fast this event happens among those still at risk of it. It is the number for mechanism — whether a treatment acts on this cause — and censoring the competitor is correct for it.
The complement of Kaplan–Meier answers what the risk would be with the competitor removed and with the latent times independent. Both conditions are assumptions about a world that was not observed, and the number should be reported, if at all, under a name that says so.
What to look for in a reported risk
A reader without the data can usually tell which number a report contains, and three checks cover most cases.
The methods sentence about other deaths. A phrase of the form “patients who died of other causes were censored at the time of death”, sitting beside a “cumulative risk” or “five-year risk” read off a curve, is the complement. A phrase naming the cumulative incidence function, the Aalen–Johansen estimator, or a competing-risks analysis is the incidence. When neither sentence is present the curve is almost always the complement, since it is what a standard survival routine produces by default.
Whether the causes add up. A table that reports a risk for each of several causes, from separate curves, can be summed by hand. If the causes are exhaustive and the sum at the last time exceeds the share of the cohort that had any event — or exceeds one — the figures are complements. In the world above the sum passes that bound from 2.812 years onward, so a table read at five years would show it plainly, and one read at two years would not.
How old and how ill the cohort is. The overstatement grows with the competitor’s strength and with time, so it is small in a young population followed briefly and large in an elderly one followed for a decade. A complement reported for a trial in patients over eighty, over ten years, is a number about a population with most of its deaths removed, and it can be several times the risk any of those patients faced. The same estimator that recovers the curve without competing events is answering a question nobody in that trial was in a position to ask.
The failure is not using the third. It is printing the third under the name of the first, which is what “five-year risk” under a Kaplan–Meier curve does, and which the data that stops early warned of in a simpler form: forcing a subject into a slot the data never gave them.
Where this goes next
The next question is about regression with competing causes. A covariate can be put into a model for the cause-specific hazard, or into a model for the cumulative incidence directly — the Fine–Gray subdistribution hazard — and the preventing-the-other-cause result above says the two can disagree in sign. A treatment with no effect at all on the cause-specific hazard of this event, and a strong effect on the competitor’s, has a cause-specific hazard ratio of exactly one for this event and a subdistribution hazard ratio well above one. Which of those is “the effect of the treatment on this outcome” is the same estimand question as here, one level up, and it is distinct from this essay’s because it is about how a comparison between groups is summarised rather than about what one group’s curve means.
That summary already has a known trap of its own in the simplest two-arm case, before a competitor is involved: a hazard ratio is an average whose weights the study’s follow-up chooses.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A horizon chosen after looking — both name censoring, estimand, kaplan–meier
- The mechanism the data cannot see — both name closed form, estimand, non-identifiability
- A space is not a relation — both name closed form, non-identifiability
- A weight fitted to balance — both name closed form, estimand
- An imputation model the analysis does not contain — both name closed form, estimand
- Dropping the incomplete rows — both name closed form, estimand
Named objects
A flat tag is an object no other essay names yet.
Aalen–JohansenCause specific hazardCensoringClosed formCompeting risksCumulative incidenceDependent censoringEstimandHazardKaplan–MeierNelson–AalenNon-identifiabilityRisk setSurvival analysis