When the data stops early

One minus Kaplan–Meier is not a risk

With two ways for observation to end, one minus Kaplan–Meier for one cause reads 0.6318 at t = 5 where the chance of actually having had that event is 0.3670. Added across the two causes, the complements reach 1.4088 — more than the whole cohort. Nothing is estimated badly: the complement estimates, correctly, the risk in a world where the other cause does not exist.

Worth reading first: The data that stops early.

A report of a long study says the five-year risk of the event it studied was sixty-three per cent. The number came from a Kaplan–Meier curve for that event, subtracted from one, with every subject who had some other ending event — a death from an unrelated cause, a transplant that removes the risk, a switch to a treatment that ends follow-up for this outcome — recorded as censored at that time. It is one of the most common constructions in applied survival analysis, and in the world measured below the share of the cohort that actually had the event by five years was thirty-seven per cent.

The risk of one cause, estimated two ways. Two causes of an ending event with constant hazards 0.2 (the one of interest) and 0.3 (the competitor), random dropout at 0.1 and follow-up to 6. The lower line is the cumulative incidence, (0.2/0.5)(1 − e^(−0.5t)), the chance of actually having had this event by t; the dots on it are the Aalen–Johansen estimate over 2000 studies of 300, 0.3670 at t = 5 against 0.3672. The upper line is 1 − e^(−0.2t), and the dots on it are one minus Kaplan–Meier with the competing event treated as censoring: 0.6318 at t = 5 against 0.6321. The second is larger by a factor of 1.722 at t = 5, and it is not an error of estimation. It estimates, correctly, the risk in a population where the competing cause does not exist.
Fig. 1 The chance of having had the event of interest by each time, as a closed curve with the Aalen–Johansen estimate on it, and one minus Kaplan–Meier with the other cause censored, as a higher closed curve with its own estimates.

Two ways out, and one way to be counted

The world is as simple as a world with competing risks can be. Each subject faces two causes of an ending event with constant hazards, 0.2 for the event of interest and 0.3 for the competitor, and random dropout at 0.1, with follow-up to six. Three hundred subjects a study, two thousand studies.

The quantity a patient, a planner or a payer means by “risk” is the cumulative incidence: the share of the cohort that has had this particular event by time t. With constant hazards it has a closed form,

F1(t)=λ1λ1+λ2(1e(λ1+λ2)t),F_1(t) = \frac{\lambda_1}{\lambda_1 + \lambda_2}\left(1 - e^{-(\lambda_1 + \lambda_2)t}\right),

because the chance of any ending event by t is 1e(λ1+λ2)t1 - e^{-(\lambda_1+\lambda_2)t} and a fixed share λ1/(λ1+λ2)\lambda_1/(\lambda_1+\lambda_2) of those endings are this cause. At t = 5 it is 0.3672. The Aalen–Johansen estimator, which is built to estimate exactly that, reads 0.3670 averaged over the two thousand studies, with a standard error of 0.0007.

One minus Kaplan–Meier, with the competitor censored, converges to something else: 1eλ1t1 - e^{-\lambda_1 t}, which at t = 5 is 0.6321. It reads 0.6318. The ratio between the two numbers is 1.722, and it is not an estimation error. The estimator has converged, precisely, on a well-defined quantity that is not the risk.

A risk in a world nobody lives in

What 1eλ1t1 - e^{-\lambda_1 t} describes is the chance of having the event by t if nobody could have the competing event first. Censoring a subject at their competing event tells Kaplan–Meier that they were still capable of having the event of interest afterwards, and that their chance of doing so was the same as that of the subjects still being watched. The estimator then multiplies forward as though the competitor had been removed from the world and its victims returned to the risk set.

That is a meaningful number for some questions. It is the quantity behind “what would the risk of this disease be if the other cause were eliminated”, which is a real question in aetiology. But it answers that question only under two assumptions, and neither can be checked. The competing cause has to be removable without changing anything else about the subjects’ hazards; and the latent time to this event for a subject who died of the other cause has to be independent of that death. The second is exactly the non-identifiability a dropout the data cannot see turns on: every competing-risks dataset is consistent with independent latent times and with dependence of any strength, and the record cannot say which.

The cumulative incidence needs neither assumption. It describes the world the cohort was actually in, with the competitor present, and it is a function of the recorded data alone. The first number asks a counterfactual question and hides the fact; the second answers a factual one.

A time that does not exist is not missing

The instruction to censor the competing event borrows its authority from ordinary censoring, and the borrowing does not survive a close look. A subject censored because the study ended was still event-free when observation stopped: the event time exists, lies beyond the censoring time, and would have been seen with longer follow-up. Its value is unrecorded, and Kaplan–Meier is a correct way of using the partial information that remains.

A subject who died of the competing cause has no later time for the event of interest to be unrecorded at. There is no longer follow-up that would reveal it. Treating that subject as censored treats a quantity that does not exist as a quantity that is merely unseen, and then asks the estimator to fill it in from the subjects who are still alive — exactly the move an imputation makes for a value that is absent, applied to a value that was never going to be there.

The distinction matters because the language of missing data then misleads. Whether censoring is “at random” is a question about a value that exists and was not recorded, the question the three missingness mechanisms are built to ask. For a competing death that question has no object. The latent time to the event of interest after a competing death is a modelling device, and every conclusion that depends on its distribution depends on something no cohort, however large and however completely followed, contains.

That is also why no amount of care in follow-up repairs the complement. A trial that loses no subject at all to administrative censoring and records every competing death with its exact date still produces, from one minus Kaplan–Meier, the risk in a population where those deaths did not happen. Complete data make the estimate precise. They do not make the population real.

More than everybody

The clearest sign that the complement is not a risk is arithmetic that a risk cannot do.

Two risks that add to more than certainty. The same estimates for both causes, added. Every subject has at most one ending event, so the chance of having had cause one plus the chance of having had cause two can never exceed one, and the cumulative incidences respect that: they add to 0.918 at t = 5, the chance of having had either (0.918 counted). One minus Kaplan–Meier for each cause, the other censored, adds to 1.409 closed and 1.409 counted, and passes one at t = 2.812. A table reporting both numbers side by side reports more than a hundred per cent of the cohort having an event.
Fig. 2 The two causes’ estimates added together: one minus Kaplan–Meier for each cause with the other censored, against the two cumulative incidences, with the line at one.

Every subject has at most one ending event, so the share who had cause one plus the share who had cause two cannot exceed the share who had either, and that cannot exceed one. The cumulative incidences obey this: at t = 5 they add to 0.9179. The two complements, each computed with the other cause censored, add to 1.4088 over the two thousand studies, against 1.4090 closed. They pass one at t = 2.812. A table reporting the “risk” of each cause this way reports that more than a hundred and forty per cent of the cohort had an event.

The reason is double counting, and it is visible in the construction. A subject who died of the competitor at t = 2 is counted, by the first complement, as still able to have the event of interest; and by the second, as having had the competitor. Both complements claim the same person’s future.

Every subject in exactly one place

The cumulative incidence partitions the cohort instead of multiplying it forward.

Every subject is in exactly one place. The cohort partitioned at each time into three shares that add to one: had the event of interest, F₁(t); had the competing event, F₂(t); neither yet, e^(−0.5t). At t = 5 they are 0.3672, 0.5507 and 0.0821. The dashed line is one minus the Kaplan–Meier limit for the event of interest with the competitor censored, 0.6321 at t = 5: it runs up through the competitor's band and into the event-free one, claiming subjects who have already had the other event, and whose chance of later having this one is nought.
Fig. 3 The cohort at each time divided into three shares that add to one — had the event of interest, had the competing event, neither yet — with one minus Kaplan–Meier for the event of interest dashed across them.

At t = 5 the three shares are 0.3672, 0.5507 and 0.0821. The dashed complement, at 0.6321, runs up through the competitor’s band and into the event-free one: of the subjects it counts as having had the event of interest, a large share have already had the other event and cannot have this one.

The Aalen–Johansen estimator is the empirical form of that partition. At each event time it adds, to the running incidence of the cause that occurred,

F^1(t)=titS^(ti)d1ini,\hat F_1(t) = \sum_{t_i \le t} \hat S(t_i^-)\,\frac{d_{1i}}{n_i},

the chance of being event-free of everything just before that time, times the share of those at risk who had this cause then. Kaplan–Meier for cause one uses the same ratio d1i/nid_{1i}/n_i and multiplies it by the wrong thing: its own complement, which ignores that the other cause has been emptying the cohort. The difference between the two constructions is one factor, and it is the whole of the error.

Because the partition is built into the estimator, it holds on every dataset rather than on average. The event-free Kaplan–Meier estimate plus the two Aalen–Johansen incidences equals one at every step of every one of the two thousand studies, with a largest gap of 2.9×10152.9\times10^{-15}. That is the second route here: the estimator’s unbiasedness is a count, and its coherence is an identity.

The same censoring is right for the hazard

None of this means censoring the competing event was a mistake in itself. For one quantity it is exactly right, and the distinction is the most useful thing to carry from this essay.

The same censoring is right for the hazard and wrong for the risk. The cumulative hazard of the event of interest, λ₁t = 0.2t, and the Nelson–Aalen estimate of it with the competing event treated as censoring — 0.1998 at t = 1, 0.3996 at t = 2, 0.6011 at t = 3, 0.8014 at t = 4, 0.9990 at t = 5, over 2000 studies. Among the subjects still at risk the competing event is exactly a reason to stop being observed, so the manoeuvre that fails for the risk is correct here. The lower curve is what the risk would imply if it were one minus the exponential of a hazard, −log(1 − F₁(t)), 0.4575 at t = 5 against the hazard's 1.0000: with a competitor, a hazard and a risk are no longer transforms of each other.
Fig. 4 The cumulative hazard of the event of interest with the Nelson–Aalen estimate on it, which censors the competitor, against minus the logarithm of one minus the cumulative incidence.

The cause-specific hazard is the rate at which subjects still at risk have this event. For a subject still at risk, a competing event is just another reason to stop being at risk, so treating it as censoring is the correct bookkeeping, and the Nelson–Aalen estimate of the cumulative hazard reads 0.9990 at t = 5 against a true λ1t\lambda_1 t of 1.0000. Every cause-specific hazard in a competing-risks analysis is estimated this way, and properly.

The mistake is only in the next step: turning the hazard into a risk by 1eΛ1 - e^{-\Lambda}. Without a competitor, a cumulative hazard and a risk are transforms of each other. With one, they are not. At t = 5 the cumulative hazard is 1.0000, and minus the logarithm of one minus the actual risk is 0.4575. The hazard describes a rate among the subjects left; the risk describes the cohort; and the competitor is what makes those two groups different.

So “censor the competitor” is a correct instruction with a scope. It is correct for estimating a cause-specific hazard, for testing whether a cause-specific hazard differs between arms, and for fitting a model to one. It is incorrect for any statement that begins “the chance of having the event by”.

Preventing the other cause raises this one’s risk

The most practical consequence is also the least intuitive, and it is where the two numbers do not merely differ in size but move in different directions.

Preventing the other cause raises this one's risk. The risk of the event of interest by t = 5 as only the competing cause's hazard is changed; the hazard of the event of interest stays at 0.2 throughout. The cumulative incidence rises from 0.2454 at λ₂ = 0.6 to 0.4721 at 0.15 and 0.6321 at nought, because subjects no longer removed by the competitor are there to have this event — Aalen–Johansen, dots, over 800 studies apiece. One minus Kaplan–Meier is flat at 0.6321 at every λ₂, counted 0.6325, 0.6319, 0.6307, 0.6307, 0.6348. A treatment that halved the competing hazard would raise the real risk by 0.1050 and move the commonly reported number by nothing.
Fig. 5 The chance of the event of interest by t = 5 as only the competing cause’s hazard is changed, with the event of interest’s own hazard held fixed: cumulative incidence and one minus Kaplan–Meier, closed lines and counted dots.

Hold the event of interest’s hazard at 0.2 and change only the competitor’s. At a competing hazard of 0.6 the cumulative incidence at t = 5 is 0.2454; at 0.3, 0.3672; at 0.15, 0.4721; at 0.075, 0.5434; and with no competitor, 0.6321. Subjects no longer removed by the other cause are still present to have this one, so the real risk of this event rises as the other cause is prevented. The Aalen–Johansen estimates, eight hundred studies at each setting, sit on the line.

One minus Kaplan–Meier is flat at 0.6321 across the whole sweep, counted at 0.6348, 0.6307, 0.6307, 0.6319 and 0.6325. It cannot see the change, because the change is in who is left rather than in the rate among those left. A treatment that halved the competing hazard from 0.3 to 0.15 would raise the actual five-year risk of the event of interest by 0.1050 and move the commonly reported number by nothing.

This is not a curiosity. A trial of a drug that reduces deaths from one cause will, in a population where several causes compete, raise the cumulative incidence of the others. Read through the cumulative incidence, that increase is real, and it is caused by success. Read through the complement of Kaplan–Meier it does not exist, and a report using the complement for the primary outcome and the cumulative incidence for a safety outcome would be comparing an effect measured in one world against a harm measured in another.

How large the overstatement is

The factor by which the complement exceeds the real risk has a closed form, and its shape settles when the distinction can safely be ignored.

How far the complement overstates, by how strong the competitor is. The factor (1 − e^(−λ₁t)) / [(λ₁/λ)(1 − e^(−λt))] by which one minus Kaplan–Meier exceeds the cumulative incidence, for the event of interest at hazard 0.2 and a competitor at 0.075, 0.15, 0.3, 0.6. All four start at one — before anybody has had either event there is nothing to double-count — and grow with time. At t = 5 they are 1.163, 1.339, 1.722, 2.576. The overstatement is a property of the competitor and the time read, not of the sample, so no study size reduces it.
Fig. 6 One minus Kaplan–Meier as a multiple of the cumulative incidence, over time, for four strengths of the competing cause.

It starts at one for every competitor, since before any events there is nothing to double-count, and grows with time and with the competitor’s hazard. At t = 5 it is 1.163 for a competing hazard of 0.075, 1.339 for 0.15, 1.722 for 0.3 and 2.576 for 0.6. A weak competitor and an early reading make it small; a strong competitor and a late reading make the complement more than twice the risk. Neither factor involves the sample size, so no study is large enough to shrink it: this is a difference of estimands, and precision only makes the wrong one more exact.

What is proved here, and what is only counted

The closed forms are exact for constant hazards, and the specific factors — 1.722 at t = 5, the crossing at 2.812 — belong to this world. The direction does not, and it needs no assumption at all. With Λ1\Lambda_1 the cause-specific cumulative hazard, the complement of Kaplan–Meier converges to 1eΛ1(t)1 - e^{-\Lambda_1(t)} and the cumulative incidence is 0tS(s)dΛ1(s)\int_0^t S(s)\,d\Lambda_1(s). The chance of being free of every event, S(s)=eΛ1(s)Λ2(s)S(s) = e^{-\Lambda_1(s) - \Lambda_2(s)}, can never exceed eΛ1(s)e^{-\Lambda_1(s)}, so the complement is at least the incidence for any hazards whatever, with equality only when the competitor never acts. Independence of the latent times is not needed for the inequality. It is needed only to interpret the complement as the risk with the competitor removed — and without it the complement is a number with no world behind it.

The estimators’ behaviour is counted: two thousand studies of three hundred, standard errors under a thousandth, and every estimate within four standard errors of its closed form at every time checked. The partition identity is exact on every dataset. What is not examined here is the variance of the Aalen–Johansen estimator or the coverage of its interval, which has its own delta-method standard error and would face the same end-of-curve problem the interval at the end of a Kaplan–Meier curve runs into.

Which number answers which question

Three quantities are estimable from a competing-risks dataset, and each is the right answer to one question.

The cumulative incidence answers how many subjects in this population, with every cause acting, will have this event by a given time. It is the number for prognosis, planning and cost. An estimator that converges cleanly on the wrong estimand is a recurring shape in statistics, and here the right estimand is almost always this one.

The cause-specific hazard answers how fast this event happens among those still at risk of it. It is the number for mechanism — whether a treatment acts on this cause — and censoring the competitor is correct for it.

The complement of Kaplan–Meier answers what the risk would be with the competitor removed and with the latent times independent. Both conditions are assumptions about a world that was not observed, and the number should be reported, if at all, under a name that says so.

What to look for in a reported risk

A reader without the data can usually tell which number a report contains, and three checks cover most cases.

The methods sentence about other deaths. A phrase of the form “patients who died of other causes were censored at the time of death”, sitting beside a “cumulative risk” or “five-year risk” read off a curve, is the complement. A phrase naming the cumulative incidence function, the Aalen–Johansen estimator, or a competing-risks analysis is the incidence. When neither sentence is present the curve is almost always the complement, since it is what a standard survival routine produces by default.

Whether the causes add up. A table that reports a risk for each of several causes, from separate curves, can be summed by hand. If the causes are exhaustive and the sum at the last time exceeds the share of the cohort that had any event — or exceeds one — the figures are complements. In the world above the sum passes that bound from 2.812 years onward, so a table read at five years would show it plainly, and one read at two years would not.

How old and how ill the cohort is. The overstatement grows with the competitor’s strength and with time, so it is small in a young population followed briefly and large in an elderly one followed for a decade. A complement reported for a trial in patients over eighty, over ten years, is a number about a population with most of its deaths removed, and it can be several times the risk any of those patients faced. The same estimator that recovers the curve without competing events is answering a question nobody in that trial was in a position to ask.

The failure is not using the third. It is printing the third under the name of the first, which is what “five-year risk” under a Kaplan–Meier curve does, and which the data that stops early warned of in a simpler form: forcing a subject into a slot the data never gave them.

Where this goes next

The next question is about regression with competing causes. A covariate can be put into a model for the cause-specific hazard, or into a model for the cumulative incidence directly — the Fine–Gray subdistribution hazard — and the preventing-the-other-cause result above says the two can disagree in sign. A treatment with no effect at all on the cause-specific hazard of this event, and a strong effect on the competitor’s, has a cause-specific hazard ratio of exactly one for this event and a subdistribution hazard ratio well above one. Which of those is “the effect of the treatment on this outcome” is the same estimand question as here, one level up, and it is distinct from this essay’s because it is about how a comparison between groups is summarised rather than about what one group’s curve means.

That summary already has a known trap of its own in the simplest two-arm case, before a competitor is involved: a hazard ratio is an average whose weights the study’s follow-up chooses.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Aalen–JohansenCause specific hazardCensoringClosed formCompeting risksCumulative incidenceDependent censoringEstimandHazardKaplan–MeierNelson–AalenNon-identifiabilityRisk setSurvival analysis