The interval at the end of the curve
Worth reading first: The data that stops early · What the 95% refers to.
A survival curve is almost never read in the middle. It is read at its end: survival at five years, the median if the curve reaches it, the last step before the follow-up runs out. That is where the number in the abstract comes from, and it is where the interval printed around the curve is weakest.
Across 4,000 simulated studies of forty subjects, with the true survival known in closed form, the interval most software prints first — the Kaplan–Meier estimate plus and minus 1.96 of Greenwood’s standard errors — contains the truth 89.7% of the time at five years. By then an average of 3.3 subjects are still under observation. The same variance, carried on a log–log scale and mapped back, contains the truth 94.8% of the time at the same point, from the same studies.
The picture above is one study, picked to show what the failure looks like, and it proves nothing about how often it happens. The rest of this essay is the count.
Where a survival curve is actually read
The construction of the estimator established that Kaplan–Meier recovers the truth at every level of censoring tried, and that its uncertainty grows to the right, steeply at the end, because each late event is divided by a small risk set. That essay checked the estimate against its own standard error. It did not ask whether the interval built from that standard error covers at the stated rate, and the two questions come apart exactly where the curve is used.
The world here is the one the field has used throughout: exponential survival at rate 0.35 a year, dropout at 0.15 a year, follow-up to six years, and cohorts of forty. The truth at five years is , which is 0.174. At every quarter from a quarter-year to five and a half years, each of the 4,000 studies builds both intervals and records whether each one contains the true value. With 4,000 studies every coverage reading carries a Monte Carlo standard error of about half a point, so a gap of two points between the intervals at one time is several standard errors and not noise.
Two things make the end of the curve different from its middle, and they have different consequences. The first is the size of the risk set: fewer subjects left means a noisier estimate, and Greenwood’s formula already knows that, which is why the interval widens — the same bookkeeping that lets a censored subject count for exactly as long as it was watched and no longer. The second is the edge: an estimate near zero cannot be symmetric about itself, and nothing in the plain interval knows that at all.
A symmetric interval, pressed against an edge
The plain interval is symmetric by construction. It puts equal width above and below the estimate whatever the estimate is. For a quantity confined to [0, 1] that is harmless in the middle and wrong near either end, where the distribution of the estimate is pushed against a wall and piles up on the side away from it.
The survival curve visits both walls. At a quarter of a year the true survival is 0.916 and most studies have seen only a handful of events, so the estimate sits close to one. At five and a half years the true survival is 0.146 and the estimate sits close to zero. The plain interval leaves [0, 1] at both ends, on opposite sides.
At a quarter of a year the upper limit exceeds one in 55.1% of studies. At five years the lower limit is below zero in 33.3%, and by five and a half years in 46.2%. Software commonly clips the interval to [0, 1] before printing it, which removes the symptom from the page and changes nothing about where the interval is. A limit of zero produced by clipping is not a statement the data supports; it is the printed form of a limit that was negative.
The escape is not in itself the loss of coverage — an interval running from −0.02 to 0.31 still contains 0.174 — but it is the visible sign of a shape that is wrong, and the shape is what costs coverage. It is also the kind of failure that a comparison of widths hides: the narrowest interval on a table is often the one that misses, and an interval that runs off the bottom of the scale looks, if anything, generous.
The width is nearly right, and the place is wrong
The natural diagnosis is that Greenwood’s standard error is too small at the end of the curve, and the natural repair is a wider interval. The counts say otherwise, and the difference matters because a wider symmetric interval would not have fixed it.
The simulation keeps every study’s estimate and every study’s Greenwood standard error, so the two can be compared directly. At five years the estimates have a standard deviation of 0.0796 across studies, and the root-mean-square of the standard errors the studies reported is 0.0761. Greenwood is short by about four per cent at that point — a real shortfall, and a small one. At a quarter of a year the two read 0.0443 and 0.0434, closer still, and the plain interval’s coverage there is worse.
What differs between the middle and the ends is the skewness of the estimate. At two years it is −0.089, and the interval covers 94.1%. At a quarter of a year it is −0.501, the estimate’s distribution leaning towards the lower values with its bulk crowded against one, and the coverage is 84.0%. At five years the skew has turned positive, 0.270, the bulk crowded against zero with a long upper tail, and the coverage is 89.7%.
A symmetric interval with the right width around a skewed estimate misses more often on the long side and less often on the short side, and the two do not cancel. That is the same mechanism the two intervals for a return level measured in a different field, where 99.24% of a symmetric interval’s misses sat entirely on one side of the truth. There the repair was an interval that moved rather than one that grew, and the same is true here.
The same failure with nobody censored
Survival curves are a specialised object, and it would be easy to read the failure as something about censoring. It is not, and there is a clean way to show it that shares no simulation with the count above.
With nobody censored, the Kaplan–Meier estimate telescopes into the plain proportion of subjects still event-free, and Greenwood’s variance telescopes with it: is exactly . The plain Greenwood interval is then, term for term, the Wald interval for a binomial proportion. The identity can be checked at every step of every draw, and across 10,000 uncensored cohorts the largest gap between the two variances is — arithmetic rounding, not a difference.
That makes the plain interval’s coverage without censoring computable exactly, as a finite sum over the forty-one possible counts of survivors, and it can be set against the count. The worst disagreement between the exact sum and 10,000 counted cohorts, over twenty-two times, is 2.08 standard errors.
The sawtooth is familiar from the interval for a proportion, whose coverage does not rise smoothly with anything but jumps as the true value crosses each limit a whole-number count can produce — the same effect that makes a sample of twenty cover worse than a sample of nineteen. Its lowest point near the start of this curve is the Wald interval failing near an edge, as it always does.
So the plain Greenwood interval inherits every defect of the Wald interval, and censoring adds a second route to the same trouble: it thins the risk set at the end, where the estimate is already near zero. Censoring does not cause the failure. It moves the curve’s small-count region to where the curve is read.
A transform that has nowhere to go but inside
The log–log interval uses exactly the same variance sum. What it changes is the scale on which the interval is symmetric. It builds the interval for , which runs over the whole real line as the survival runs over (0, 1), and maps the two limits back. The result is compact:
Raising a number between zero and one to a positive power keeps it between zero and one, so the interval cannot leave [0, 1] at all, and the limits are asymmetric about the estimate in the direction the edge requires — shorter towards the wall, longer away from it. Near zero, the interval stretches upward; near one, downward.
It holds its level for most of the curve: 96.5% at half a year, 95.5% at two years, 94.8% at five years. Across the quarters from half a year to five years it is never more than two points from 95%, which is the standard set for it here, and the plain interval, held to the same standard, fails it at five years by more than five points.
The log–log interval has two failures of its own, and both are honest ones. At a quarter of a year, 3.52% of studies have not yet seen a single event. The estimate is exactly one, the log–log scale has no finite value for it, and the interval is undefined; those studies count as misses, which is why the interval reads 92.5% there. At five and a half years, 4.95% of studies have already seen their last subject fail, the estimate is exactly zero, and the same thing happens at the other end — 92.5% again. An interval that declines to exist when the estimate sits on the boundary is not wrong. It is reporting that no interval of this family can be drawn from that study, which is true, and a zero-width plain interval at the same point is the far worse report, a claim of certainty from a study that has none.
It is the number left, not the number that started
The failure at the end of the curve is easy to misattribute to a small study. The count by cohort size says the relevant quantity is not how many subjects began but how many are still there at the time being read.
At twenty subjects, 1.6 remain at five years on average, and both intervals fail: 85.2% for the plain and 85.7% for the log–log. No transform repairs an interval drawn from one or two people, and the essay should not suggest otherwise. At three hundred and twenty subjects, 26.3 remain, the plain interval covers 94.5% and the log–log 95.0%.
The comparison that separates the two readings sits across the rows. A study of forty read at two years has 14.8 subjects still at risk and its plain interval covers 94.1%. A study of one hundred and sixty read at five years has 13.2 at risk and covers 93.0% — four times the cohort and much the same answer, because it is being read at a point with much the same number left and a lower survival. The size of a trial printed in its abstract says little about how well its five-year figure is determined; the numbers-at-risk row under the curve says it directly.
Pooling every study at every time by how many were still being watched gives the same message from a different cut, and it has to be read with care. With four or fewer under observation the plain interval covers 85.8% and the log–log 92.8%. That is a conditional reading — a study with few left at a given time is a study that happened to lose subjects fast, which is itself informative about its estimate — and it is the same distinction a marginal coverage guarantee draws against a conditional one. It is reported here as a description of where the misses concentrate, not as the coverage any one study can claim.
Reading along a band is reading many intervals
Every interval so far has been read at one time. That is not how a survival figure is read. A reader looks along the curve, sees where the shaded band excludes some value or where two bands separate, and draws a conclusion from the place along the curve where that happens. Each vertical slice of the band is a 95% interval; the band as a whole is not a 95% statement about the curve.
Checked at a single time in the middle of that stretch, the plain band misses 6.7% of the time and the log–log band 5.0%. Checked at ten evenly spaced times, they miss somewhere on 33.1% and 23.2% of studies. Checked everywhere — which can be computed exactly rather than approached, because between two event times the band is flat while the truth only falls, so a miss inside a stretch can happen only at one of its two ends — they miss somewhere on 49.4% and 37.7%.
So a band that is correct at every point, in the sense of the log–log band above, fails to contain the true curve somewhere along its length on more than a third of honest studies of forty. That is the arithmetic of twenty honest analyses of nothing turned sideways: many correct statements, read together and reported at the most striking one. The slices are highly correlated, which is why the rate climbs to 37.7% rather than to anything like the near-certainty that the same number of independent intervals would give, and it still climbs a long way.
Which of these are theorems and which are counts
The claims in this essay are of three kinds, and they should not be read as equally firm.
Proved. With nobody censored, Greenwood’s variance equals the binomial variance exactly, so the plain interval is the Wald interval; the counted gap of checks the implementation of that identity, not the identity. The log–log interval cannot leave [0, 1], because its limits are powers of a number in that range. A band misses the curve somewhere at least as often as it misses at any one time, which needs no simulation.
Counted, against a known truth. Every coverage figure — 89.7% and 94.8% at five years, the escape rates, the rates by cohort size, the 49.4% and 37.7% — is a count over simulated studies from one stated world, with Monte Carlo standard errors near half a point at 4,000 studies. They are statements about that world. Another survival shape, another dropout rate, or another cohort size would move them, and the direction of the plain interval’s failure near each edge is the part expected to carry over, because it follows from the Wald interval’s behaviour near an edge.
Not established here. The counts put the log–log interval within two points of 95% from half a year to five years in this world. They do not show that the log–log interval is the best choice among the transforms in use — the arcsine-square-root interval and the likelihood-ratio interval are both defensible and neither was counted. And nothing here establishes a rule for how many subjects at risk are enough; the counts only show that the answer is set by that number rather than by the cohort size.
What a reader can check on a published curve
Three things follow directly, none of which require the data.
Which interval was drawn. Packages differ in their default — the log–log interval, the plain one, or a log interval that behaves between the two — and a figure’s caption rarely says which. A band whose limits are symmetric about the step curve near the end is the plain one. A band that stops above zero and stretches upward away from a low estimate is the log–log one. The figure alone usually says which.
The number at risk at the time quoted. The coverage at five years is governed by how many subjects are left at five years, and the risk table beneath the curve prints it. A five-year figure resting on a single digit at risk is not well determined whatever interval surrounds it, and a symmetric interval around it misplaces what little it knows.
Whether a conclusion was read off the band at a place chosen by looking. A statement that “the curves separate after three years” is a statement about the band at a time selected from the figure, and the pointwise band does not support it at its nominal level. That does not make the statement false. It makes its evidence weaker than the shading suggests: read along its whole length, even the log–log band misses somewhere seven and a half times as often as it promises to miss at one time.
Where this goes next: a band that misses five per cent of the time somewhere
The last count opens the question the next essay on censoring should answer, and it is distinct from everything above. A pointwise band is an honest answer to a question about one time. Most readers are asking a question about the whole curve — does the true curve lie inside this shaded region everywhere in the follow-up — and that question has its own constructions: bands built so that the probability of a miss anywhere in a stated range is 5%, of which the Hall–Wellner and equal-precision bands are the classical pair.
What such a band costs is the thing worth measuring, and it cannot be read off this essay. It has to be wider than the pointwise band everywhere, and the widening is not uniform: it is spent disproportionately at the ends, where the pointwise band was already widest and where the curve is read. The measurement that would settle whether it is worth using is its width at five years against the pointwise band’s, in the same world, next to the counted rate at which it misses the curve somewhere — which should come out at 5% where this essay’s log–log band reads 37.7%. A censoring mechanism that the data cannot see is the other direction the field takes from here, and it sits upstream of every interval above: all of them assume the dropout says nothing about the event.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name binomial proportion, confidence interval, coverage, monte carlo, multiple comparisons
- The same draws for both methods — both name binomial proportion, confidence interval, coverage, monte carlo, wald interval
- A simulation that stops when it looks settled — both name binomial proportion, confidence interval, coverage, monte carlo
- An interval that covers and says nothing — both name binomial proportion, confidence interval, coverage, wald interval
- Intervals for the findings — both name confidence interval, coverage, monte carlo, multiple comparisons
- Robust is not free — both name confidence interval, coverage, monte carlo, wald interval
Named objects
A flat tag is an object no other essay names yet.
Binomial proportionCensoringConditional coverageConfidence intervalCoverageGreenwood's formulaKaplan–MeierMonotone transformationMonte CarloMultiple comparisonsRisk setSurvival curveWald interval