When the data stops early

One test for a curve that might cross

The log-rank statistic and the differences in restricted mean survival at eleven horizons are one Gaussian vector under no effect, and its correlations are closed — from −0.461 at half a year to −0.949 at three, where the log-rank statistic is nearly the restricted mean itself. So one critical value, 2.374, prices the largest of all twelve. At 100 a side it keeps 93.6% of the log-rank test's power when the hazard is halved throughout and 98.5% of the best horizon's when the curves cross, where the log-rank test has none. Running both tests and keeping either rejection rejects 8.67% of trials with no effect.

Worth reading first: The data that stops early · The curve that survives censoring.

A horizon chosen after looking priced one freedom an analysis of survival curves has. A difference in restricted mean survival can be read at any horizon, and the horizon at which it looks most convincing is not a horizon fixed in advance. The essay showed that the differences at eleven horizons converge jointly to a Gaussian process whose correlation is closed. The largest of their absolute values then has a critical value that can be drawn from that process, 2.317, instead of the 1.96 a single horizon would use. Read against it, the best horizon holds its size and loses little power to the horizon nobody could have known to fix.

That is one way of not having to know in advance what shape a treatment’s effect will take. The log-rank test is another, and it is the one most trials pre-specify. It is the most powerful test when the hazards are proportional and is blind when the curves cross, because then early benefit and late harm cancel in its sum. The hazard ratio the follow-up chose showed what that cancellation does to an estimate. The obvious move is to compute both and keep the stronger, and the obvious objection is that keeping the stronger of two tests is a test of its own, with its own size. This essay builds that test properly, so its size is the stated one, and measures what it costs against each of its components in the shapes where each is best.

The log-rank statistic is a restricted mean, nearly

Every statistic in the comparison is a linear functional of the same thing: the difference between events counted and events expected in each arm, as time runs. The log-rank numerator adds that difference over the whole of follow-up with equal weight at each event. A difference in restricted means to a horizon aa weights the same differences by A(t,a)A(t, a), the area under the survival curve from tt to aa, divided by the share still at risk. Under no effect, with two equal arms, an at-risk fraction y(t)y(t) and a hazard λ\lambda, their correlation has a closed form:

corr⁡(ZLR,Za)=−∫0aA(t,a) λ dt∫0Ty λ dt  ⋅  ∫0aA(t,a)2 λ/y dt.\operatorname{corr}(Z_{\text{LR}}, Z_{a}) = -\frac{\int_0^a A(t,a)\,\lambda\,dt}{\sqrt{\int_0^T y\,\lambda\,dt \;\cdot\; \int_0^a A(t,a)^2\,\lambda / y\,dt}}.

The sign says only that a treatment with more events has a smaller restricted mean. The size is the interesting part.

How closely the log-rank statistic tracks the restricted mean at each horizon. The correlation, under no treatment effect, between the log-rank z and the z for the difference in restricted means to each horizon: closed form -0.461, -0.554, -0.629, -0.691, -0.744, -0.790, -0.830, -0.865, -0.896, -0.924, -0.949 at horizons 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.25, 2.5, 2.75, 3; counted on 6,000 null trials of 200 a side -0.463, -0.553, -0.627, -0.689, -0.743, -0.790, -0.831, -0.866, -0.897, -0.925, -0.950. The critical value for the largest of the eleven horizons alone is 2.317; adding the log-rank statistic raises it to 2.374.
Fig. 1 The correlation, under no treatment effect, between the log-rank statistic and the restricted-mean difference at each of eleven horizons — the closed form as a line, and the same correlation counted on six thousand null trials of 200 a side as points.

In the world this field has used throughout — a control hazard of 0.35 a year, dropout at 0.1, follow-up to three years — the correlation is −0.461 at half a year and grows in size with every horizon, to −0.949 at three. Counted on six thousand trials with no effect, the correlations are −0.463 and −0.950, and every one of the eleven agrees with the closed form to within 0.003. At the last horizon the log-rank statistic and the restricted mean are close to the same statistic. Both weigh the whole of follow-up, and they differ only in how they weigh early events against late ones.

That near-identity is what makes the combined test cheap. Adding a statistic that is almost one of the existing eleven adds almost nothing to the largest of them under the null. The critical value for the largest of the eleven restricted-mean statistics alone is 2.317. With the log-rank statistic added, drawn from the twelve-dimensional Gaussian vector in the same way, it is 2.374. The insurance costs 0.057 on the critical value.

The size it holds, and the size the habit does not

How often each test reports an effect when the treatment does nothing, at 50, 100 and 200 a side. 6,000 trials at each setting, dropout 0.1 a year, follow-up to three years. no effect, 50 a side: log-rank 5.0%, best of eleven horizons 5.5%, restricted mean at 3 5.4%, one test for all twelve 5.4%, log-rank or horizons, either 8.2%; no effect, 100 a side: log-rank 5.4%, best of eleven horizons 5.8%, restricted mean at 3 5.6%, one test for all twelve 5.9%, log-rank or horizons, either 8.7%; no effect, 200 a side: log-rank 4.9%, best of eleven horizons 5.0%, restricted mean at 3 4.6%, one test for all twelve 4.8%, log-rank or horizons, either 7.8%.
Fig. 2 How often each test reports an effect when the treatment does nothing, at 50, 100 and 200 a side: the log-rank test, the best of eleven horizons on its own critical value, the restricted mean at three years, the combined test, and the habit of running the log-rank test and the horizons each at their own level and reporting a rejection if either rejects.

At 200 a side the combined test rejects 4.82% of trials with no effect. The log-rank test alone rejects 4.90%, the best horizon 5.02%. The habit it replaces is to run the log-rank test at 1.96 and the horizons at 2.317, reporting whichever rejects. It rejects 7.82%. That habit is not unusual. It is what a report does when it shows a log-rank pp-value and a restricted-mean comparison side by side and leads with whichever is smaller. At 100 a side it rejects 8.67%.

At smaller arms every statistic built on restricted means drifts slightly above its level, because their standard errors come from the Greenwood-type variance of a Kaplan–Meier area. That variance is a little small when few subjects remain at the late horizons, as the interval at the end of the curve found for the curve itself. At 100 a side the best horizon rejects 5.75% and the combined test 5.87%. At 50 a side, 5.53% and 5.37%. That is a property of the restricted-mean statistics, not of combining them: the log-rank test, whose variance does not lean on the tail, stays nearer its level, at 4.95% and 5.38%. The joint critical value is exact in the limit. At small arms it inherits the drift of the statistics it combines and adds nothing to it.

Power in three shapes

Three treatment effects are run, each the same as in the essays before it. The first halves the hazard throughout, the shape the log-rank test is built for. The second halves it for a year and then does nothing. In the third the curves cross: the hazard is cut to 0.12 for a year and then raised above the control’s, to the value at which a three-year hazard ratio is exactly one.

Power of four tests of two survival curves, in three shapes of treatment effect, at 50, 100 and 200 a side. 6,000 trials at each setting, dropout 0.1 a year, follow-up to three years. hazard halved throughout, 50 a side: log-rank 64.9%, best of eleven horizons 51.7%, restricted mean at 3 61.1%, one test for all twelve 56.6%; hazard halved throughout, 100 a side: log-rank 91.8%, best of eleven horizons 81.0%, restricted mean at 3 88.4%, one test for all twelve 86.0%; hazard halved throughout, 200 a side: log-rank 99.7%, best of eleven horizons 98.4%, restricted mean at 3 99.4%, one test for all twelve 99.3%; halved for a year, then nothing, 50 a side: log-rank 18.5%, best of eleven horizons 32.1%, restricted mean at 3 26.6%, one test for all twelve 30.8%; halved for a year, then nothing, 100 a side: log-rank 29.3%, best of eleven horizons 51.7%, restricted mean at 3 42.5%, one test for all twelve 50.2%; halved for a year, then nothing, 200 a side: log-rank 51.4%, best of eleven horizons 81.4%, restricted mean at 3 70.3%, one test for all twelve 80.2%; curves that cross, 50 a side: log-rank 5.1%, best of eleven horizons 45.7%, restricted mean at 3 10.4%, one test for all twelve 44.0%; curves that cross, 100 a side: log-rank 5.2%, best of eleven horizons 76.7%, restricted mean at 3 16.4%, one test for all twelve 75.5%; curves that cross, 200 a side: log-rank 5.2%, best of eleven horizons 96.7%, restricted mean at 3 26.5%, one test for all twelve 96.6%.
Fig. 3 Power of the log-rank test, the best of eleven horizons, the restricted mean at three years and the combined test, in the three shapes of treatment effect, at 50, 100 and 200 a side.

With the hazard halved throughout and 100 a side, the log-rank test rejects 91.82% of trials. The best horizon on its own rejects only 81.05%, because it pays the 2.317 critical value for flexibility this shape does not need. The combined test rejects 85.97%, keeping 93.6% of the log-rank power and recovering about half of what the horizons alone gave up. The restricted mean at the last horizon, fixed in advance, rejects 88.38%. That is close to the log-rank test, because at three years it nearly is the log-rank test. At 200 a side every test is above 98% and the differences vanish. At 50 a side the log-rank test leads by eight points over the combined test, 64.88% against 56.63%.

With the curves crossing, the picture inverts. The log-rank test rejects 5.23% at 100 a side, its own size, and 5.22% at 200: by construction the early benefit and the late harm cancel in its sum. The restricted mean at three years rejects only 16.40%, because by three years the curves have crossed back towards each other. The best horizon finds the gap and rejects 76.67%. The combined test rejects 75.50%, 98.5% of that. At 200 a side the two are 96.73% and 96.63%.

The effect that wanes is between the two. The log-rank test rejects 29.33% at 100 a side and the best horizon 51.73%, because the benefit sits early and the restricted mean at an early horizon sees it undiluted. The combined test rejects 50.17%.

The waning shape is the one a pre-specified log-rank test handles worst without failing outright. The treatment has a real effect, and the log-rank test finds it in fewer than a third of trials of 100 a side, because two of its three years of follow-up add events from a period in which the arms are alike, diluting the year in which they are not. When it does reject, the hazard ratio printed beside it is an average over that dilution, set by how long the trial happened to run. The combined test finds the effect in half the trials and, in nineteen of every twenty of its rejections, through one of the restricted means rather than the log-rank statistic.

So the combined test is never the best test in any shape, and never far from it. Its largest loss in these nine settings is to the log-rank test under proportional hazards, eight points at 50 a side and six at 100. Its largest loss to the best horizon is under 1.7 points. The log-rank test, used alone, loses everything under the crossing shape. The best horizon, used alone, loses eleven points under the proportional one at 100 a side. The combined test’s worst case is better than either component’s.

Which statistic decides

Which statistic decides the combined test's rejections, at 100 a side. The share of the combined test's rejections in which the log-rank statistic, rather than one of the restricted means, was the largest: hazard halved throughout 61.2%, halved for a year, then nothing 5.3%, curves that cross 1.2%. Under no effect it is 21.3%.
Fig. 4 At 100 a side, the share of the combined test’s rejections in which the log-rank statistic was the largest of the twelve, in each shape of treatment effect.

The combined test does not choose a statistic in advance, but its rejections can be read afterwards for which statistic carried them. With the hazard halved throughout, the log-rank statistic is the largest of the twelve in 61.2% of the combined test’s rejections. When the effect wanes it leads in 5.3%, and when the curves cross in 1.2%. The test is using the log-rank statistic where the log-rank statistic is the right one and ignoring it where it is not, without having been told which shape it faces.

That reading is not a second test and should not be reported as one. The combined test’s pp-value is the probability that the largest of the twelve exceeds what was seen. Which of the twelve was largest is a description of the data, like the horizon the earlier essay found the gap at. But it tells a reader something the single pp-value does not: whether the evidence is a difference in overall hazard or a difference at some particular time. These are different findings about a treatment, and one minus Kaplan–Meier is a reminder that what a survival statistic measures has to be named before it can be interpreted.

Why one critical value, rather than two adjusted ones

A familiar alternative to a joint critical value is to split the 5% between the two tests: the log-rank test at 2.5% and the horizons at 2.5%. That is the Bonferroni reading, and it is valid. It also ignores what the correlation offers. Under the null the log-rank statistic and the latest horizon’s restricted mean move together with a correlation of −0.949, so the events on which one exceeds its critical value are mostly events on which the other does too. A split of the size counts those events twice. The joint critical value counts them once, which is why adding a twelfth statistic moved it by only 0.057.

Measured rather than argued, the split puts the log-rank test at 2.241 and the horizons’ maximum at 2.589, the 97.5th percentile of their process. It holds its size with room to spare — 4.62% at 100 a side and 3.95% at 200 — and the room is the waste. Its log-rank threshold is lower than the combined test’s 2.374, so under proportional hazards it is slightly ahead, 86.88% against 85.97% at 100 a side. Its horizon threshold is higher, so it falls behind wherever the horizons do the work: 68.12% against 75.50% when the curves cross, 43.32% against 50.17% when the effect wanes. The split is a test that leans towards the log-rank statistic without having been asked to. The joint critical value leans nowhere, because it spends the 5% where the twelve statistics actually disagree.

This is the same arithmetic as combining p-values from studies that share a control, where a correlation that a combination rule ignores mis-states the size. Here the correlation is known exactly, so it can be used rather than guarded against. It is also the same reasoning that warned against testing for a flat point first and then choosing a procedure. A rule that chooses between statistics after looking at them is a statistic of its own. The way to keep its size is to price the choice, not to pretend it was not made.

What dropout does to the price

The correlation was computed at one dropout rate, 0.1 a year, and dropout enters it twice: through the at-risk fraction y(t)y(t), which weights the log-rank sum, and through the variance of each restricted mean, which grows as fewer subjects remain. Heavier dropout shrinks the late part of the follow-up that the log-rank statistic and the latest restricted mean share. They become more alike, not less: the correlation at the last horizon is −0.928 with no dropout, −0.949 at 0.1, −0.978 at 0.3 and −0.985 at 0.6.

The critical value for the combined test barely moves: 2.377 with no dropout, 2.374 at 0.1, 2.368 at 0.3 and 2.374 at 0.6. The critical value for the horizons alone moves more, from 2.310 to 2.358, because heavier dropout decorrelates the early horizons from the late ones and makes eleven horizons more like eleven separate looks. Adding the log-rank statistic then costs less and less, since it is increasingly a copy of the last horizon. A trial can fix 2.374 before it knows its dropout rate and be within a hundredth of the right value anywhere in that range.

That is the useful kind of robustness, and it holds only for dropout of the kind assumed throughout this field: unrelated to the outcome. Dropout that carries information changes what every statistic here estimates, and no critical value repairs that. Nor does any of this change the first lesson of censoring, that a subject who left early is known to have lasted at least that long and is counted as exactly that by every statistic combined here.

What a trial analysing survival curves should pre-specify

The statistics to be combined and the critical value for their maximum. Here, the log-rank statistic and the restricted-mean differences at eleven horizons from half a year to three, read against 2.374. Both the set and the value can be fixed before any data are seen, because the correlation is closed under the null and depends on the control hazard and the dropout, not on the treatment.

Not the habit of reporting both. A log-rank pp-value and a restricted-mean comparison shown side by side, with the stronger one leading, is a test whose size is 7.82% at 200 a side and 8.67% at 100. The two numbers are fine to report. The rule that one of them is the result is what costs the size.

Which statistic carried the rejection, as a description. It says whether the evidence is a difference in hazard over the whole of follow-up or a gap at some time, and in which shape the treatment acted.

The arm size at which the restricted-mean statistics hold their level. At 100 a side and below they run half a point to a point above 5%. A trial that small should expect the combined test to do the same, or use a variance for the late horizons that does not rely on the last few subjects at risk.

What is claimed and what is not

Closed. The correlation between the log-rank statistic and each restricted-mean difference, and among the restricted-mean differences, under no effect with equal arms, exponential control survival and exponential dropout. The two critical values are drawn from that correlation by Cholesky factorisation, two hundred thousand draws each.

Counted. Size and power of every test in four worlds at three arm sizes, six thousand trials a setting, every test computed on the same trials. The counted correlations under no effect, against the closed ones, as the second route.

Not claimed. The correlation was computed under the control arm’s own hazard, which is the null. Under an alternative it changes, and that is why the critical value is fixed at the null: it is the size the test has to hold. Unequal arms, or a log-rank test weighted towards late or early events, change the correlation and have their own closed forms; neither is measured here.

Still open: a horizon chosen by the design rather than the data

The combined test covers both the shapes the field has measured, and it pays for flexibility it might not need. A trial usually knows something about when a treatment acts — a drug’s mechanism, a surgery’s recovery period — and could pre-specify a horizon, or a small set of them, rather than eleven. A grid chosen from that knowledge would have a smaller critical value than 2.374 and keep more of the log-rank test’s power under proportional hazards. The cost is the power lost when the knowledge is wrong. Where between one horizon and eleven the trade balances, as a function of how confident the prior knowledge of the effect’s timing is, is a question about design. The closed correlation can answer it, and it has not been asked.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Closed formCritical valueGaussian processKaplan–MeierLog-rank testMonte CarloMultiple testingProportional hazardsRestricted mean survival timeStatistical power