One test for a curve that might cross
Worth reading first: The data that stops early · The curve that survives censoring.
A horizon chosen after looking priced one freedom an analysis of survival curves has. A difference in restricted mean survival can be read at any horizon, and the horizon at which it looks most convincing is not a horizon fixed in advance. The essay showed that the differences at eleven horizons converge jointly to a Gaussian process whose correlation is closed. The largest of their absolute values then has a critical value that can be drawn from that process, 2.317, instead of the 1.96 a single horizon would use. Read against it, the best horizon holds its size and loses little power to the horizon nobody could have known to fix.
That is one way of not having to know in advance what shape a treatment’s effect will take. The log-rank test is another, and it is the one most trials pre-specify. It is the most powerful test when the hazards are proportional and is blind when the curves cross, because then early benefit and late harm cancel in its sum. The hazard ratio the follow-up chose showed what that cancellation does to an estimate. The obvious move is to compute both and keep the stronger, and the obvious objection is that keeping the stronger of two tests is a test of its own, with its own size. This essay builds that test properly, so its size is the stated one, and measures what it costs against each of its components in the shapes where each is best.
The log-rank statistic is a restricted mean, nearly
Every statistic in the comparison is a linear functional of the same thing: the difference between events counted and events expected in each arm, as time runs. The log-rank numerator adds that difference over the whole of follow-up with equal weight at each event. A difference in restricted means to a horizon weights the same differences by , the area under the survival curve from to , divided by the share still at risk. Under no effect, with two equal arms, an at-risk fraction and a hazard , their correlation has a closed form:
The sign says only that a treatment with more events has a smaller restricted mean. The size is the interesting part.
In the world this field has used throughout — a control hazard of 0.35 a year, dropout at 0.1, follow-up to three years — the correlation is −0.461 at half a year and grows in size with every horizon, to −0.949 at three. Counted on six thousand trials with no effect, the correlations are −0.463 and −0.950, and every one of the eleven agrees with the closed form to within 0.003. At the last horizon the log-rank statistic and the restricted mean are close to the same statistic. Both weigh the whole of follow-up, and they differ only in how they weigh early events against late ones.
That near-identity is what makes the combined test cheap. Adding a statistic that is almost one of the existing eleven adds almost nothing to the largest of them under the null. The critical value for the largest of the eleven restricted-mean statistics alone is 2.317. With the log-rank statistic added, drawn from the twelve-dimensional Gaussian vector in the same way, it is 2.374. The insurance costs 0.057 on the critical value.
The size it holds, and the size the habit does not
At 200 a side the combined test rejects 4.82% of trials with no effect. The log-rank test alone rejects 4.90%, the best horizon 5.02%. The habit it replaces is to run the log-rank test at 1.96 and the horizons at 2.317, reporting whichever rejects. It rejects 7.82%. That habit is not unusual. It is what a report does when it shows a log-rank -value and a restricted-mean comparison side by side and leads with whichever is smaller. At 100 a side it rejects 8.67%.
At smaller arms every statistic built on restricted means drifts slightly above its level, because their standard errors come from the Greenwood-type variance of a Kaplan–Meier area. That variance is a little small when few subjects remain at the late horizons, as the interval at the end of the curve found for the curve itself. At 100 a side the best horizon rejects 5.75% and the combined test 5.87%. At 50 a side, 5.53% and 5.37%. That is a property of the restricted-mean statistics, not of combining them: the log-rank test, whose variance does not lean on the tail, stays nearer its level, at 4.95% and 5.38%. The joint critical value is exact in the limit. At small arms it inherits the drift of the statistics it combines and adds nothing to it.
Power in three shapes
Three treatment effects are run, each the same as in the essays before it. The first halves the hazard throughout, the shape the log-rank test is built for. The second halves it for a year and then does nothing. In the third the curves cross: the hazard is cut to 0.12 for a year and then raised above the control’s, to the value at which a three-year hazard ratio is exactly one.
With the hazard halved throughout and 100 a side, the log-rank test rejects 91.82% of trials. The best horizon on its own rejects only 81.05%, because it pays the 2.317 critical value for flexibility this shape does not need. The combined test rejects 85.97%, keeping 93.6% of the log-rank power and recovering about half of what the horizons alone gave up. The restricted mean at the last horizon, fixed in advance, rejects 88.38%. That is close to the log-rank test, because at three years it nearly is the log-rank test. At 200 a side every test is above 98% and the differences vanish. At 50 a side the log-rank test leads by eight points over the combined test, 64.88% against 56.63%.
With the curves crossing, the picture inverts. The log-rank test rejects 5.23% at 100 a side, its own size, and 5.22% at 200: by construction the early benefit and the late harm cancel in its sum. The restricted mean at three years rejects only 16.40%, because by three years the curves have crossed back towards each other. The best horizon finds the gap and rejects 76.67%. The combined test rejects 75.50%, 98.5% of that. At 200 a side the two are 96.73% and 96.63%.
The effect that wanes is between the two. The log-rank test rejects 29.33% at 100 a side and the best horizon 51.73%, because the benefit sits early and the restricted mean at an early horizon sees it undiluted. The combined test rejects 50.17%.
The waning shape is the one a pre-specified log-rank test handles worst without failing outright. The treatment has a real effect, and the log-rank test finds it in fewer than a third of trials of 100 a side, because two of its three years of follow-up add events from a period in which the arms are alike, diluting the year in which they are not. When it does reject, the hazard ratio printed beside it is an average over that dilution, set by how long the trial happened to run. The combined test finds the effect in half the trials and, in nineteen of every twenty of its rejections, through one of the restricted means rather than the log-rank statistic.
So the combined test is never the best test in any shape, and never far from it. Its largest loss in these nine settings is to the log-rank test under proportional hazards, eight points at 50 a side and six at 100. Its largest loss to the best horizon is under 1.7 points. The log-rank test, used alone, loses everything under the crossing shape. The best horizon, used alone, loses eleven points under the proportional one at 100 a side. The combined test’s worst case is better than either component’s.
Which statistic decides
The combined test does not choose a statistic in advance, but its rejections can be read afterwards for which statistic carried them. With the hazard halved throughout, the log-rank statistic is the largest of the twelve in 61.2% of the combined test’s rejections. When the effect wanes it leads in 5.3%, and when the curves cross in 1.2%. The test is using the log-rank statistic where the log-rank statistic is the right one and ignoring it where it is not, without having been told which shape it faces.
That reading is not a second test and should not be reported as one. The combined test’s -value is the probability that the largest of the twelve exceeds what was seen. Which of the twelve was largest is a description of the data, like the horizon the earlier essay found the gap at. But it tells a reader something the single -value does not: whether the evidence is a difference in overall hazard or a difference at some particular time. These are different findings about a treatment, and one minus Kaplan–Meier is a reminder that what a survival statistic measures has to be named before it can be interpreted.
Why one critical value, rather than two adjusted ones
A familiar alternative to a joint critical value is to split the 5% between the two tests: the log-rank test at 2.5% and the horizons at 2.5%. That is the Bonferroni reading, and it is valid. It also ignores what the correlation offers. Under the null the log-rank statistic and the latest horizon’s restricted mean move together with a correlation of −0.949, so the events on which one exceeds its critical value are mostly events on which the other does too. A split of the size counts those events twice. The joint critical value counts them once, which is why adding a twelfth statistic moved it by only 0.057.
Measured rather than argued, the split puts the log-rank test at 2.241 and the horizons’ maximum at 2.589, the 97.5th percentile of their process. It holds its size with room to spare — 4.62% at 100 a side and 3.95% at 200 — and the room is the waste. Its log-rank threshold is lower than the combined test’s 2.374, so under proportional hazards it is slightly ahead, 86.88% against 85.97% at 100 a side. Its horizon threshold is higher, so it falls behind wherever the horizons do the work: 68.12% against 75.50% when the curves cross, 43.32% against 50.17% when the effect wanes. The split is a test that leans towards the log-rank statistic without having been asked to. The joint critical value leans nowhere, because it spends the 5% where the twelve statistics actually disagree.
This is the same arithmetic as combining p-values from studies that share a control, where a correlation that a combination rule ignores mis-states the size. Here the correlation is known exactly, so it can be used rather than guarded against. It is also the same reasoning that warned against testing for a flat point first and then choosing a procedure. A rule that chooses between statistics after looking at them is a statistic of its own. The way to keep its size is to price the choice, not to pretend it was not made.
What dropout does to the price
The correlation was computed at one dropout rate, 0.1 a year, and dropout enters it twice: through the at-risk fraction , which weights the log-rank sum, and through the variance of each restricted mean, which grows as fewer subjects remain. Heavier dropout shrinks the late part of the follow-up that the log-rank statistic and the latest restricted mean share. They become more alike, not less: the correlation at the last horizon is −0.928 with no dropout, −0.949 at 0.1, −0.978 at 0.3 and −0.985 at 0.6.
The critical value for the combined test barely moves: 2.377 with no dropout, 2.374 at 0.1, 2.368 at 0.3 and 2.374 at 0.6. The critical value for the horizons alone moves more, from 2.310 to 2.358, because heavier dropout decorrelates the early horizons from the late ones and makes eleven horizons more like eleven separate looks. Adding the log-rank statistic then costs less and less, since it is increasingly a copy of the last horizon. A trial can fix 2.374 before it knows its dropout rate and be within a hundredth of the right value anywhere in that range.
That is the useful kind of robustness, and it holds only for dropout of the kind assumed throughout this field: unrelated to the outcome. Dropout that carries information changes what every statistic here estimates, and no critical value repairs that. Nor does any of this change the first lesson of censoring, that a subject who left early is known to have lasted at least that long and is counted as exactly that by every statistic combined here.
What a trial analysing survival curves should pre-specify
The statistics to be combined and the critical value for their maximum. Here, the log-rank statistic and the restricted-mean differences at eleven horizons from half a year to three, read against 2.374. Both the set and the value can be fixed before any data are seen, because the correlation is closed under the null and depends on the control hazard and the dropout, not on the treatment.
Not the habit of reporting both. A log-rank -value and a restricted-mean comparison shown side by side, with the stronger one leading, is a test whose size is 7.82% at 200 a side and 8.67% at 100. The two numbers are fine to report. The rule that one of them is the result is what costs the size.
Which statistic carried the rejection, as a description. It says whether the evidence is a difference in hazard over the whole of follow-up or a gap at some time, and in which shape the treatment acted.
The arm size at which the restricted-mean statistics hold their level. At 100 a side and below they run half a point to a point above 5%. A trial that small should expect the combined test to do the same, or use a variance for the late horizons that does not rely on the last few subjects at risk.
What is claimed and what is not
Closed. The correlation between the log-rank statistic and each restricted-mean difference, and among the restricted-mean differences, under no effect with equal arms, exponential control survival and exponential dropout. The two critical values are drawn from that correlation by Cholesky factorisation, two hundred thousand draws each.
Counted. Size and power of every test in four worlds at three arm sizes, six thousand trials a setting, every test computed on the same trials. The counted correlations under no effect, against the closed ones, as the second route.
Not claimed. The correlation was computed under the control arm’s own hazard, which is the null. Under an alternative it changes, and that is why the critical value is fixed at the null: it is the size the test has to hold. Unequal arms, or a log-rank test weighted towards late or early events, change the correlation and have their own closed forms; neither is measured here.
Still open: a horizon chosen by the design rather than the data
The combined test covers both the shapes the field has measured, and it pays for flexibility it might not need. A trial usually knows something about when a treatment acts — a drug’s mechanism, a surgery’s recovery period — and could pre-specify a horizon, or a small set of them, rather than eleven. A grid chosen from that knowledge would have a smaller critical value than 2.374 and keep more of the log-rank test’s power under proportional hazards. The cost is the power lost when the knowledge is wrong. Where between one horizon and eleven the trade balances, as a function of how confident the prior knowledge of the effect’s timing is, is a question about design. The closed correlation can answer it, and it has not been asked.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The cliff that is a slope — both name closed form, critical value, monte carlo, statistical power
- A boundary for giving up — both name closed form, monte carlo, statistical power
- A coverage table with its own error — both name closed form, monte carlo, statistical power
- A detector built for the ordering — both name critical value, monte carlo, statistical power
- A look the trend asked for — both name closed form, monte carlo, statistical power
- A simulation that stops when it looks settled — both name closed form, monte carlo, statistical power
Named objects
A flat tag is an object no other essay names yet.
Closed formCritical valueGaussian processKaplan–MeierLog-rank testMonte CarloMultiple testingProportional hazardsRestricted mean survival timeStatistical power