The hazard ratio the follow-up chose
Worth reading first: The data that stops early.
A trial reports a hazard ratio of 0.76 for its treatment after three years of follow-up, with a tight interval and a small p-value. The same treatment, in the same population, would have been reported at 0.5000 had the trial stopped at one year, and at 0.8194 had it run to eight. The treatment did not change between those three trials. It halves the hazard for the first year and then stops working, and each trial summarised that with one number.
A hazard ratio is one number only when the hazards in the two arms keep a fixed ratio for the whole of follow-up. When they do not — an effect that wanes, a benefit that turns into a harm, a treatment that works only after a delay — a Cox model still returns one number, still converges, and still produces an honest interval. What it converges on is an average of a ratio that changes over time, and the weights of that average are set by when the events happen, which depends on how long the trial ran and how many subjects it lost along the way.
Three treatments and one control
The world is kept small enough that every number has a closed form behind it. The control arm has a hazard of 0.35 a year throughout. Each of three treatments changes its hazard once, at one year, and dropout runs at 0.1 a year in both arms, independent of everything — the kind of censoring Kaplan–Meier handles without bias.
The proportional treatment halves the hazard throughout, to 0.175. It has a hazard ratio, one half, in the only sense the phrase has a meaning.
The waning treatment halves it for the first year and then has no effect at all: its hazard returns to 0.35. Its ratio is one half and then one.
The crossing treatment cuts the hazard to 0.12 for a year and then raises it to 0.5779, above the control. That late value is not arbitrary. It was found by bisection as the one at which a Cox model with three years of follow-up converges on a hazard ratio of exactly one, so that this treatment is, for a three-year trial, the null result it will be reported as.
Only the first treatment has a hazard ratio. The other two have a ratio function — a hazard ratio at each time — and every one-number summary of it has to decide how to weight the times.
What a Cox model converges to when there is no single ratio
A Cox model estimates its coefficient by solving the score equation of the partial likelihood: at every event time, the arm the event happened in is compared with what a common ratio would have predicted given who was at risk. Summed over events, the observed-minus-expected sums to zero at the estimate.
When the model is right, that sum is zero at the true ratio. When it is wrong, the estimate still converges — to the value at which the expected score is zero. For two arms with at-risk proportions and and hazards and , that value solves
which is the result Struthers and Kalbfleisch published for a misspecified Cox model. It says what the hazard ratio is when there is no hazard ratio: the constant ratio that balances the observed and expected events over the follow-up actually run. The upper limit of the integral is the length of follow-up. The terms contain the dropout. Both are inside the definition of the quantity being estimated.
Two routes check that this is what the fits do. The integral is solved numerically, and separately, 400 trials of 400 subjects an arm are simulated at each length of follow-up and fitted by Newton’s method on the partial likelihood. For the waning treatment the limits at one, two, three, five and eight years are 0.5000, 0.7009, 0.7617, 0.8034 and 0.8194; the counted means are 0.5014, 0.7020, 0.7635, 0.8034 and 0.8208, each within four of its own standard errors. A third route checks the fitting code itself: the Cox score statistic at a ratio of one is, by algebra, the log-rank statistic, and in the simulated trials the two agree to .
The control that stops this from being read as a property of the Cox model is the proportional treatment. Its counted hazard ratio is 0.5014, 0.5012, 0.5002, 0.5001 and 0.5005 at the same five lengths of follow-up — flat, at the true value, from trials run identically. The drift belongs to the waning treatment, and to the question a single ratio asks of it.
Where the events fall is where the ratio is read
The weights in that integral are easiest to see through a quantity that has nothing to do with the model: the share of all events in the trial that happen during the year the treatment works.
A trial that stops at one year has all of its events inside the effective year, and it reports one half. At two years 52.5% of the events are inside it; at three years 40.3%; at five 32.4%; at eight 29.5%. Every extension adds events from the years in which the arms have the same hazard, and each of those events votes for a ratio of one. The estimate moves towards one not because the treatment’s benefit is being diluted by noise but because the trial is being asked, more and more, about years in which there was no benefit to find.
This is the same shape as what a wrong model estimates in the sandwich field, where a straight line fitted to a curve converges to the design’s own tangent: two honest studies with different designs, each estimating something precise, neither estimating the other’s quantity. Here the design variable is time. It is also the shape of whose effect an instrument estimates — a weighted average over a subpopulation the method chose, reported under a name that sounds like a property of the treatment.
An interval that is honest about an average
It would be natural to hope that the interval around the hazard ratio at least warns when the number is an average of something that changes. It does not, and it is correct not to, because the interval is doing its job.
Around the limit it converges to, the model-based interval covers 96.0%, 93.5%, 95.8%, 95.8% and 94.8% at the five lengths of follow-up — each within a standard error or so of 95% at 400 trials. It is a well-calibrated interval about a precisely defined quantity.
Around one half, the ratio during the year the treatment actually works, the same intervals cover 96.0% at one year, when the two are the same number, then 14.5% at two years and 0.3% at three. By five years not one of the 400 intervals contains it. A longer trial has more events and a narrower interval, and the narrowing tightens the interval around the average and further away from the effect, the same shape as an estimate that becomes more precise about the wrong estimand as the sample grows.
None of this is a coverage failure in the usual sense. It is an interval that answers a question about the follow-up-weighted average, read by somebody who asked about the treatment.
The dropout rate is part of the answer too
Follow-up is a design decision somebody made. Dropout is not, and it enters the same integral.
With no dropout, the waning treatment’s three-year hazard ratio is 0.7789. At the dropout rate used above, 0.1 a year, it is 0.7617. At a dropout hazard of one a year it is 0.6289. The crossing treatment moves from 1.0445 to exactly 1.0000 to 0.6655 over the same range — from a trial that would report harm to one that would report a substantial benefit.
The mechanism is the same as for follow-up. Dropout removes subjects before the late events can happen, so heavier dropout shifts the share of events towards the early months and pulls the average towards the early ratio. And this is dropout of the harmless kind: independent of the event, identical in both arms, no bias in any Kaplan–Meier curve in the trial. It is not the case the data cannot see, where dropout carried information about the event. Here it carries none and still changes the estimand, because the estimand was defined over whoever remained.
Two trials of the same treatment with different retention therefore report different hazard ratios, neither of them biased, and a meta-analysis that pools them is pooling two different averages.
A trial stopped early reports the early ratio
There is a third way the follow-up gets chosen, and it is the one least likely to be noticed: the trial stops itself. A trial with interim analyses and a stopping boundary for efficacy ends when the accumulated evidence crosses the boundary, and under a waning effect that is most likely to happen early, while nearly every event is still inside the year in which the treatment works.
Read through the integral, a trial that stops at eighteen months is a trial whose weights sit mostly on the first year, and it reports a hazard ratio near the early one half rather than the 0.76 a three-year trial would have reported. A trial of the same treatment that did not cross its boundary runs to completion and reports the diluted average. So among trials of one waning treatment, the ones stopped early for benefit report stronger effects than the ones that ran their course — not because stopping inflated a noisy estimate, which is the selection a stopping rule is known to add, but because the stopped trials were estimating a different average. Both effects push in the same direction, and a comparison of early-stopped and completed trials cannot tell them apart.
That is argued from the integral here rather than counted, and the size of the gap depends on where the boundary sits and how fast events accrue, neither of which was simulated. The direction does not depend on either. For an effect that fades, a shorter follow-up puts more weight on the period in which it was present, and a stopping rule is a rule for choosing the follow-up after the data have started to arrive.
A ratio of one from two curves that differ
The crossing treatment makes the problem visible in a single picture, because its three-year hazard ratio reads exactly null.
At one year the treated arm’s survival is 0.8869 against the control’s 0.7047, a gap of 0.1822: eighteen points more of the cohort alive at the end of the first year. The treated curve then falls faster and crosses the control at 2.009 years, ending at 0.2792 against 0.3499 at three. A single hazard ratio of one describes both curves.
The area under each curve to three years — the restricted mean survival time, the expected number of years lived out of the first three — is 1.9939 for the treated arm and 1.8573 for the control. Averaged over the whole cohort, the treatment bought 0.1366 years of life inside the trial’s window. Whether that is worth an extra year of higher hazard afterwards is a judgement about the patients; but it is not a null result, and a hazard ratio of one reports it as one.
It is the survival-analysis version of four datasets with one summary: a single statistic that is exactly the same for pictures a reader would never confuse.
A test that cannot see it
If the effect estimate is an average that can cancel, so is the test built from the same arithmetic.
The log-rank test rejects 5.4%, 5.6%, 5.1% and 6.2% of the time at 100, 200, 400 and 800 subjects an arm. That is close to its own level at every size. More data does not help it, because it weights each event time the way a Cox model does, and the early benefit and the late harm cancel in its numerator exactly as they cancel in the estimate.
A comparison of the two curves at one year, using Greenwood’s variance for each, rejects 89.2% of the time at 100 an arm and 99.8% at 200. It is not a better test in general. It is a test of a different hypothesis — that survival at one year is the same — and against these two arms that hypothesis is false by eighteen points while the log-rank test’s hypothesis is, by construction, true.
The log-rank test is the most powerful test available against proportional alternatives, and it is performing exactly as designed. Its null is not “the curves are the same”. It is “the weighted average of the hazard ratio is one”, and that null can be true of two treatments that are very different.
What proportional hazards was buying
It is worth saying what the assumption does when it holds, because it is easy to hear the above as an argument against Cox models, and it is not.
Under proportional hazards, the integral that defines the limit has one root whatever the terms are, because the ratio in its second factor is constant and the weights cannot move it. That is exactly what the flat line for the proportional treatment shows. Follow-up changes the precision of the estimate and nothing else. Dropout changes the precision and nothing else. Two trials with different lengths and retention estimate the same number, and pooling them is legitimate. The assumption is what makes the hazard ratio a property of the treatment instead of a property of the trial.
So the real question is never whether a Cox model is “valid” for non-proportional data. It always converges to something well defined. The question is whether that something is the quantity the report names, and the answer depends on whether the ratio was constant — which the data can speak to, weakly, and which the reader of a single number cannot check at all.
Which of these are theorems and which are counts
Proved. The characterisation of the limit as the root of the expected score is a theorem about misspecified Cox models, not a finding of this essay; the numerical integration here computes it and the simulated fits confirm the computation. That the Cox score at a ratio of one is the log-rank statistic is algebra, and the two are computed separately here only to confirm that they agree. That the limit is the same at every follow-up under proportional hazards follows from the integral.
Computed for these worlds. Every specific number — 0.7617 at three years, the event shares, 0.6289 at heavy dropout, the eighteen-point gap under a null hazard ratio, 0.1366 years of restricted mean survival — belongs to the three hazard shapes chosen. They were chosen to be simple rather than typical. A real waning effect declines gradually rather than switching off, which would change the numbers and not the direction: the limit moves towards the ratio in whichever period contributes the most events.
Counted. The coverage figures and the power figures are simulations, with standard errors of about a point at 400 trials and under a point at 1,000. The log-rank test’s rejection rate under crossing hazards is not guaranteed to be exactly 5%, since the limit being one makes the numerator’s expectation zero without making its variance the null variance; the counts read between 5.1% and 6.2%, which is what was measured and all that is claimed.
Not examined. How often a test of the proportional-hazards assumption would flag these trials was not counted, and it should not be assumed to be high: such tests are known to have limited power at modest sample sizes, and a check that has never been measured against a specific departure has not earned the reassurance people take from it passing.
What a report can print instead
Three quantities carry their time horizon on their face, and each is a better primary summary when the ratio may not be constant.
Survival at stated times, in each arm. “88.7% against 70.5% at one year, 27.9% against 35.0% at three” is what the crossing trial found, and nobody reading it would call it a null result.
The restricted mean survival time to a stated horizon. The difference, 0.1366 years to three years, is an average over the window that says which window; it is estimable directly from the Kaplan–Meier curves and does not change meaning when the hazards cross.
The hazard ratio with its follow-up named, and a plot of the ratio over time. If a hazard ratio is reported, reporting it as “the average hazard ratio over three years of follow-up, with dropout of so much” is accurate, and a smoothed estimate of the ratio against time shows the reader whether that average is standing in for a constant.
None of these needs a different trial. They are different readings of the same curves, and the difference between them is whether the number printed carries the choices that produced it.
Where this goes next: a horizon chosen after looking
The restricted mean survival time repairs the problem this essay measured by making the horizon explicit, and that opens the next question. A horizon can be chosen after the curves have been seen. Under the crossing treatment, the difference in restricted means is positive at every horizon up to some point past the crossing and shrinks afterwards; a report free to pick the horizon can pick the one where the difference looks largest, and an interval computed as though the horizon had been fixed in advance will not cover at its stated rate.
That is a distinct argument from this one. This essay concerned an estimand whose weights the design chose silently; the next concerns an estimand whose one free parameter an analyst chooses visibly, and the thing to measure is how much the freedom costs — the coverage of a restricted-mean interval at a horizon chosen to maximise its difference, against one fixed in the protocol, in the same three worlds. It is twenty analyses of nothing with a continuous dial instead of twenty discrete choices, and when the looking happens decides, as it does for a stopping rule, what the reported number is relative to.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The interval at the end of the curve — both name censoring, kaplan–meier, risk set, survival curve
- A weight fitted to balance — both name estimand, model misspecification
- An imputation model the analysis does not contain — both name estimand, model misspecification
- The moments a balance is told — both name estimand, model misspecification
- The repair that keeps the question — both name estimand, statistical power
- When the order matters — both name model misspecification, statistical power
Named objects
A flat tag is an object no other essay names yet.
CensoringCox modelEstimandHazardHazard ratioKaplan–MeierLog-rank testModel misspecificationProportional hazardsRestricted mean survivalRisk setStatistical powerSurvival curve