Every patient at one half
Worth reading first: Simpson's reversal is a region, not a table.
The change that is not confounding put an odds ratio of exactly 2.5 in each of five strata, allocated the treatment by a coin, and found an odds ratio of 1.789 in the pooled table — nothing confounded, because an odds ratio is not an average of odds ratios. It ended by naming the survival version of the same property as the one that “has the property and hides it better”: a hazard ratio is non-collapsible for the same reason, and on top of that it is computed on a population that survival itself is selecting as time goes on.
The hazard ratio the follow-up chose has already shown one way a Cox hazard ratio depends on follow-up: a treatment whose effect wanes produces a hazard ratio that is an average over time, weighted by where the events fall. That drift is real — the treatment’s effect changes — and the hazard ratio reports an average of a changing thing. The case here is the opposite and more troubling. The treatment’s effect does not change, for any patient, at any moment, and the reported hazard ratio drifts anyway.
This essay measures how much better it hides. The setting is as clean as a trial gets: randomised arms, a treatment that multiplies every patient’s hazard by exactly one half at every moment, no censoring but the end of follow-up. The only complication is that patients are not all at the same risk, which is true of every trial ever run.
Every patient at one half
Half the patients have a baseline hazard of 0.1 events a year and half have a hazard of 1. The treatment halves both: 0.05 and 0.5. For any single patient the hazard ratio is 0.5 at every moment, and randomisation puts equal shares of each risk group in each arm.
At the moment of randomisation the arms’ hazards are averages over the same mixture of patients, so their ratio is exactly 0.5. After that the arms stop being the same mixture. The control arm’s high-risk patients have events at a rate of one a year and its low-risk patients at a tenth of that, so the control arm’s survivors become low-risk quickly. The treated arm’s high-risk patients have events at half the rate, so its survivors stay high-risk for longer.
After one year 28.9% of the control arm’s survivors are high-risk and 38.9% of the treated arm’s; after three years, 6.3% and 20.6%. A hazard at time is the event rate among those still at risk at , so a hazard ratio at compares two groups that differ in their mix of patients — and the treated group is the riskier one, precisely because the treatment worked. The ratio of the arms’ hazards rises from 0.5 to 0.929 at 3.6 years, as the treated arm carries the high-risk patients the control arm has lost. Then it falls back towards 0.5, as the treated arm loses them too and both arms are left with low-risk patients alone.
With a continuous spread of risk, a gamma frailty of variance 1, there is no second group to come back to, and the ratio climbs steadily: 0.600 at one year, 0.714 at three and 0.857 at ten, towards 1. A treatment that halves every patient’s hazard produces arms whose hazards, after a decade, are within fifteen per cent of each other.
How much heterogeneity it takes
The size of the drift is set by how different the patients are, and the continuous version makes that a single dial.
A gamma frailty’s variance is the squared coefficient of variation of patients’ baseline hazards. At a variance of one half — patients’ risks spread by about 70% of their average — the ratio of the arms’ hazards drifts from 0.5 to 0.778 over ten years at this base rate; at one, to 0.857; at two, to 0.917. None of these is an exotic population: a trial enrolling across stages of a disease, or across ages, routinely has patients whose risks differ by factors of several, and the drift grows with the spread.
The two-group version shows something the continuous one cannot: the drift is not monotone in general. It depends on the whole distribution of risk, and a mixture with a long, thin high-risk tail produces a ratio that rises while that tail is being depleted and falls once it is gone. A reader who sees an unadjusted hazard ratio change with follow-up cannot tell from its direction alone whether it is heading towards one, back towards the individual ratio, or somewhere else, because the direction is a property of patients the trial did not describe.
What a trial reports
Nobody reads the moment-by-moment ratio. A trial reports one hazard ratio from a Cox model, which is a weighted average of that ratio over the follow-up, weighted by where the events fall. Its value therefore depends on how long the trial ran.
| follow-up | unadjusted, two risk groups | unadjusted, gamma frailty | stratified by risk group |
|---|---|---|---|
| 6 months | 0.524 | 0.527 | 0.500 |
| 1 year | 0.550 | 0.548 | 0.500 |
| 3 years | 0.635 | 0.598 | 0.500 |
| 5 years | 0.672 | 0.624 | 0.500 |
| 10 years | 0.673 | 0.655 | 0.500 |
These are the values the analyses converge to with unlimited patients, so none of the spread is sampling error. Two trials of the same treatment on the same population, one following patients for six months and one for five years, report hazard ratios of 0.524 and 0.672, and a reader comparing them would conclude that the treatment’s benefit wanes with time, or that the longer trial enrolled a population in which it works less well. Neither is true. Every patient’s hazard was halved at every moment in both.
The stratified analysis, which compares each patient with others at the same baseline risk, returns 0.500 at every follow-up. That is not because it corrects for confounding — there is none — but because within a risk group everyone has the same hazard, so survival selects nobody, and the ratio within each group is the ratio for every patient in it. It is the same reason adjusting an odds ratio for a prognostic variable moves it away from the unadjusted one in a randomised trial: the adjusted and unadjusted numbers answer different questions, and only one of them is a property of individual patients.
Why the odds ratio’s problem is worse here
The odds-ratio version of non-collapsibility — the survival-free cousin of Simpson’s reversal, in which a pooled table disagrees with its strata for reasons that are arithmetic rather than causal — is a property of a single table: the pooled odds ratio is closer to one than the stratum odds ratios because odds are not averaged the way probabilities are. The hazard ratio has that property too, and a second one on top of it.
The second is the selection that conditioning on what the treatment caused described in another form. Being alive at time is an outcome of the treatment, and the hazard at is computed among those for whom that outcome went one way. Comparing the arms’ hazards at is therefore comparing two groups defined by a post-randomisation variable, which is the one comparison randomisation does not protect. The two-group trial above is balanced on risk at time zero by design and unbalanced on it at every later time by the treatment’s own effect.
The two properties compound. With a continuous frailty the drift is monotone and heads to one; with two groups it rises and falls as the selection first separates the arms and then empties both of their high-risk patients. The shape depends on the unmeasured distribution of risk, which a trial cannot see, so the drift cannot be corrected by a formula applied to the unadjusted number. It can only be avoided by adjusting for the risk — when the risk is measured — or by reporting a summary that does not have the property.
The two-group curve also shows why the selection cannot be read off the data after the fact. By eight years almost no high-risk patients are left in either arm — 0.07% of the control arm’s survivors and 2.66% of the treated arm’s — and the ratio of the arms’ hazards is back near 0.6 and falling towards 0.5. A trial that began at that point, enrolling only the survivors, would see nearly the individual ratio, because its patients would be nearly homogeneous. The drift is largest in the middle of follow-up, when the arms are most different in composition, and a trial’s own hazard ratio averages over exactly that stretch.
A summary without the property
The restricted mean survival time — the average time alive over the first years — is a mean, and means are collapsible. In the two-group trial the difference in restricted mean survival between the arms is 0.089 years at one year of follow-up and 0.666 at five, and it is exactly the average of the two risk groups’ own differences at every horizon, to every digit computed. Adjusting for risk group changes its precision and not its value, which is what a randomised comparison ought to deliver.
It is not free of the follow-up: a longer horizon lets more of the benefit accumulate, so the difference grows with , and the choice of horizon has to be stated — and stated in advance, since a horizon chosen after looking buys a false-positive rate of its own. But it grows because there is more time in which to be alive, not because the population it is measured on changes composition, and two trials with different horizons report numbers whose difference means something a reader can state. The absolute risk at a fixed time is collapsible for the same reason. At five years 30.7% of the control arm and 43.0% of the treated arm are alive, and those are averages of the risk groups’ survivals with weights fixed by randomisation.
The hazard ratio’s appeal is that it is a single number independent of follow-up — when the proportional-hazards assumption holds. What the trial above shows is that the assumption can hold for every patient and fail for every arm. Proportional hazards at the level of individuals and proportional hazards at the level of randomised groups are different assumptions, and only the first is plausible for a treatment that acts the same way on everyone.
Which number belongs to a patient
The trial’s two hazard ratios are answers to different questions, and it is worth being exact about whose question each one answers.
The stratified ratio, 0.500, is what the treatment does to a patient’s hazard: at any moment, a patient on treatment has half the event rate they would have had without it. That is the number a patient who knows their own risk group needs, and it is the same number for both groups. It is also, in this trial, a property of every patient, which is as close as a trial comes to a causal quantity that applies to an individual.
The unadjusted ratio at three years, 0.635, is what the treatment does to the event rate among those still alive at each moment, averaged over the follow-up. It is not any patient’s number. No patient’s hazard is reduced by 36.5%; each has theirs halved. The unadjusted ratio mixes the treatment’s effect with the treatment’s effect on who survives to be counted, and the second of those has no meaning for a patient deciding whether to take the drug.
The absolute benefit is a third answer, and it is the one most useful for that decision. At one year 77.9% of treated and 63.6% of control patients are alive, a difference of 14.3 points; at three years the difference is 14.7 points; at five, 12.4. Those differences are collapsible — they are the average of the risk groups’ own differences — and they state what the treatment buys in the units a patient experiences. A report of a survival trial that gave the stratified hazard ratio for the mechanism and the survival differences for the decision would be giving each question its own number.
What this means for reading trials
A hazard ratio that weakens with longer follow-up is not evidence that the treatment’s effect wanes. A waning effect does produce that pattern, as the essay on follow-up measured, but so does a constant effect on patients who differ in risk, and the unadjusted curves of the two are the same shape. It is what a constant effect on a heterogeneous population produces, and the heterogeneity that produces it is the ordinary kind — some patients sicker than others. Separating the two needs either the risk measured or a summary that does not select, and the curves themselves — the two arms’ survival probabilities, which are collapsible at every time — are the most direct such summary a trial already prints.
Adjusted and unadjusted hazard ratios in a randomised trial will differ, and the adjusted one is further from one. A reader who expects adjustment in a randomised trial to change only the precision — as it does for a difference in means — will be surprised by an adjusted hazard ratio of 0.50 beside an unadjusted 0.64, and may suspect imbalance. The gap here is entirely non-collapsibility and selection. A reversal a coin cannot prevent found chance imbalance moving a randomised comparison; this is a gap that no imbalance, chance or otherwise, is needed to produce.
A meta-analysis pooling hazard ratios from trials with different follow-up is pooling different quantities. The table above is the between-trial heterogeneity a meta-analysis would report for one treatment in one population, with no heterogeneity in the treatment at all.
A constant individual ratio and a drifting trial ratio, in closed form
With the treatment halving every patient’s hazard and patients in two risk groups, the ratio of the randomised arms’ hazards rises from 0.50 to 0.929 at 3.6 years and falls back; with a gamma frailty of variance 1 it rises to 0.857 at ten years.
The unadjusted Cox hazard ratio a large trial reports is 0.524 at six months of follow-up, 0.635 at three years and 0.672 at five, and the analysis stratified by risk group reports 0.500 at every follow-up.
The restricted mean survival difference is the same whether computed on the whole trial or averaged over the risk groups, at every horizon.
Every quantity is exact: the arms’ survivals, densities and hazards are closed-form mixtures, and the Cox limits solve the expected score equation by numerical integration over the follow-up, which is what a Cox model converges to as the trial grows. No trial is simulated, so none of the spread in the table is noise.
Not claimed: that the frailties here are realistic in size. A tenfold difference in baseline risk between two halves of a trial is large but not unusual for a trial enrolling a mixed population, and a gamma frailty of variance one is a common default in the literature on this effect. The drift is smaller with less heterogeneity and larger with more, and its shape depends on the distribution of risk, which is not identified from the unadjusted data alone.
Still open: the drift that looks like a waning effect
The drift and a genuinely waning effect produce the same unadjusted hazard ratio curve. Separating them needs the risk measured, and most trials measure only part of it — a stage, an age, a biomarker — leaving the rest as a frailty the analysis cannot see.
Adjusting for the measured part removes some of the drift and leaves the rest, so a trial’s adjusted hazard ratio sits somewhere between the individual ratio and the unadjusted one, closer to the former the more of the risk is measured. How much of the drift a prognostic variable of a given strength removes, and whether the remainder is large enough that a trial with good covariates should still expect its hazard ratio to depend on follow-up, is a calculation with the same closed forms as the one above and a partly observed frailty, and it has not been done.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Guessing one arm in three — both name randomisation, selection bias
- The data that stops early — both name selection bias, survival analysis
- The rule that can be guessed — both name randomisation, selection bias
Named objects
A flat tag is an object no other essay names yet.
Cox modelFrailtyHazard ratioNon-collapsibilityRandomisationRestricted mean survival timeSelection biasSurvival analysis