A horizon chosen after looking
Worth reading first: The data that stops early.
The hazard ratio the follow-up chose ended by recommending a summary that carries its window on its face. The restricted mean survival time to a horizon τ is the area under a survival curve up to τ — the expected time lived out of the first τ — and the difference between two arms’ restricted means is an average over a window that says which window. It does not change meaning when hazards cross, and it can be read directly off two Kaplan–Meier curves, with a standard error from the same arithmetic Greenwood’s formula uses.
The window is still one number, and nothing in the arithmetic fixes it. A protocol can state it; a report can also compute the difference at a range of horizons, notice where it looks strongest, and print that one. The difference at any single horizon is estimated honestly and its interval covers at its stated rate. The interval beside the horizon that was chosen because it looked best does not, and this essay measures by how much, what the right critical value for that choice is, and what a choice paid for honestly still buys.
The worlds are the ones the hazard-ratio essay built: a control arm with a hazard of 0.35 a year, dropout at 0.1 a year in both arms, and a treatment that changes its hazard once, at one year — proportionally, by waning, or by crossing over the control. A fourth world, in which the treatment arm is a copy of the control, is the null. Trials enrol two hundred subjects an arm unless stated, follow them to three years, and read the difference in restricted means at eleven horizons, from half a year to three in steps of a quarter.
Three worlds read at every horizon
The truth at each horizon is a closed form. The proportional treatment’s advantage grows steadily, to 0.4766 years by τ = 3, and the waning treatment’s grows more slowly, to 0.2675, because after the first year it adds nothing and the survival gap it opened simply persists. The crossing treatment’s difference climbs to 0.1776 at τ = 2 — close to the time the two survival curves cross — and then falls back to 0.1366 at τ = 3, the figure the earlier essay reported, because after the crossing the treated arm is dying faster and every extra month of window subtracts.
So the horizon changes the answer in the crossing world by nearly a third, and in all three worlds a report at a later horizon reports a larger number of years with a larger standard error. Neither of those is a defect of the restricted mean. They are what naming a window means. The trouble begins when the window is named after the data have been seen.
A trial of nothing chooses its horizon
Run ten thousand trials in which the treatment arm is a copy of the control. At each of the eleven horizons, compute the difference in restricted means and its z, and report the horizon at which |z| is largest. A test at the protocol’s horizon of three years rejects 4.70% of these trials, as a 5% test should. The same test at the chosen horizon rejects 11.24%.
Where the trials put their chosen horizon is not even. 26.8% of them choose the first horizon, half a year, and 21.4% the last, three years, against the 9.1% each of eleven horizons would get if the choice were spread evenly. The middle horizons are chosen 3.5% to 10.6% of the time. The reason is the correlation between neighbouring horizons. A difference in restricted means to 1.5 years and one to 1.75 share almost all their data, so an interior horizon is hemmed in by neighbours that nearly always move with it, and it is rarely the single largest. The two end horizons each have a neighbour on only one side, and the first is the least correlated with everything after it.
Under the waning treatment the choice gathers instead: 28.4% of trials choose τ = 1.25, the horizon at which the treatment’s early benefit is largest relative to the noise. A trial with a real effect picks a horizon for a reason; a trial of nothing picks one at an end.
A price in closed form
The eleven z statistics of one trial are strongly correlated, and the correlation does not need to be simulated. For one arm, the covariance of the restricted means at two horizons a and b is an integral from zero to the smaller horizon of the area under the survival curve from t to a, times the area from t to b, times the hazard over the chance of still being observed at t. With the treatment arm a copy of the control, the difference’s covariance is twice that, and its correlation depends on neither the number of subjects nor anything but the control’s hazard and the dropout.
Computed that way, neighbouring horizons a quarter of a year apart are correlated at 0.9514, and the first horizon with the last at 0.570. The correlation with the three-year horizon runs from 0.570 at half a year to 0.997 at 2.75 years.
That correlation defines a Gaussian process, the one the standardised differences converge to, and drawing two hundred thousand vectors from it by Cholesky factorisation — without drawing a single subject — gives the size of reading the largest |z| against 1.96: 11.31% for eleven horizons, beside the counted 11.24%. The same route gives 8.61% for three horizons at one, two and three years, counted at 8.47%; 11.38% for twenty-six horizons a tenth of a year apart, counted at 11.46%; and 11.60% for 251 horizons a hundredth apart, which is as close to choosing the horizon continuously as the grid needs to come.
The price stops growing. Going from three horizons to eleven costs nearly three points of size, and going from eleven to 251 costs a third of a point more, because horizons closer together than a quarter of a year are so correlated that adding them adds almost no new chances. The freedom to choose any horizon is worth about the same as the freedom to choose among eleven.
The process also supplies the repair. The 95th percentile of its largest absolute coordinate is the critical value a chosen horizon should be read against: 2.197 for three horizons, 2.317 for eleven, 2.319 for twenty-six and 2.344 for 251. The counted trials give their own 95th percentile, 2.316 for eleven, and read against the process’s 2.317 they reject 4.99% — the price paid in full and no more. It is the same move the smallest of three combinations needed: a report that keeps the most convincing of several correlated statistics is honest when its critical value is computed for the maximum rather than for the one it kept.
The structure is also the one a group-sequential trial has. A restricted mean to a longer horizon contains every observation the shorter one used and adds the months after it, so the eleven differences are nested looks at accumulating information, and choosing the largest is choosing the look at which the evidence was strongest. The difference from an error rate spent across interim looks is that nothing stops at a horizon: every horizon is computed at the end, from the same finished data, which is why the whole cost can be paid by one constant rather than a boundary that changes from look to look. And the constant is smaller than a group-sequential one — 2.317 for eleven horizons against 2.413 for Pocock’s five equally spaced looks — because horizons a quarter of a year apart share far more of their information than looks a fifth of a trial apart.
What the price buys back when the curves cross
A critical value of 2.317 is a stricter bar than 1.96, and the question is what the chosen horizon can still find over it. Under the crossing treatment the answer is almost everything. A protocol that fixed the horizon at three years — the natural choice, the length of follow-up — has 26.31% power, because by then the treated arm’s late excess of deaths has eaten most of its early gain. The best horizon to have fixed is 1.25 years, with 98.49%, and nobody writing the protocol could have known to fix it there.
The chosen horizon, priced at 2.317, has 96.92% power: 1.57 points short of the best fixed horizon, and seventy points above the protocol’s. The unpriced version has 98.76%, and those extra 1.84 points are exactly what a test of size 11.24% buys. Choosing honestly does not give up the benefit of looking; it gives up the part of the benefit that was an error rate.
The horizon the tests prefer is not the horizon at which the difference is largest. The true difference peaks at two years and power peaks at 1.25, because the standard error of a restricted mean grows with its window, and a z statistic is a difference over its standard error. 61.0% of trials choose 1.25 years and 33.3% choose one year. A report that picks the most significant horizon picks one earlier than the largest effect, and reports a smaller gain in years than the window at which the gain is largest would have shown.
A waning effect, where the choice is hardest
Under the waning treatment the power curve across horizons is flatter, and the choice has less to gain. A horizon fixed at three years has 70.68% power; the best fixed horizon is again 1.25 years, with 85.37%. The chosen horizon priced at 2.317 has 81.37% — ten points above the protocol’s horizon and four below the best one — and unpriced it has 89.43%.
A flat power curve is also the case in which selection distorts the estimate most, because many horizons are nearly tied and the one that wins is the one whose noise happened to point upwards. At the chosen horizon the reported difference averages 0.1550 years, where the true difference at the horizons chosen averages 0.1390: the report overstates the gain at its own horizon by 11.5%. Under the crossing treatment, where one region of horizons is clearly best, the overstatement is 2.5%.
What the interval at the chosen horizon covers
The cost of choosing lands differently on an interval than on a test. With no difference between the arms, the interval at the chosen horizon misses zero exactly when the test at that horizon rejects, so its coverage is one minus the size: 85.80% with 25 subjects an arm, 87.36% with 50, 87.83% with 100, 88.76% with 200 and 88.65% with 400. More subjects do not repair it, because the correlation between horizons, and so the price of the maximum, does not depend on the number of subjects.
With a real effect the damage fades as the trial grows, because a strong signal decides the horizon and the noise stops deciding it. Under the waning treatment the chosen interval covers 86.21% at 25 subjects an arm, 89.03% at 50, 92.27% at 100, 94.72% at 200 and 94.63% at 400, and the overstatement of the gain at the chosen horizon falls with it, from 33.2% at 25 subjects an arm to 5.9% at 400. Under the crossing treatment coverage reaches 94.55% by 100. The interval at a horizon fixed at three years covers between 93.78% and 95.22% at every size.
So the two failures are not the same failure. A trial of nothing that chooses its horizon reports an interval that excludes no effect on more than a tenth of occasions, however large it is. A small trial of a real, fading effect reports a gain up to a third larger than the truth at the horizon it chose. A large trial of a real effect loses little by choosing at all.
Small trials and an asymptotic price
The Gaussian process is the large-sample limit, and at small sizes the trials drift from it. With 25 subjects an arm, even the fixed three-year horizon rejects 6.26% of null trials, because a restricted mean from a curve with few subjects at risk late has a standard error that is itself noisy. The counted 95th percentile of the largest |z| is 2.436 there, above the process’s 2.317, and read against 2.317 the chosen horizon rejects 6.79%. By 100 subjects an arm that is 5.62%, by 200 it is 4.99%, and at 400 5.23%.
A report from a small trial that wants to choose its horizon honestly should take its critical value from simulated trials of its own size rather than from the process, in the same way that an interval at the end of a survival curve needs more than Greenwood’s formula when few subjects remain at risk.
What a restricted-mean report should state
The horizon, and when it was fixed. A horizon named in the protocol makes a difference in restricted means an exact 5% test. A horizon named after the curves were seen is a draw from a test of size near 11% unless its critical value was raised to pay for the choice.
If the horizon was chosen, how many were looked at, and the critical value used. Eleven horizons over three years call for 2.317 at these dropout and hazard rates; any grid can be priced the same way from the closed correlation, before a single subject is enrolled.
The difference at the protocol’s horizon beside the chosen one. Under the crossing treatment the true differences at three years and at 1.25 years are 0.1366 and 0.1364 — the same gain in years — and a trial reading them has 26.31% power at one and 98.49% at the other. A reader shown both knows that the choice moved the significance and not the size of the gain.
The broader point is the one twenty analyses of nothing makes with discrete choices and a look the trend asked for makes with the timing of an interim analysis: when the looking happens is part of the result. What is particular to a horizon is that the choice is continuous and its price is finite — about 11.6% however finely the window is searched — because neighbouring windows share almost all their data.
Closed, simulated and counted
Closed. The true difference in restricted means at every horizon in each world, and the correlation of the standardised differences across horizons under the null, by integration.
Simulated without subjects. The size of reading the largest of the process’s coordinates against 1.96, and the critical value that restores 5%, from Gaussian vectors with that correlation.
Counted. Every size, power, coverage and overstatement, on ten thousand trials at each setting, with subjects drawn from the hazards and censored at random, and the restricted means and their standard errors read off each trial’s own Kaplan–Meier curves.
Particular to these worlds. One control hazard, one dropout rate, treatments that change once at a year, a window up to three years, and a choice made by the largest |z| in either direction, with dropout that carries nothing about the event. Dropout that does — the kind the data cannot see — would move every restricted mean before any horizon was chosen. A choice made by the largest difference in years rather than the largest z, or in one direction only, is a different maximum with a different price.
Still open: one test for a curve that might cross
The chosen horizon is one way of not having to know in advance what shape the treatment’s effect takes. The log-rank test is another, optimal when hazards are proportional and blind when they cross, and a report could compute both and keep the stronger — the pattern of several correlated statistics with one maximum that this essay priced for horizons. The joint distribution of a log-rank statistic and restricted-mean differences at several horizons is again a Gaussian process with a computable correlation, so the critical value for their maximum is a closed-form draw rather than a simulation of trials. Whether that combined test keeps the log-rank test’s power under the proportional treatment while recovering the crossing treatment’s 96.92%, and what the correlation between a log-rank statistic and a restricted mean is in these worlds, is the next measurement.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A coverage table with its own error — both name coverage, monte carlo, multiple comparisons, statistical power
- The rank is a decision — both name critical value, monte carlo, multiple comparisons, statistical power
- A detector built for the ordering — both name critical value, monte carlo, statistical power
- A simulation that stops when it looks settled — both name coverage, monte carlo, statistical power
- An order that spends the error rate — both name monte carlo, multiple comparisons, statistical power
- Estimating how many nulls are true — both name monte carlo, multiple comparisons, statistical power
Named objects
A flat tag is an object no other essay names yet.
CensoringCoverageCritical valueEstimandHazard ratioKaplan–MeierLog-rank testMonte CarloMultiple comparisonsRestricted mean survivalStatistical powerSurvival curve