The spread, and its own uncertainty

A history that stepped

When sixteen earlier control arms stepped once — the first eight a fifth of a standard deviation below the last — borrowing them as exchangeable flatters a treatment with no effect by 0.063 and stops borrowing, and fitting a line through them penalises it by 0.052. A model with a change point is unbiased at every step size, borrows more the larger the step, 512 patients at a fifth, and puts two thirds of its belief on the right place for the change. Against a steady drift it is the one that fails, so each of the three is safe only against the shape of history it assumes.

Worth reading first: Eight groups, one population.

A history that has been moving borrowed a control arm from earlier trials whose control mean had been drifting, and found that an exchangeable model turns the drift into a spread centred on the history’s middle, while a line in time removes a straight-line drift exactly. It ended on the history a line handles worst. A new standard of care arrives between two trials, and the controls before it sit at one level and the controls after it at another. Fitted with a line, that is a slope, and the line extrapolates it past the current trial. Fitted as exchangeable, it is a spread, and the prior sits between the two levels.

The model that matches a single change has a change point. It is measured here beside the other two, on histories that stepped, histories that drifted and histories that did neither.

Three models for sixteen earlier trials

The setting is the earlier essay’s. Sixteen earlier trials each contribute a control mean from two hundred patients; the current trial has fifty patients an arm; every control mean is in units of the patients’ standard deviation, and the current trial’s true control mean is zero. A treatment is declared a success when the posterior probability that it beats the control exceeds 97.5%, so with no effect and no borrowing the false-positive rate would be 2.5%.

The exchangeable model treats the sixteen control means as draws from one normal distribution with a spread τ, under a half-Cauchy prior of scale 0.05 on τ, and the current control as one more draw. The line adds a slope in time. The change point says that the earlier trials from some position onwards share the current trial’s level and those before it have a level of their own, with the same spread τ about each. The position is unknown and gets a uniform prior over all sixteen possibilities, one of which — every earlier trial in the current regime — is the exchangeable model. With flat priors on the two levels, each position’s evidence is the likelihood of two groups with their means integrated out, and the current control’s prior is the later group’s mean. Everything is integrated over a grid on τ and summed over positions exactly.

In the stepped histories, the first eight earlier trials sit a step below the last eight, and the last eight sit at the current level. That is the direction in which borrowing flatters a treatment: a history partly lower than the current control pulls the control’s estimate down and the treatment’s advantage up.

What a step of a fifth looks like

The units hide how ordinary these changes are. In a trial of blood pressure with a patient-to-patient standard deviation of fifteen millimetres of mercury, a step of a fifth of a standard deviation is three millimetres — the size of change a new first-line drug or a tighter target in routine care produces in the patients who enter trials as controls. In each earlier trial of two hundred the control mean is measured to about a millimetre, so the step is three standard errors of a single trial’s control: plainly there if anybody plots the sixteen means in the order they were run, and invisible in a table that reports them as a pooled estimate and a spread.

That is the case the three models were run on, and it is also the case a reader of a borrowing analysis is least able to check, because the analyses that borrow most are the ones that report the earlier trials as an unordered set. A tenth of a standard deviation — a millimetre and a half — is within a single trial’s noise and only shows across several, and half a standard deviation is a change nobody would miss and nobody would borrow across.

The bias each model puts into the effect

The bias each model puts into the treatment effect when sixteen earlier controls stepped onceSixteen earlier trials of two hundred, the first eight a step below the last eight, which sit at the current control's level; the current trial has fifty patients an arm and no treatment effect; 1,500 trials at each step. Exchangeable trials: -0.003, 0.020, 0.041, 0.056, 0.063, 0.065, 0.054. A line in time: -0.003, -0.016, -0.029, -0.041, -0.052, -0.067, -0.075. A change point: -0.003, 0.003, 0.003, -0.001, -0.002, -0.003, -0.003 — at steps of 0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5.-0.05000.05000.1000.2000.3000.4000.500the step between the first eight earlier trials and the last eight, in standard deviationsbias of the effect estimateexchangeable trialsa line in timea change point1,500 trials a point; half-Cauchy prior, scale 0.05above zero flatters the treatment
Fig. 1 The bias of the treatment effect, with no true effect, against the size of the step between the first eight earlier trials and the last eight. The exchangeable model flatters the treatment, the line penalises it, and the change point is unbiased at every step size.

The three models err in three different ways. The exchangeable model’s bias rises with the step, to 0.041 at a step of a tenth of a standard deviation and 0.063 at a fifth, and then levels off and falls slightly — 0.054 at half — because a large step convinces it the trials are far apart and it stops borrowing. The line’s bias has the opposite sign and keeps growing: −0.029 at a tenth, −0.052 at a fifth, −0.075 at half. It reads the step as a slope and extrapolates the slope a trial past the end of the history, predicting a current control above the level the last eight trials actually sit at.

The change point is unbiased throughout, between −0.003 and 0.003 at every step from none to half a standard deviation. That is the earlier essay’s result for the line, transposed: a model whose structure matches the history’s removes the bias the history causes, whatever its size.

The patients each model borrows

The control patients each model borrows from sixteen earlier trials that stepped once. Sixteen earlier trials of two hundred, the first eight a step below the last eight, which sit at the current control's level; the current trial has fifty patients an arm and no treatment effect; 1,500 trials at each step. Exchangeable trials: 777, 685, 460, 236, 111, 41, 14. A line in time: 375, 368, 344, 306, 256, 151, 46. A change point: 464, 430, 400, 449, 512, 561, 579 — at steps of 0, 0.05, 0.1, 0.15, 0.2, 0.3, 0.5.
Fig. 2 The control patients each model borrows — one over the current control’s posterior variance, less its own fifty — against the step’s size. The exchangeable model and the line borrow less as the step grows; the change point borrows more.

With no step at all the exchangeable model borrows 777 control patients, the line 375 and the change point 464: the change point pays for having considered sixteen places for a step that is not there — 312 of the exchangeable model’s patients, where the line’s slope costs it 401.

As the step grows the two misfitting models borrow less and less. At a fifth of a standard deviation the exchangeable model borrows 111 patients and the line 256; at half, 14 and 46. Each sees a history that does not fit its shape, puts the misfit into a larger spread, and a larger spread borrows less. The change point does the reverse: 400 at a tenth, 512 at a fifth and 579 at half. A larger step is easier to locate, and once it is located the eight later trials are a clean, agreeing history at exactly the current level, from which the model borrows confidently. The right model of a history does not merely avoid the bias; it recovers the borrowing the wrong models gave up, and more of it the clearer the step.

Where the step was

How sure the change-point model is of where the step was, by the step's size. The change-point model's posterior probability that the current regime began at the ninth of sixteen earlier trials, where it did, averaged over 1,500 histories at each step size: 0.06 at 0.05, 0.19 at 0.1, 0.43 at 0.15, 0.67 at 0.2, 0.92 at 0.3, 1.00 at 0.5. With sixteen positions to choose among, a uniform prior puts 1⁄16 on each.
Fig. 3 The change-point model’s posterior probability that the current regime began at the ninth of the sixteen earlier trials, where it did, averaged over histories, against the step’s size. A uniform prior puts one sixteenth on each position.

How well the step is located is set by its size against the noise in each earlier trial’s control mean, whose standard error is 0.071 at two hundred patients. At a step of a twentieth of a standard deviation the model puts 0.06 of its belief on the true position, barely above the prior’s one sixteenth; at a tenth, 0.19; at a fifth, 0.67; at three tenths, 0.92; and at half, all of it. A step of one and a half standard errors is located a fifth of the time, one of nearly three standard errors two thirds of the time.

The model does not need to locate the step to be unbiased. At a tenth of a standard deviation it is unsure of the position and still unbiased, because its uncertainty is spread over positions near the truth, and every one of those keeps the current regime’s trials apart from the earlier ones. What locating the step buys is borrowing: an unsure model averages over positions that keep only a few trials, and borrows 400 patients; a sure one keeps all eight and borrows 579.

Rates and power

The bias translates into the rate at which a treatment with no effect is declared a success. With no change, the three models sit at 2.7%, 2.3% and 2.5%. At a step of a tenth the exchangeable model’s rate is 4.1% and at a fifth 4.9%, nearly double the nominal, while the line’s falls to 1.8% and 1.1% — too strict, which is not a safety but a cost. The change point stays at 2.3% and 2.2%.

The cost of the line’s strictness is power. With a true effect of half a standard deviation and a step of a tenth, the line has 88.4% power, the change point 92.5% and the exchangeable model 95.6% — the last inflated by the same bias that inflates its false positives, and so not a power anybody should want. With no step the three are 93.2%, 92.1% and 92.3%: the change point gives up less than a point against the exchangeable model when the history is flat, and buys four points against the line when it stepped.

Where in the history the step falls

A step in the middle of the history is the easy case: eight trials on each side, so both levels are well estimated. A step near the end leaves few trials at the current level.

At a step of a tenth with only the last two earlier trials at the current level, the change point’s rate rises to 3.9%, the line’s to 4.2% and the exchangeable model’s to 6.5%. With only two trials in the current regime the change point cannot tell a step two trials from the end from noise, its posterior on the true position is 0.15, and it borrows 282 patients against the 400 it borrows when the step is in the middle. With the last four at the current level it is at 3.2%; with twelve, 2.5%, the nominal rate. A recent change is the case that matters most in practice, since the standard of care that moved the controls is usually the one still in force, and it is the case where sixteen earlier trials say least.

Each model is safe against its own history

How often each model declares a treatment with no effect a success, when the history did not change, stepped once or drifted. Sixteen earlier trials, 1,500 histories each; the false-positive rate at a nominal one-sided 2.5%. No change: exchangeable trials 2.7%, a line in time 2.3%, a change point 2.5%. A step of 0.1: exchangeable trials 4.1%, a line in time 1.8%, a change point 2.3%. A step of 0.2: exchangeable trials 4.9%, a line in time 1.1%, a change point 2.2%. A drift of 0.02 a trial: exchangeable trials 9.0%, a line in time 2.3%, a change point 5.7%. A drift of 0.05 a trial: exchangeable trials 7.1%, a line in time 2.3%, a change point 8.7%.
Fig. 4 The false-positive rate of each model, at a nominal 2.5%, on histories with no change, a step of a tenth or a fifth, and a steady drift of 0.02 or 0.05 a trial. The line holds its rate on any drift and the change point on any step; neither holds it on the other’s shape.

The change point is not a general repair. On a history that drifted steadily, it reads the drift as a step and keeps only the trials it places in the current regime — but in a drifting history those trials are still below the current level. At a drift of 0.02 a trial its rate is 5.7%, better than the exchangeable model’s 9.0% and far worse than the line’s 2.3%. At a drift of 0.05 a trial it is 8.7%, worse than the exchangeable model’s 7.1%: the steeper drift makes it confident that only the last few trials share the current level, and the last few still sit below it.

So each model is safe against exactly one shape. The line holds its rate on any straight drift and penalises a step; the change point holds it on any single step and fails on a drift; the exchangeable model holds it only when nothing moved. That is the situation a break that was looked for described for a single series: a search for a break finds one whether or not the series has one, and the model that searches pays for the searching when the series did something else. Here the price is paid in borrowed patients when nothing changed and in bias when the change was a drift. It is also the pattern one population or two found in a hierarchical model fitted to two clusters: every model reports its own fit as adequate, and the shape it assumed away is visible only to a model that allows it.

The change-point model is a search over positions with the search’s cost built in. Its uniform prior spreads one sixteenth over each position, and the posterior concentrates only where the data favour a split; on a flat history no position is favoured and the model ends up averaging over all of them, including positions that keep only the last one or two earlier trials. That average is what costs it 312 patients on a history that never moved, and it is the same price a split that depends on the order put on searching a series for the point where it divides: a search is paid for on the data where there is nothing to find.

The price is fair in a way the alternatives’ prices are not. The exchangeable model pays nothing on a flat history and an uncontrolled bias on a stepped one; the line pays 401 patients on a flat history and a bias of the opposite sign on a stepped one. The change point pays 312 on the flat history and nothing in bias on the stepped one, and from a step of 0.15 upwards it borrows more than either of the others. Of the three, it is the only one whose cost is bounded in advance and whose failure on its own shape of history is zero.

What a borrowing analysis should state about its history

Which shapes of change it allows. An exchangeable model assumes the control never moved; a line, that it moved steadily; a change point, that it moved once. Each is wrong in a known direction on the others’ histories, and the report should say which risk it took.

Where the change point was placed, and how surely. A posterior concentrated on one position is a finding about the history — a date at which the control changed — and one spread over many positions is a model averaging over its own ignorance, borrowing less as a result.

How many trials sit after the change. Two trials in the current regime are a thin basis for borrowing however clean the step, and that is the case a recent change of practice produces.

Both of the other analyses. The three models agree on a history that did not move; where they disagree the gap is an estimate of the bias the wrong ones carry, and a history that agrees with itself is the warning that agreement among the earlier trials alone is not evidence that the current trial belongs with them.

Counted, and summed exactly within each trial

Every rate, bias and borrowed count is counted over 1,500 simulated current trials at each setting, each with its own sixteen earlier trials of two hundred controls and fifty patients an arm in the current trial, one stated seed per trial. Within each trial the analysis is exact: a grid of two hundred values of τ from zero to three under a half-Cauchy prior of scale 0.05, and for the change point a sum over all sixteen positions, with the two levels integrated out under flat priors. The standard error of a rate near 2.5% is about 0.4 points, so the differences between the change point and the nominal rate here are within the count; the exchangeable model’s 4.9% and the line’s 1.1% at a step of a fifth are not. Only steps upward towards the current level are measured, which is the direction that flatters a treatment; a step downward reverses the signs of the biases and, by the symmetry of every prior and every error here, leaves their sizes.

Still open: a history whose shape is unknown

Each model is right for one shape, and a real history’s shape is not announced. The natural response is to let the data choose: weight the three models by how well each explains the history, or use a model that contains all three — a line with a change point in it, or a spread that is allowed to differ before and after a change. Either borrows less than the right model would have and should be safe against all three shapes, and the measurement that decides whether the trade is worth it is the one the relation the table has to estimate made for a league table: what choosing the shape from the data costs against knowing it.

Whether sixteen earlier trials can tell a step of a fifth of a standard deviation from a drift of 0.02 a trial often enough to choose between them, what an averaged model’s false-positive rate is on each shape, and how many patients it borrows when nothing moved, are counts the same models and draws can make. Borrowing towards a line and a control borrowed from the last trial mark the two ends of what is being chosen between: borrowing shaped by a model of how the history moved, and borrowing from the one trial that is closest to now.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasChange pointError rateThe half-Cauchy priorHierarchical modelHistorical controlModel misspecificationPartial poolingPriorPrior sensitivityStatistical powerVariance components