The spread, and its own uncertainty

A history that agrees with itself

Borrowing a control from sixteen earlier trials that agree with each other should be safer than borrowing from one, and it is the opposite. Sixteen agreeing trials estimate the spread between trials as small, and a small spread estimated confidently is a licence to pool: under a half-Cauchy prior that let go of a single disagreeing trial, a current control 0.6 standard deviations from a sixteen-trial history turns a treatment with no effect into a success 80.7% of the time, against 10.5% at worst with one earlier trial. Moving a fifth of the prior onto a vague component brings the worst case to 16.9% and keeps power at 91.0%, against 70.5% without borrowing.

Worth reading first: What a prior is worth · The weight that decides.

A control borrowed from the last trial found that with one earlier trial the prior on the spread between trials is the whole borrowing rule, because one difference between two control means says almost nothing about that spread. A half-normal prior never lets go of a history that disagrees; a half-Cauchy of the same scale does, after an error peak of about 11% at a moderate drift. It ended on the obvious escape: more history. With several earlier trials of the same control the spread between trials is estimated from data, and the prior should matter less, as it did for eight groups.

It does matter less. The borrowing does not get safer. It gets far more dangerous, and the reason is that the history’s agreement with itself is exactly the evidence that licenses pooling it with a current trial that does not agree.

How often a trial borrowing from 16 earlier trials that agree declares a treatment with no effect a success, by its own control's driftFifty patients an arm, each earlier control of two hundred, a one-sided 2.5% threshold. With 16 earlier trials, the half-normal prior peaks at 88.1%, the half-Cauchy at 80.7% near a drift of 0.6, and the half-Cauchy with a fifth of the prior on a vague component at 16.9%. At a drift of 1.5 the three read 73.3%, 25.5% and 3.7%.00.2500.5000.750100.50011.50drift of the current control above the earlier trials', in standard deviationsshare of trials with no effect declared a successhalf-normal, scale 0.05half-Cauchy, scale 0.05half-Cauchy with a fifth vague1,500 trials a point, earlier trials identical in truthdashed red: the nominal 2.5%
Fig. 1 The share of trials of fifty patients an arm that declare a treatment with no effect a success, borrowing their control from sixteen earlier trials of two hundred that agree with each other, against how far the current control sits above the history. A half-normal and a half-Cauchy prior of scale 0.05 on the spread between trials, and the half-Cauchy with a fifth of the prior moved to a vague component. The slider sets the number of earlier trials.

The prior built from the history

With K earlier trials the analysis is the same hierarchical model: every trial’s control mean is a draw from N(μ,τ2)N(\mu, \tau^2). At each value of τ, the earlier trials give an estimate of μ and its variance, and that estimate plus τ2\tau^2 is a prior for the current control — the meta-analytic predictive prior, the standard construction for borrowing from a history of trials. The current control’s data update it, and τ is integrated over the posterior that all K + 1 trials give it. With one earlier trial this is exactly the two-trial analysis of the essay before, and the two calculations agree to the last digit.

What changes with K is what the history says about τ. One earlier trial says nothing: the spread’s posterior from it is the prior. Sixteen earlier trials that agree say a great deal.

What one, four and sixteen agreeing earlier trials say about the spread between trials. Under a half-Cauchy prior of scale 0.05, the 90th percentile of the spread's posterior is 0.181 from one earlier trial (which says nothing beyond the prior), 0.084 from four and 0.063 from sixteen, drawn with no true spread.
Fig. 2 The posterior for the spread between trials from the earlier trials alone, under a half-Cauchy prior of scale 0.05, for one, four and sixteen earlier trials drawn with no true spread between them.

The 90th percentile of the spread’s posterior is 0.181 standard deviations from one earlier trial — the prior’s own tail — 0.084 from four and 0.063 from sixteen. The history has learned, correctly, that its trials are alike. It has learned nothing about whether the current trial is like them, because that is a question about one more trial, and one trial is one difference.

Why agreement tightens the grip

The current control’s estimate is a weighted combination of its own mean and the history’s, and the weight depends on how large τ is believed to be. With one earlier trial, a current control far from the history is best explained by a large τ, which the heavy-tailed prior allows, and the borrowing stops. With sixteen earlier trials, a large τ is no longer an available explanation: sixteen trials agreeing to within their sampling error make a large spread between trials implausible, and the only remaining explanation for a far current control is chance in the current trial. So the analysis keeps borrowing, harder the further the current trial drifts, until the drift is so large that even chance cannot cover it.

How far a history of agreeing trials pulls the current control's estimate, against how far the current control sits from it. Under a half-Cauchy prior of scale 0.05 on the spread, a current control 1.5 standard deviations above the history is pulled back by 0.029 with one earlier trial, 0.070 with four and 0.266 with sixteen; with sixteen and a fifth of the prior on a vague component, by 0.029. At 0.5 the four read 0.131, 0.362, 0.482 and 0.043.
Fig. 3 How far the current control’s estimate is pulled towards a history sitting exactly at zero, against how far the current control’s observed mean sits from it, under a half-Cauchy prior of scale 0.05: one, four and sixteen earlier trials, and sixteen with a fifth of the prior on a vague component.

The picture is the one-trial picture stretched. With one earlier trial the shift peaks at about 0.17 and falls to 0.029 at a disagreement of 1.5 standard deviations. With four it peaks near 0.36 and falls to 0.070. With sixteen it tracks the disagreement almost one for one, reaching about 0.84 before it turns, and at 1.5 standard deviations it is still 0.266 — nine times the one-trial shift. A current control half a standard deviation from the history is pulled 0.482 of the way back by sixteen agreeing trials, against 0.131 by one.

This is what the essay on one population or two found from the other side: a hierarchical model reports its estimated spread, weights every group by it, and does not flag the group the spread does not describe. Sixteen trials describe themselves well. The seventeenth is described by the same number.

The error that follows

The false-positive rates follow the shift. With one earlier trial, the half-Cauchy’s worst rate is 10.5%, near a drift of 0.4 standard deviations, and at 1.5 it is back to 3.5%. With four earlier trials the worst is 28.1%. With sixteen it is 80.7%, near a drift of 0.6, and at 1.5 standard deviations it is still 25.5%. The half-normal is worse throughout — 88.1% at worst with sixteen trials and 73.3% at a drift of 1.5 — but the point of the comparison is that the half-Cauchy, which rescued the one-trial analysis, does not rescue this one. Its tail was an escape route from the prior; the history has closed it by making the posterior for τ, not the prior, the thing that decides.

The drifts in question are not exotic. Half a standard deviation is the treatment effect this trial was designed to detect, and a current control that differs from a historical one by that much is exactly what a change in the standard of care, the case mix or the assessment would produce.

Four trials, where the two regimes meet

Four earlier trials sit between the regime in which the prior decides and the one in which the history does, and they show both at once.

How often a trial borrowing from 4 earlier trials that agree declares a treatment with no effect a success, by its own control's drift. Fifty patients an arm, each earlier control of two hundred, a one-sided 2.5% threshold. With 4 earlier trials, the half-normal prior peaks at 56.4%, the half-Cauchy at 28.1% near a drift of 0.5, and the half-Cauchy with a fifth of the prior on a vague component at 12.1%. At a drift of 1.5 the three read 47.5%, 5.4% and 2.7%.
Fig. 4 The same three analyses borrowing from four agreeing earlier trials. The half-Cauchy still lets go eventually, but later and from a higher peak than with one trial; the vague component holds the error down throughout.

With four earlier trials the half-Cauchy’s error peaks at 28.1% near a drift of half a standard deviation and falls to 5.4% at 1.5: it still lets go, because four trials leave enough doubt about the spread for a large one to explain a large disagreement, but it needs a larger disagreement to do it and passes through a higher peak on the way. The half-normal peaks at 56.4% and stays near 47.5%. The vague fifth peaks at 12.1% and returns to 2.7%.

So the transition is gradual, and it runs in the direction nobody wants. Each earlier trial added to an agreeing history narrows the spread’s posterior a little, removes a little of the tail’s room to explain a conflict, and raises the peak a little. There is no number of earlier trials at which the history becomes safe; there is only a number at which the prior’s tail stops being able to help, and it is small.

What the history buys

The attraction is real. At no drift and an effect of half a standard deviation the trial alone has 70.5% power. Borrowing under the half-Cauchy from one earlier trial raises it to 83.9% and borrows 66.6 control patients; from four, 90.3% and 241.7 patients; from sixteen, 94.3% and 761.3 patients — a control arm of fifty analysed as if it had eight hundred. How many subjects sized a trial at sixty-four an arm for 80% power at this effect; borrowing on this scale runs a control arm of fifty and reports the precision of one of eight hundred and eleven, sixteen times larger, on the strength of an assumption the fifty cannot check.

That is the reason histories are pooled, and it is also the measure of the exposure. A control arm treated as eight hundred patients is eight hundred patients’ worth of confidence in a number that describes other trials, and when the current trial differs, the confidence is spent on the wrong number.

When the earlier trials genuinely differ from each other — a true spread between trials of 0.1 standard deviations, which is a modest heterogeneity — the history estimates that too and borrows less: 152.3 patients from sixteen trials rather than 761.3, and power 89.5%. Its worst false-positive rate is 41.1%, lower than the agreeing history’s because a history that shows its own spread is less willing to pool, and still sixteen times the nominal level. A history that agrees with itself is the most dangerous kind, because agreement is what earns the most borrowing.

The plug-in version is worse again

Everything above integrates over the spread between trials. A simpler analysis estimates the spread once and plugs the estimate in, which is what the plug-in forgets warned against for eight groups: it treats an estimate as known and produces intervals too narrow by the uncertainty it dropped. With a history of agreeing trials the dropped uncertainty is precisely the part that matters. The usual moment estimate of the spread is exactly zero on a large share of datasets whose trials are alike, and at an estimated spread of zero the plug-in analysis pools the current control with the history outright. Pooling outright with a single earlier trial of two hundred already declared 73.1% of null treatments successes at a drift of half a standard deviation; pooling with sixteen gives the history more weight still, and the plug-in analysis adopts that rule on every dataset where the estimate happens to land on zero.

The full analysis is better than that only by the width of the spread’s posterior near zero, and sixteen agreeing trials have made that width small. The two analyses converge as the history grows, towards the one that pools.

A way out that is priced

The standard repair is to stop letting the history’s agreement close every exit. Put most of the current control’s prior on the history-derived component and a share on a vague component — here a normal centred on the history’s prediction with the spread of a single patient. When the current trial agrees with the history, the vague component is outweighed and the borrowing proceeds. When it disagrees, the vague component explains the current data better than a tight history-based prior can, the posterior moves its weight there, and the current trial is analysed essentially on its own data.

What a vague component in the prior buys back, borrowing from 16 agreeing earlier trials. With no vague component the analysis has power 94.3% and a worst false-positive rate of 80.7%. A vague share of 0.1 gives 92.6% and 22.1%; 0.2, 91.0% and 16.9%; 0.5, 86.4% and 9.2%. The trial alone has 70.5%.
Fig. 5 Power at no drift against the worst false-positive rate over the drifts, for the sixteen-trial analysis with a vague share of the prior running from nothing to a half. The labels are the vague share.

With a fifth of the prior on the vague component, the sixteen-trial analysis has a worst false-positive rate of 16.9%, against 80.7% without it, and is back to 3.7% at a drift of 1.5 standard deviations. Its power at no drift falls from 94.3% to 91.0%, and it borrows 325.0 patients instead of 761.3. A tenth gives a worst rate of 22.1% and power of 92.6%; a half, 9.2% and 86.4%.

The vague share is doing what the heavy tail did for one earlier trial: it gives the analysis a second explanation for a disagreement, one the history’s own agreement cannot rule out, because the vague component is not a statement about the spread between trials at all. It is a statement that the current trial might be something else. The curve shows the price of saying so. No point on it reaches both the borrowed power and the nominal error — the same theorem holds for any number of earlier trials — but a fifth of the prior buys back four fifths of the excess error for 3.3 points of power.

What the vague share is a statement about

The vague component can look like a fudge — a share chosen to make an error curve behave. It has a more direct reading. A fifth of the prior on a vague component is a prior probability of one in five that the current trial is not exchangeable with the history at all: that whatever made the earlier controls alike does not apply to this one. What a prior is worth turned the question of a prior’s influence into a count of observations; this turns the question of borrowing into a probability that can be argued about with the evidence a protocol has — how similar the populations are, whether the standard of care has moved, whether the endpoint was assessed the same way.

Stated that way, the choice is not arbitrary and not small. A history of sixteen trials run over a decade, borrowed for a trial in a population whose care has changed since, deserves a larger share than a history of four trials run last year at the same centres. The error curve says what each share costs; the protocol has to say why the share is believed.

What changes with the number of earlier trials, and what does not

The prior on the spread matters less. With sixteen trials the half-normal and half-Cauchy analyses behave alike, both badly; with one they were the difference between a plateau of 31% and a return to 3.5%. The history has taken over the job the prior did, which is the behaviour a reader of the prior on the spread would expect.

The exposure to a drifting current trial grows. The history’s consistency is evidence about the history, and a hierarchical model spends it on the current trial, where it is not evidence of anything. The worst false-positive rate under the half-Cauchy runs 10.5%, 28.1%, 80.7% for one, four and sixteen earlier trials.

What the design needs to say is the same in both cases. The error curve over the current trial’s drift, its peak and where it sits; the number of patients the analysis borrowed in the trial that ran; and — with a long history — the share of the prior that allows the current trial to be different, which is now the borrowing rule in the way the prior’s tail was with one earlier trial.

Why this is a question about the current trial, not the history

It is tempting to read the result as a warning about bad histories, and it is not one. Every history here is a perfect history: the earlier trials are drawn with no spread at all, the model describes them exactly, and the analysis estimates their common control mean with the precision sixteen trials of two hundred deserve. What goes wrong is entirely in the step from the history to the current trial, which the model takes on the assumption that the current trial is one more draw from the same population as the others.

That assumption cannot be checked by the history, however long, and it can be checked by the current trial only with the power of fifty patients against a standard error of 0.143 on its difference from the history’s mean — barely better than the 0.158 that limited a single earlier trial, because the current trial’s own fifty patients are most of that error either way. When borrowing goes wrong found that nothing in a hierarchical model’s output flags the group that was never from the population. A long agreeing history makes the output more confident and changes nothing else.

What is measured here and what is not

Under a half-Cauchy prior of scale 0.05, the worst false-positive rate of a trial borrowing its control is 10.5% from one earlier trial, 28.1% from four and 80.7% from sixteen that agree, and power at no drift rises from 83.9% to 94.3% over the same range, against 70.5% without borrowing.

A fifth of the prior on a vague component brings the sixteen-trial worst rate to 16.9% and power to 91.0%.

When the earlier trials genuinely differ by a spread of 0.1, sixteen of them lend 152.3 patients and the worst rate is 41.1%.

The rates are counted over fifteen hundred simulated trials a point, fifty patients an arm and two hundred in each earlier control, a normal outcome with known unit variance and a one-sided posterior threshold of 97.5%, with every posterior exact on a grid of three hundred values of the spread. The one-trial case agrees with the two-trial analysis of the essay before to twelve digits.

Not measured: histories whose trials differ in size, in time or in a trend, where a model that weights recent trials more is the usual proposal; vague components other than a single patient’s worth; and binary outcomes, where the meta-analytic predictive prior is usually applied to log-odds and the arithmetic is approximate.

Still open: a history with a trend

Every history here is exchangeable: the earlier trials are interchangeable draws, and the current trial is one more. Histories of real controls usually are not. Standards of care improve, populations change and assessments drift, so the earlier trials sit along a trend and the current trial, being the latest, sits at its end — beyond every earlier trial rather than among them.

A model that treats a trending history as exchangeable estimates a spread that is really a slope, and borrows towards the history’s average when the current trial should be predicted from its end. How large the resulting bias is for a plausible trend, whether a model with a trend in time repairs it without spending the borrowing, and how many earlier trials a trend needs before it can be estimated at all, are the measurements a history of real controls would need and none of them has been made here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Error rateThe half-Cauchy priorHierarchical modelPartial poolingPriorPrior sensitivityStatistical powerVariance components