A history that agrees with itself
Worth reading first: What a prior is worth · The weight that decides.
A control borrowed from the last trial found that with one earlier trial the prior on the spread between trials is the whole borrowing rule, because one difference between two control means says almost nothing about that spread. A half-normal prior never lets go of a history that disagrees; a half-Cauchy of the same scale does, after an error peak of about 11% at a moderate drift. It ended on the obvious escape: more history. With several earlier trials of the same control the spread between trials is estimated from data, and the prior should matter less, as it did for eight groups.
It does matter less. The borrowing does not get safer. It gets far more dangerous, and the reason is that the history’s agreement with itself is exactly the evidence that licenses pooling it with a current trial that does not agree.
The prior built from the history
With K earlier trials the analysis is the same hierarchical model: every trial’s control mean is a draw from . At each value of τ, the earlier trials give an estimate of μ and its variance, and that estimate plus is a prior for the current control — the meta-analytic predictive prior, the standard construction for borrowing from a history of trials. The current control’s data update it, and τ is integrated over the posterior that all K + 1 trials give it. With one earlier trial this is exactly the two-trial analysis of the essay before, and the two calculations agree to the last digit.
What changes with K is what the history says about τ. One earlier trial says nothing: the spread’s posterior from it is the prior. Sixteen earlier trials that agree say a great deal.
The 90th percentile of the spread’s posterior is 0.181 standard deviations from one earlier trial — the prior’s own tail — 0.084 from four and 0.063 from sixteen. The history has learned, correctly, that its trials are alike. It has learned nothing about whether the current trial is like them, because that is a question about one more trial, and one trial is one difference.
Why agreement tightens the grip
The current control’s estimate is a weighted combination of its own mean and the history’s, and the weight depends on how large τ is believed to be. With one earlier trial, a current control far from the history is best explained by a large τ, which the heavy-tailed prior allows, and the borrowing stops. With sixteen earlier trials, a large τ is no longer an available explanation: sixteen trials agreeing to within their sampling error make a large spread between trials implausible, and the only remaining explanation for a far current control is chance in the current trial. So the analysis keeps borrowing, harder the further the current trial drifts, until the drift is so large that even chance cannot cover it.
The picture is the one-trial picture stretched. With one earlier trial the shift peaks at about 0.17 and falls to 0.029 at a disagreement of 1.5 standard deviations. With four it peaks near 0.36 and falls to 0.070. With sixteen it tracks the disagreement almost one for one, reaching about 0.84 before it turns, and at 1.5 standard deviations it is still 0.266 — nine times the one-trial shift. A current control half a standard deviation from the history is pulled 0.482 of the way back by sixteen agreeing trials, against 0.131 by one.
This is what the essay on one population or two found from the other side: a hierarchical model reports its estimated spread, weights every group by it, and does not flag the group the spread does not describe. Sixteen trials describe themselves well. The seventeenth is described by the same number.
The error that follows
The false-positive rates follow the shift. With one earlier trial, the half-Cauchy’s worst rate is 10.5%, near a drift of 0.4 standard deviations, and at 1.5 it is back to 3.5%. With four earlier trials the worst is 28.1%. With sixteen it is 80.7%, near a drift of 0.6, and at 1.5 standard deviations it is still 25.5%. The half-normal is worse throughout — 88.1% at worst with sixteen trials and 73.3% at a drift of 1.5 — but the point of the comparison is that the half-Cauchy, which rescued the one-trial analysis, does not rescue this one. Its tail was an escape route from the prior; the history has closed it by making the posterior for τ, not the prior, the thing that decides.
The drifts in question are not exotic. Half a standard deviation is the treatment effect this trial was designed to detect, and a current control that differs from a historical one by that much is exactly what a change in the standard of care, the case mix or the assessment would produce.
Four trials, where the two regimes meet
Four earlier trials sit between the regime in which the prior decides and the one in which the history does, and they show both at once.
With four earlier trials the half-Cauchy’s error peaks at 28.1% near a drift of half a standard deviation and falls to 5.4% at 1.5: it still lets go, because four trials leave enough doubt about the spread for a large one to explain a large disagreement, but it needs a larger disagreement to do it and passes through a higher peak on the way. The half-normal peaks at 56.4% and stays near 47.5%. The vague fifth peaks at 12.1% and returns to 2.7%.
So the transition is gradual, and it runs in the direction nobody wants. Each earlier trial added to an agreeing history narrows the spread’s posterior a little, removes a little of the tail’s room to explain a conflict, and raises the peak a little. There is no number of earlier trials at which the history becomes safe; there is only a number at which the prior’s tail stops being able to help, and it is small.
What the history buys
The attraction is real. At no drift and an effect of half a standard deviation the trial alone has 70.5% power. Borrowing under the half-Cauchy from one earlier trial raises it to 83.9% and borrows 66.6 control patients; from four, 90.3% and 241.7 patients; from sixteen, 94.3% and 761.3 patients — a control arm of fifty analysed as if it had eight hundred. How many subjects sized a trial at sixty-four an arm for 80% power at this effect; borrowing on this scale runs a control arm of fifty and reports the precision of one of eight hundred and eleven, sixteen times larger, on the strength of an assumption the fifty cannot check.
That is the reason histories are pooled, and it is also the measure of the exposure. A control arm treated as eight hundred patients is eight hundred patients’ worth of confidence in a number that describes other trials, and when the current trial differs, the confidence is spent on the wrong number.
When the earlier trials genuinely differ from each other — a true spread between trials of 0.1 standard deviations, which is a modest heterogeneity — the history estimates that too and borrows less: 152.3 patients from sixteen trials rather than 761.3, and power 89.5%. Its worst false-positive rate is 41.1%, lower than the agreeing history’s because a history that shows its own spread is less willing to pool, and still sixteen times the nominal level. A history that agrees with itself is the most dangerous kind, because agreement is what earns the most borrowing.
The plug-in version is worse again
Everything above integrates over the spread between trials. A simpler analysis estimates the spread once and plugs the estimate in, which is what the plug-in forgets warned against for eight groups: it treats an estimate as known and produces intervals too narrow by the uncertainty it dropped. With a history of agreeing trials the dropped uncertainty is precisely the part that matters. The usual moment estimate of the spread is exactly zero on a large share of datasets whose trials are alike, and at an estimated spread of zero the plug-in analysis pools the current control with the history outright. Pooling outright with a single earlier trial of two hundred already declared 73.1% of null treatments successes at a drift of half a standard deviation; pooling with sixteen gives the history more weight still, and the plug-in analysis adopts that rule on every dataset where the estimate happens to land on zero.
The full analysis is better than that only by the width of the spread’s posterior near zero, and sixteen agreeing trials have made that width small. The two analyses converge as the history grows, towards the one that pools.
A way out that is priced
The standard repair is to stop letting the history’s agreement close every exit. Put most of the current control’s prior on the history-derived component and a share on a vague component — here a normal centred on the history’s prediction with the spread of a single patient. When the current trial agrees with the history, the vague component is outweighed and the borrowing proceeds. When it disagrees, the vague component explains the current data better than a tight history-based prior can, the posterior moves its weight there, and the current trial is analysed essentially on its own data.
With a fifth of the prior on the vague component, the sixteen-trial analysis has a worst false-positive rate of 16.9%, against 80.7% without it, and is back to 3.7% at a drift of 1.5 standard deviations. Its power at no drift falls from 94.3% to 91.0%, and it borrows 325.0 patients instead of 761.3. A tenth gives a worst rate of 22.1% and power of 92.6%; a half, 9.2% and 86.4%.
The vague share is doing what the heavy tail did for one earlier trial: it gives the analysis a second explanation for a disagreement, one the history’s own agreement cannot rule out, because the vague component is not a statement about the spread between trials at all. It is a statement that the current trial might be something else. The curve shows the price of saying so. No point on it reaches both the borrowed power and the nominal error — the same theorem holds for any number of earlier trials — but a fifth of the prior buys back four fifths of the excess error for 3.3 points of power.
What the vague share is a statement about
The vague component can look like a fudge — a share chosen to make an error curve behave. It has a more direct reading. A fifth of the prior on a vague component is a prior probability of one in five that the current trial is not exchangeable with the history at all: that whatever made the earlier controls alike does not apply to this one. What a prior is worth turned the question of a prior’s influence into a count of observations; this turns the question of borrowing into a probability that can be argued about with the evidence a protocol has — how similar the populations are, whether the standard of care has moved, whether the endpoint was assessed the same way.
Stated that way, the choice is not arbitrary and not small. A history of sixteen trials run over a decade, borrowed for a trial in a population whose care has changed since, deserves a larger share than a history of four trials run last year at the same centres. The error curve says what each share costs; the protocol has to say why the share is believed.
What changes with the number of earlier trials, and what does not
The prior on the spread matters less. With sixteen trials the half-normal and half-Cauchy analyses behave alike, both badly; with one they were the difference between a plateau of 31% and a return to 3.5%. The history has taken over the job the prior did, which is the behaviour a reader of the prior on the spread would expect.
The exposure to a drifting current trial grows. The history’s consistency is evidence about the history, and a hierarchical model spends it on the current trial, where it is not evidence of anything. The worst false-positive rate under the half-Cauchy runs 10.5%, 28.1%, 80.7% for one, four and sixteen earlier trials.
What the design needs to say is the same in both cases. The error curve over the current trial’s drift, its peak and where it sits; the number of patients the analysis borrowed in the trial that ran; and — with a long history — the share of the prior that allows the current trial to be different, which is now the borrowing rule in the way the prior’s tail was with one earlier trial.
Why this is a question about the current trial, not the history
It is tempting to read the result as a warning about bad histories, and it is not one. Every history here is a perfect history: the earlier trials are drawn with no spread at all, the model describes them exactly, and the analysis estimates their common control mean with the precision sixteen trials of two hundred deserve. What goes wrong is entirely in the step from the history to the current trial, which the model takes on the assumption that the current trial is one more draw from the same population as the others.
That assumption cannot be checked by the history, however long, and it can be checked by the current trial only with the power of fifty patients against a standard error of 0.143 on its difference from the history’s mean — barely better than the 0.158 that limited a single earlier trial, because the current trial’s own fifty patients are most of that error either way. When borrowing goes wrong found that nothing in a hierarchical model’s output flags the group that was never from the population. A long agreeing history makes the output more confident and changes nothing else.
What is measured here and what is not
Under a half-Cauchy prior of scale 0.05, the worst false-positive rate of a trial borrowing its control is 10.5% from one earlier trial, 28.1% from four and 80.7% from sixteen that agree, and power at no drift rises from 83.9% to 94.3% over the same range, against 70.5% without borrowing.
A fifth of the prior on a vague component brings the sixteen-trial worst rate to 16.9% and power to 91.0%.
When the earlier trials genuinely differ by a spread of 0.1, sixteen of them lend 152.3 patients and the worst rate is 41.1%.
The rates are counted over fifteen hundred simulated trials a point, fifty patients an arm and two hundred in each earlier control, a normal outcome with known unit variance and a one-sided posterior threshold of 97.5%, with every posterior exact on a grid of three hundred values of the spread. The one-trial case agrees with the two-trial analysis of the essay before to twelve digits.
Not measured: histories whose trials differ in size, in time or in a trend, where a model that weights recent trials more is the usual proposal; vague components other than a single patient’s worth; and binary outcomes, where the meta-analytic predictive prior is usually applied to log-odds and the arithmetic is approximate.
Still open: a history with a trend
Every history here is exchangeable: the earlier trials are interchangeable draws, and the current trial is one more. Histories of real controls usually are not. Standards of care improve, populations change and assessments drift, so the earlier trials sit along a trend and the current trial, being the latest, sits at its end — beyond every earlier trial rather than among them.
A model that treats a trending history as exchangeable estimates a spread that is really a slope, and borrows towards the history’s average when the current trial should be predicted from its end. How large the resulting bias is for a plausible trend, whether a model with a trend in time repairs it without spending the borrowing, and how many earlier trials a trend needs before it can be estimated at all, are the measurements a history of real controls would need and none of them has been made here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a two-unit study should report — both name hierarchical model, partial pooling, prior sensitivity, variance components
- Eight groups, one population — both name hierarchical model, partial pooling, prior
- Pooling a proportion — both name hierarchical model, partial pooling, prior
- The fewest groups that can borrow — both name hierarchical model, partial pooling, variance components
- The prior the data estimates — both name partial pooling, prior, variance components
- Where the borrowing goes — both name hierarchical model, partial pooling, variance components
Named objects
A flat tag is an object no other essay names yet.
Error rateThe half-Cauchy priorHierarchical modelPartial poolingPriorPrior sensitivityStatistical powerVariance components