A control borrowed from the last trial
Worth reading first: What a prior is worth · The weight that decides.
A prior on the spread measured what the prior on a hierarchical model’s population spread does with eight groups: it decides the answer where the groups are nearly identical, and its influence fades, without vanishing, as they separate. Eight groups is a comfortable number for that question. The setting where a spread between groups is estimated from the fewest possible groups is also one of the commonest in practice, and in it the prior does not fade at all.
A trial comparing a treatment with a control can borrow control patients from an earlier trial of the same control. The earlier trial’s control mean is informative about the current one if the two populations, protocols and eras are similar, and misleading if they are not. The standard way to borrow honestly is to treat the two control means as draws from one population with a spread τ between trials, estimate τ, and borrow in proportion to how small it looks. With two trials there is exactly one difference to estimate τ from.
The model, and what one difference can say
Write the two control means as , from the earlier trial of two hundred patients, and , from the current trial’s fifty. Both are noisy measurements of their own trial’s true control mean, and the two true means are draws from with unknown. At a given τ the earlier trial is simply one more measurement of the current control’s mean, with variance , and the current control’s posterior is the precision-weighted combination of the two. At τ = 0 the earlier trial gets weight 0.800 — its two hundred patients against the current fifty — and the analysis behaves as if the control arm had two hundred and fifty patients. As τ grows the weight falls towards zero.
Everything therefore depends on τ, and the data carry one number about it: the difference , whose variance is . A single difference cannot distinguish “the trials differ by 0.2 because τ is 0.15” from “the trials differ by 0.2 by chance and τ is zero”. The likelihood for τ is nearly flat over any range the problem cares about, so the posterior for τ is nearly the prior.
There is a sharper way to say it. At large τ the likelihood falls only like , because one difference is one observation of a variance. With eight groups it fell like , which is why a flat prior on the spread was proper there. With two trials diverges: a flat prior on τ gives no posterior at all, the same pathology the 1/τ prior had with eight groups, now arriving for the prior that was safe. Every borrowing analysis of one earlier trial has to use a proper prior on τ, and that prior is the borrowing rule.
Two units, again
The situation is not new to this site; what is new is the use it is put to. A level with two units measured a variance estimated from two units as a scaled chi-square on one degree of freedom — exactly zero on a quarter of studies, spread over a factor of a hundred and seventy between its deciles. What a two-unit study should report found that a random-effects interval for the mean of a population sampled at two sites covers at its level only by being more than eleven times wider than the fixed-effect one. And the fewest groups that can borrow found that the estimator which shrinks towards its own data’s mean does not shrink at all below three groups, and at two expands.
Those were all descriptions of what two units can say about a spread. Borrowing a control turns the description into a decision: the spread between two trials is not reported, it is used, and the current trial’s conclusion inherits whatever the prior filled the silence with. The arithmetic that made two units a poor basis for an interval makes them a poor basis for a weight, and the weight is what the trial’s result is built from.
Two priors with one scale
The priors usually proposed for a between-trial spread are half-normal and half-Cauchy, each with a scale that says how large a spread between trials is thought plausible. Set both to a scale of 0.05 standard deviations — a strong belief that the earlier trial’s control is essentially the current one — and they agree about the typical spread. They disagree about the atypical one.
Both put most of their mass below 0.10. The half-normal says a spread above 0.5 has probability — not unlikely but, for any purpose, impossible. The half-Cauchy says 0.063: unlikely, and entirely possible. Nothing in the two priors’ descriptions, “scale 0.05”, tells a reader which of these beliefs a protocol has committed to, and it is the only one that matters when the trials disagree.
What each prior does with a disagreement
The mechanism is visible without simulating anything. For each observed difference between the two control means, the analysis moves the current control’s estimate some distance towards the earlier one.
Pooled outright, the estimate moves by 0.8 of every difference, however large. Under the half-normal prior it moves by 0.297 at a difference of half a standard deviation and by 0.292 at a difference of one and a half: the shift rises, flattens and stays. The half-normal’s tail is so thin that a large difference cannot be explained by a large τ — the prior forbids it — so the analysis explains it as a moderate τ plus bad luck and keeps borrowing. Under the half-Cauchy the shift rises to about 0.17, then falls: 0.131 at half a standard deviation and 0.029 at one and a half. A large difference is explained by a large τ, which the heavy tail permits, and a large τ borrows nothing.
This is the property robust Bayesian analysis calls conflict resolution, and it is a property of the prior’s tail rather than of its scale. A thin-tailed prior never lets go of an information source that contradicts the data; a heavy-tailed one lets go once the contradiction is large enough that the tail is the better explanation. The same distinction decided when borrowing goes wrong for a single group from outside its population — shrinkage towards a normal centre punished the outlier six times over — and it decides here whether a trial is analysed against its own control or against somebody else’s.
The error the borrowing does
The frequentist cost follows from the shift. When the earlier control mean sits below the current one — the earlier population was sicker, the standard of care has improved, measurement drifted — borrowing drags the current control’s estimate down, the treatment looks better than it is, and a treatment with no effect is declared a success more often than 2.5%.
Without borrowing the rate is 2.5% at every drift. Pooling outright reaches 73.1% at a drift of half a standard deviation. Under the half-normal prior of scale 0.05 the rate rises to 35.2% near a drift of 0.7 standard deviations and is still 31.8% at a drift of 1.5, a disagreement so large that nobody looking at the two trials would call them comparable. Under the half-Cauchy of the same scale it peaks at 11.0% near a drift of 0.4 and falls back to 3.4% at 1.5, close to the nominal rate: the prior has let go.
So the half-Cauchy’s error has a peak and the half-normal’s has a plateau. The peak sits at the drift where a disagreement is too large to be harmless and too small to be recognised — near 0.4 standard deviations here, about two and a half times the standard error of the difference between the two control means — and every borrowing rule that lets go has such a peak somewhere, because letting go needs evidence and a moderate drift does not provide it.
The location of the peak can be read off the model. The observed difference between the two controls has a standard error of 0.158 when the trials agree. A drift below about one standard error is invisible in the difference and small in its consequences: borrowing moves the control’s estimate by a fraction of a small bias. A drift of four or more standard errors produces differences that a heavy-tailed prior attributes to a large spread, and the borrowing falls away. Between them lies a band of drifts large enough to bias the control’s estimate by a substantial fraction of the treatment effect and small enough to be read as chance, and the rules differ only in how wide that band is and how high the error inside it climbs.
Widening the scale shrinks the band. At a scale of 0.2 the half-normal’s worst rate falls to 7.0% and the half-Cauchy’s to 5.7%, and the two curves nearly coincide. The price is in the borrowing: the half-normal now borrows 31.4 patients at no drift and the half-Cauchy 27.8 — 30% and 42% of what the same families took at a scale of 0.05.
What the borrowing buys
The reason to borrow is power. At no drift and a true effect of half a standard deviation, the trial alone has 70.5% power. Borrowing under the half-normal prior of scale 0.05 raises it to 87.9%, and under the half-Cauchy of the same scale to 84.1%; the half-normal analysis borrows 106 of the earlier trial’s two hundred control patients on average and the half-Cauchy 66.5.
Laid out as power against worst-case error, the rules form two curves, and the half-Cauchy’s lies above the half-normal’s throughout: for the same power it pays less in its worst error, and for the same worst error it buys more power. At the small scales that buy most power the gap is large — the half-normal at scale 0.02 reaches a worst rate of 95.6% — and at a scale of 0.5 the two nearly coincide, because a prior that wide borrows little and its tail hardly matters.
And no rule reaches the corner. Every rule that raises power above the trial’s own 70.5% raises its worst false-positive rate above 2.5% somewhere. That is not a feature of these priors. It is a theorem about any analysis that borrows: at the drift where the earlier trial is biased by just enough to be indistinguishable from agreement, the borrowing that helps at zero drift has to hurt, because the data cannot tell the two situations apart and the rule has to treat them alike. The only questions a design gets to choose are how much power to buy and how the error it costs is distributed over drifts — as a plateau that never ends, or as a peak that does.
The confident prior, seen from the other side
A half-normal prior that never lets go is a special case of a problem when the prior is confident and wrong measured for a prior on a mean: a prior worth thirty-five observations and centred in the wrong place produced an interval that covered nothing and took seventeen thousand observations to repair. There the confidence was in a location. Here it is in the similarity of two trials, and the repair is not available at all, because the current trial cannot grow into a second history: however large it is, it contributes one difference to the estimate of τ.
That is why the tail matters more than the scale. With a prior on a mean, enough data eventually overwhelm any proper prior. With a prior on the spread between two trials, the data never become enough; the likelihood for τ stays as flat as one difference makes it, and the prior’s shape in the region the data point to is the whole of the answer. A half-Cauchy lets go because its shape there is permissive, not because the data grew.
Why the scale is a promise about the history
A protocol that borrows an earlier control has to state its prior on the spread, and “a half-normal with scale 0.05” is usually defended as a statement about how similar the trials are expected to be. The numbers say it is a statement about something else as well: how the analysis will behave if the expectation is wrong. A scale small enough to borrow heavily is a scale that, with a thin tail, commits the analysis to borrowing heavily however badly the trials disagree.
That is why the choice of family matters more than the choice of scale. What a prior is worth found that a conjugate prior for a proportion is worth a fixed number of observations whatever the data say. A prior on the spread between trials is worth a number of the earlier trial’s patients that depends on the data — 106 at no drift under the half-normal of scale 0.05 and 66.5 under the half-Cauchy, near nothing at a large drift under the half-Cauchy — and the family decides whether that number can fall.
What a design that borrows should report
The error curve over drift, not the error at no drift. Every borrowing rule holds its nominal level when the trials agree, roughly, and every one exceeds it somewhere. The peak of that curve, and where it sits, is the honest summary of what the borrowing costs, and it should sit beside the power it buys.
The family of the prior on the spread, and not only its scale. A half-normal and a half-Cauchy with the same scale agree on typical spreads and disagree completely about the case the design most needs to survive.
How much was borrowed in the trial that ran. A borrowing analysis reports its effect estimate; it should also report how many of the earlier trial’s patients that estimate used, which the interval that integrates makes available from the same posterior. A trial whose analysis borrowed 106 patients when the two controls disagreed by a third of a standard deviation is a different result from one that borrowed 106 when they agreed.
A prior that cannot be flat. With one earlier trial the flat prior on the spread has no posterior, and an analysis that appears to use one has used whatever its software’s upper bound on τ happened to be. That bound is then the borrowing rule, stated nowhere.
What is measured here and what is not
With two trials, a flat prior on the spread between them is improper, because the likelihood from one difference falls like 1/τ.
At a scale of 0.05, a half-normal prior’s analysis declares a null treatment a success 35.2% of the time at worst and 31.8% at a drift of 1.5 standard deviations; a half-Cauchy’s peaks at 11.0% and returns to 3.4%. The two buy 17.4 and 13.6 points of power at no drift.
No rule that buys power over the trial’s own 70.5% keeps its worst false-positive rate at 2.5%.
The rates are counted over three thousand simulated trials a point, fifty patients an arm and two hundred in the earlier control, a normal outcome with known unit variance and a one-sided posterior threshold of 97.5%. The posterior at each trial is exact on a grid of four hundred values of the spread.
Not measured: outcomes with estimated variance, binary outcomes, and robust mixture priors that put the earlier trial’s information and a vague component side by side rather than on a spread between them. Not claimed either that the drifts where the error peaks are the drifts real historical controls show; they are where a two-trial model is most vulnerable, whatever the history is.
Still open: more than one earlier trial
Everything here is the two-trial case, where the spread between trials is barely estimable and the prior decides. The obvious escape is more history: with four, eight or sixteen earlier trials of the same control, the spread between them is estimated from the data, and the choice of prior should stop mattering as it did with eight groups.
Whether more history makes the borrowing safer is a different question, and it is not obviously yes. Sixteen earlier trials that agree with each other estimate a small spread confidently, and a small spread confidently estimated is exactly the belief that makes an analysis borrow heavily from a history the current trial disagrees with. How the power, the worst-case error and the prior’s influence move as the number of earlier trials grows, has not been measured.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Eight groups, one population — both name hierarchical model, partial pooling, prior
- Pooling a proportion — both name hierarchical model, partial pooling, prior
- What the plug-in forgets — both name hierarchical model, partial pooling, prior
- A boundary for giving up — both name error rate, statistical power
- A group from the population's own tail — both name hierarchical model, partial pooling
- A league table of a hundred — both name hierarchical model, partial pooling
Named objects
A flat tag is an object no other essay names yet.
Error rateThe half-Cauchy priorHierarchical modelImproper posteriorPartial poolingPriorPrior sensitivityStatistical power