A history that has been moving
Worth reading first: What a prior is worth · The weight that decides.
A history that agrees with itself found that borrowing a control arm from sixteen earlier trials is more dangerous than borrowing from one, because sixteen trials that agree make a large spread between trials implausible and leave the analysis nothing to explain a drifting current trial with but chance. Its histories were all exchangeable: the earlier trials were interchangeable draws, and the current trial was one more. It ended on the way real histories of controls usually differ from that. Standards of care improve, populations change, assessments drift, so the earlier trials sit along a trend and the current trial, being the latest, sits at its end.
A trend is a specific kind of disagreement, and the exchangeable model’s response to it is specific too. It does not see a slope. It sees a spread, centred on the middle of the history, and it borrows towards the middle.
Two priors for the next trial
Put the earlier trials at times one to and the current trial at , and let the control mean have been rising by a fixed amount a trial, so that the current trial’s control sits at zero and the earlier ones below it. That is the direction in which borrowing flatters a treatment: a history lower than the current control pulls the control’s estimate down and the treatment’s advantage up.
The exchangeable analysis is the one the previous essay used. Every trial’s control mean is a draw from ; the earlier trials estimate , and that estimate plus , integrated over the posterior for , is the meta-analytic predictive prior for the current control. The alternative adds one term. Every trial’s control mean is with ; the earlier trials estimate the line, and the prior for the current control is the line’s prediction at plus and the line’s own uncertainty there. With flat priors on and the history’s evidence about is its likelihood with the line integrated out, and with no trend term the construction is the exchangeable one exactly.
The picture shows the whole difference. From eight earlier trials drifting 0.05 standard deviations a trial, the exchangeable prior for the current control has mean −0.213 and standard deviation 0.138; the line’s has mean 0.034 and standard deviation 0.071. The truth is zero. The exchangeable prior is centred a quarter of a standard deviation low, and its width, though generous, is not generous enough to stop a current control at zero being pulled towards it.
A slope is read as a spread
The exchangeable model does not ignore the trend. It absorbs it into , the only place a systematic difference between trials can go, and the amount it absorbs is predictable.
Trials at times one to spread with standard deviation , so a line of slope puts the trials’ true means that far apart times . The history’s posterior for follows that reference once the slope is large enough to be distinguished from sampling error: at a drift of 0.1 a trial, sixteen earlier trials give a posterior median of 0.472 against the reference 0.461, and eight give 0.231 against 0.229. Below that the half-Cauchy’s scale of 0.05 dominates, and a gentle trend is read as trials that agree.
Neither reading places the current trial correctly. A history read as agreeing borrows hard towards its middle. A history read as widely spread borrows less, but it still predicts the current trial from the middle, and the current trial is not a draw from around the middle: under a line it is beyond every earlier trial, times the spread the trend puts between the earlier trials above the history’s mean — 2.24 times at four trials, 1.96 at eight, 1.84 at sixteen, and never less than . So the exchangeable model is wrong in location when it borrows much and right only in the sense of borrowing little.
The trend that does the damage is the gentle one
That shape produces the curves at the top, and their most important feature is where they peak. From sixteen earlier trials the exchangeable model’s false-positive rate is 2.7% with no trend, 6.5% at a drift of 0.01 a trial, 9.9% at 0.02, 9.1% at 0.03, and then falls — 7.5% at 0.05 and 5.0% at 0.1. From eight trials the peak is later and similar in height, 9.9% at 0.05; from four, the rate is still rising at 9.8% at 0.1.
The peak sits where the trend is large enough to put the current trial well above the history’s middle and small enough to be read as agreement. With sixteen trials that is a drift of 0.02 a trial: a total movement of 0.32 standard deviations over the history, and a step between neighbouring trials of 0.02 against each trial’s own standard error of 0.071. No comparison of adjacent trials could see a step of that size, and no test for a spread between trials would find one — the history’s posterior median for the spread at that drift is 0.090, well inside what a half-Cauchy of scale 0.05 considers ordinary. The trend is invisible trial by trial and decisive in total.
Steeper trends are safer for the exchangeable model only in the sense that they are visible. At a drift of 0.1 a trial the sixteen-trial history is read as spread by nearly half a standard deviation, the borrowing falls from 760.1 patients with no trend to 2.8, and the false-positive rate returns towards the nominal. The analysis has given up borrowing without learning where the current trial is.
What the bias is
The false-positive rate counts decisions. The bias counts how far the estimate is off, which is what a trial that reports an effect size passes on.
From eight earlier trials the exchangeable model overstates the effect by 0.046 standard deviations at a drift of 0.01, 0.082 at 0.02 and 0.130 at 0.05 — a quarter of the half-standard-deviation effect the trial was designed to detect. From sixteen, 0.116 at a drift of 0.02 and 0.121 at 0.03. From four, the bias is still growing at 0.122 at a drift of 0.1. The vague fifth of the prior, which brought the sixteen-trial worst error rate from 80.7% to 16.9% against the previous essay’s drifting current trial, helps less here: at eight trials and a drift of 0.05 it brings the bias from 0.130 to 0.104 and the false-positive rate from 9.9% to 7.6%. The vague component is a way out for a current trial that disagrees sharply with the history, and a trend’s current trial disagrees only mildly — enough to bias every estimate, not enough to trigger the escape.
A trend in the other direction
Everything above has the control mean rising, so that a history below the current trial flatters the treatment. A control mean that has been falling — a population getting sicker, an outcome scale drifting the other way — puts the history above the current trial, and the same pull towards the middle now works against the treatment.
The error changes sign and the damage moves from the false positives to the power. With a true effect of half a standard deviation and the control mean falling 0.02 a trial, the exchangeable model borrowing from sixteen earlier trials finds the effect 68.1% of the time, against 94.3% with no trend and the trial’s own 70.5% without borrowing, and it understates the effect by 0.113 standard deviations. At a fall of 0.05 a trial from eight earlier trials, power is 58.5% and the understatement 0.129. The borrowing has made the trial worse than not borrowing at all, while reporting the precision of a control arm several times the size of its own. The line’s power is 93.0% from sixteen trials and 89.9% from eight at either drift, since it does not see the trend in either direction.
So the bias of an exchangeable history is not a regulator’s worry alone. It flatters a treatment when controls have been improving and hides one when they have been worsening, and how many subjects sized a trial for 80% power on a calculation that assumed nothing of the kind. A prior that is confident and centred in the wrong place is the failure when the prior is confident and wrong measured for a single prior worth thirty-five observations; a trending history supplies exactly such a prior, built from data, and reports its confidence as the history’s agreement.
A line removes it exactly
The model with a line in time has the flat rows in every figure, and they are flat for an exact reason. Adding to every earlier trial’s control mean moves the fitted line by exactly , and its prediction at by exactly , with nothing else changed — neither the residuals, nor the history’s evidence about , nor the prior’s width. Every quantity the analysis computes is the same on the same draws whatever the slope, so the false-positive rate is 2.9% at every drift from eight earlier trials, 2.7% from sixteen and 1.5% from four, and the bias is 0.007, −0.001 and −0.006 at every drift. The figures do not show the line doing well against the trend; they show it not seeing the trend at all.
That is the right behaviour for a straight-line trend and it is a claim about straight lines only. A history whose control mean rose and then levelled, or stepped once when the standard of care changed, is not a line, and the line will extrapolate a slope into a current trial that has none. Borrowing towards a line found the same property in a league table: shrinking groups towards a fitted relation rather than a common centre removes the bias the relation causes, exactly when the relation is the one fitted, and the relation the table has to estimate found what it costs when the relation’s shape has to be chosen from the table.
What the line costs
The line is paid for twice, in the same currency. It estimates a slope from the history, which leaves fewer degrees of freedom for the spread; and it predicts at the end of the history, where a fitted line is least certain, rather than at its middle, where a fitted mean is most certain.
With no trend at all, the exchangeable model borrows 238.1 control patients from four earlier trials, 438.8 from eight and 760.1 from sixteen. The line borrows 69.7, 174.9 and 370.4. From three earlier trials, the fewest a line and a spread can be estimated from, it borrows 44.1 against the exchangeable 183.2. At every length the line borrows under half of what the exchangeable model does, and the ratio improves only slowly with the length of the history, from about a quarter at three trials to about a half at sixteen.
Power follows. With no trend and a true effect of half a standard deviation, the trial alone has 70.5%; the exchangeable model borrowing from sixteen trials, 94.3%; the line, 93.0%. From eight, 92.9% against 89.9%; from four, 90.3% against 84.5%. Most of what borrowing buys in power survives the line when the history is long, because by sixteen trials both models have borrowed enough that more makes little difference. The borrowed patients are the more honest measure of the cost: sixteen earlier trials are worth 760 patients to an analysis that assumes the control has not moved and 370 to one that allows it to have moved in a straight line.
How many trials a trend needs
The previous essay’s question was how many earlier trials make the history’s agreement safe to borrow on, and the answer was none: agreement only tightens the grip. The question here has a more useful answer. Three earlier trials are the minimum from which a line and a spread about it can be estimated at all, and at three the line borrows 44 patients — the equivalent of one more trial of fifty in the current study, and the power it buys is 81.6% against the trial’s own 70.5%. By eight trials the line borrows more than three current control arms’ worth.
So a line is affordable from a history of about six earlier trials onwards, where it borrows 123.0 patients, and it is protective at any length. Against it, the exchangeable model borrows two to four times as much and is protected against nothing that moves steadily. The choice between them is a choice about which risk a trial can carry, and a history of controls is exactly where the steady movement is expected.
Why a history is not a sample
The exchangeable model’s assumption is that the order of the earlier trials carries no information. A control borrowed from the last trial had one earlier trial and no order to ignore, and a prior on the spread was about eight groups observed at the same time. A history of trials is different in kind: it was produced in order, and the current trial comes after all of it. An analysis that treats its members as interchangeable throws away the one covariate every history has.
One population or two found that a hierarchical model reports its estimated spread and weights every group by it without flagging the group the spread does not describe. A trending history is the sharpest case of that: the spread the model estimates is real, the trials do differ, and the group it does not describe is precisely the one being analysed — the last, at the end of a line the model has turned into a bell. The weight that decides established that the posterior mean’s weight is exactly; with a trend, the weight is computed correctly and applied to the wrong centre.
What a borrowing analysis should state about its history
The order of the earlier trials and the time each was run. Without them nobody can check for a trend, and with them the check is a line.
The fitted drift per trial and its uncertainty. A drift of 0.02 a trial is invisible in any single comparison and decisive over sixteen trials; a report that gives only the history’s spread cannot distinguish it from agreement.
Both analyses, exchangeable and with a line. Where they agree, the history was not moving detectably and the exchangeable model’s extra borrowing is available. Where they disagree, the difference is an estimate of the bias the exchangeable model is carrying.
The patients borrowed under each. Sixteen earlier trials worth 760 patients under one assumption and 370 under a weaker one is a statement about how much of the borrowing rests on the control never having moved.
What is exact and what is counted
Exact: a model with a line in time gives identical answers, on the same draws, whatever the slope of a linear trend; with no trend term it is the exchangeable meta-analytic predictive analysis to the last digit.
Counted, over 1,500 trials a point with fifty patients an arm and earlier controls of two hundred: the exchangeable model’s false-positive rate of 9.9% at a drift of 0.02 a trial from sixteen earlier trials and at 0.05 from eight; its bias of 0.130 at 0.05 from eight; the line’s 2.9% and 2.7% at every drift; and the borrowed patients, 438.8 against 174.9 from eight trials and 760.1 against 370.4 from sixteen.
Not claimed: that control arms drift in straight lines. The line is the simplest model with an order in it and the one whose repair is exact, so it measures the size of the problem cleanly; a history that stepped once is the case it handles worst. Not claimed either that 0.05 is the right scale for the spread about the line; it is the scale the earlier borrowing essays used throughout, and a larger one would borrow less under both models.
Still open: a history that stepped
The line is exact against a steady drift and wrong against a single change: a history whose first eight trials sit at one level and whose last eight sit at another, because a new standard of care arrived between them. Fitted with a line, that history is a slope, and the line extrapolates it past the current trial — overcorrecting in the direction that makes a treatment look worse. Fitted as exchangeable, it is a spread, and the current trial sits at the upper level while the prior sits between the two.
The model that matches it has a change point in the history and a prior on where the change happened, and it would borrow only from the trials after it, which may be few. How much a change-point model borrows, how badly the line and the exchangeable model do against a step of a plausible size, and whether a history of sixteen trials can tell a step from a slope with enough confidence to choose between them, are measurements of the same kind as these and have not been made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- What a two-unit study should report — both name hierarchical model, partial pooling, prior sensitivity, variance components
- Eight groups, one population — both name hierarchical model, partial pooling, prior
- Pooling a proportion — both name hierarchical model, partial pooling, prior
- The fewest groups that can borrow — both name hierarchical model, partial pooling, variance components
- The prior the data estimates — both name partial pooling, prior, variance components
- What the plug-in forgets — both name hierarchical model, partial pooling, prior
Named objects
A flat tag is an object no other essay names yet.
Error rateThe half-Cauchy priorHierarchical modelPartial poolingPriorPrior sensitivityStatistical powerVariance components