What a sample-size calculation was given

The correlation a pilot can see

Two models of repeated visits agree that adjacent visits correlate 0.7 and disagree about how many follow-ups a trial should buy: three under equal correlation, one when the correlation fades. Planning on the wrong one costs 28.0% or 9.9%. Two follow-ups cost at most 3.2% more than the best under either, and at most 11.1% under any curve with that adjacent correlation. A pilot of twenty patients at four visits chooses worse on average than that rule when the correlation fades; it takes about forty to draw level with it, and at an effect of half a standard deviation those forty patients' extra visits cost fifty times what the better schedule saves.

Worth reading first: What a p-value does not say.

A visit before or a visit after worked out where a trial should spend repeated measurements of its outcome, and ended on two models that agreed about adjacent visits and disagreed about everything a schedule depends on. Under compound symmetry every pair of a patient’s visits correlates 0.7, and with a patient costing five visits’ worth to recruit the cheapest design takes one baseline and three follow-ups. Under a model in which a stable patient level carries 0.4 of the variance and the rest fades by half from one visit to the next, adjacent visits still correlate 0.7, and the cheapest design takes one follow-up. A plan has to choose between them, and it usually has a pilot to choose with.

The question left open was how large that pilot must be to tell the two apart, and what a plan built on the wrong one costs. Both have exact answers, and together they say something the lead did not expect: the schedule matters less than the model does, because there is a schedule that is nearly right under both, and a pilot large enough to improve on it costs more than it saves unless the main trial is very large.

The extra cost of each follow-up schedule over the cheapest, across every correlation curve with adjacent visits at 0.7, when a patient costs 5 visits to recruitWith one baseline, 1 follow-up costs at most 28.0% more than the cheapest schedule, 2 follow-ups cost at most 11.1% more than the cheapest schedule, 3 follow-ups cost at most 24.1% more than the cheapest schedule, 4 follow-ups cost at most 34.8% more than the cheapest schedule, over every stable share from 0 to 0.7. The schedule whose worst case is smallest is 2 follow-ups; its worst case is at a stable share of 0.00. Under compound symmetry the cheapest is 3; under a stable share of 0.4 it is 1.0%10%20%30%00.2000.4000.600stable share of the variance — the correlation between distant visitsextra cost over the cheapest schedulethe fading modelequal correlationone baseline, 1 follow-upone baseline, 2 follow-upsone baseline, 3 follow-upsone baseline, 4 follow-upsa patient costs 5 visits to recruit; exactlines leaving the top are cut at 39%
Fig. 1 The extra cost of reaching the same power with one baseline and one to four follow-up visits, over the cheapest schedule, across every correlation curve whose adjacent visits correlate 0.7 — indexed by the stable share of the variance, from a curve with no stable level at the left to compound symmetry at the right. The vertical lines mark the two models of the earlier essay. The slider sets how many visits a patient costs to recruit.

One family of curves, indexed by what does not fade

The two models are two members of one family. Write the correlation between two visits kk apart as

ρk=s+(1−s) d k,\rho_k = s + (1 - s)\,d^{\,k},

where ss is the stable share — the part of the variance that is the patient’s own level and never fades — and dd is how fast the rest decays from one visit to the next. Fixing the adjacent correlation at ρ1=0.7\rho_1 = 0.7 leaves one free number, and it is ss: compound symmetry is s=0.7s = 0.7, where nothing fades; the fading model is s=0.4s = 0.4 with d=0.5d = 0.5; and s=0s = 0 is a curve with no stable level at all, whose correlations fall towards zero as the visits separate. The stable share is also what the correlation between distant visits tends to, which is why it is the quantity a pilot has to see: adjacent visits say nothing about it, because every member of the family has the same adjacent correlation.

The cost of a schedule at fixed power is the adjusted analysis’s variance, which sets the number of patients, times what each patient costs: V(1,r) (K+1+r)V(1, r)\,(K + 1 + r) for one baseline and rr follow-ups with a patient costing KK visits to recruit. The variances are the general covariance sums of the earlier essay, adjusting for the baseline, and every cost below is exact.

What each model charges for the other

Planning on the wrong model has a price, and the price is not symmetric.

If the truth is compound symmetry and the plan assumed fading, the plan buys one follow-up where three were cheapest, and the trial costs 28.0% more than it had to — 33 patients an arm where 20 would have done, at an effect of half a standard deviation and 80% power, each measured twice rather than four times. If the truth is the fading model and the plan assumed compound symmetry, the plan buys three follow-ups where one was cheapest and costs 9.9% more.

The asymmetry has a reason. Under compound symmetry every extra follow-up is as informative about the patient’s level as the first, so a plan that stops at one leaves a large and cheap reduction in variance unbought; the variance with one follow-up is 0.510 of the unadjusted analysis’s and with three 0.310. Under fading, the extra follow-ups are further from the baseline and share less with each other, and the variance only falls from 0.510 to 0.436 — so over-buying them wastes visits without losing much else. Under-buying precision where it was cheap is dearer than over-buying it where it was not.

A schedule that does not need the curve

The hero figure shows the other thing the two costs hide. Near its minimum the cost of a schedule is flat: under compound symmetry two follow-ups cost 3.2% more than three, and under the fading model two follow-ups cost 2.9% more than one. A plan that takes two follow-ups without knowing which model is true is within about three per cent of the best under both.

That is the maximin idea protecting one parameter over a range applied to a design’s unknown: choose the schedule whose worst case over the plausible truths is smallest. Over the whole family — every stable share from 0 to 0.7 — the worst cases are 28.0% for one follow-up, 11.1% for two, 24.1% for three and 34.8% for four. Two follow-ups are the minimax schedule, and their worst case is at s=0s = 0, the curve with no stable level at all, which is further from compound symmetry than the fading model is. Against the two named models alone, two follow-ups never cost more than 3.2%.

The minimax schedule depends on the recruitment cost and the adjacent correlation, and the guarantee it gives varies a good deal:

recruitment cost adjacent 0.5 adjacent 0.7 adjacent 0.9
2 visits 2 follow-ups, at most 1.6% 1 follow-up, at most 13.3% 1 follow-up, at most 8.6%
5 visits 6 follow-ups, at most 6.7% 2 follow-ups, at most 11.1% 1 follow-up, at most 19.8%
10 visits 7 follow-ups, at most 4.4% 2 follow-ups, at most 9.5% 2 follow-ups, at most 24.9%

At a low adjacent correlation the curve barely matters, because the adjusted analysis gains little from the baseline whatever the model and the question is mostly how many noisy follow-ups to average; the minimax guarantee is a few per cent. At a high adjacent correlation it matters a great deal: if nearly all of the correlation is stable, extra follow-ups are cheap precision, and if none of it is, they are nearly worthless, and no single schedule is within twenty per cent of both. The slider on the hero figure shows the recruitment cost’s share of it. When patients are cheap relative to visits, one follow-up is the only sensible buy and the curve matters only near compound symmetry; when patients are dear, the family’s members disagree more and the safe choice gives up more.

What a pilot sees of the curve

A pilot that measures its patients at several visits with no treatment estimates the correlation at every lag, and a curve fitted to those correlations estimates the stable share. The difference it is looking for is at the longest lags: at four visits, three apart, the fading model’s correlation is 0.475 and compound symmetry’s is 0.7.

The stable share a pilot of 20 patients at 4 visits fits, under a fading correlation and under equal correlation. Over four hundred pilots each, the fitted stable share has a median of 0.40 when the truth is 0.4 and 0.68 when it is 0.7; the two histograms share 37% of their mass. A pilot drawn from the fading model is fitted with a share above 0.6 14% of the time.
Fig. 2 The stable share fitted to each of four hundred pilots of twenty patients measured at four visits, when the truth is the fading model and when it is compound symmetry. Each fit minimises the squared error of the family’s curve against the pilot’s correlations at each lag, weighted by the number of pairs of visits at that lag. The dashed lines are the true stable shares.

The fitted shares are centred in the right places — medians of 0.40 and 0.68 — and they are wide. A correlation estimated from twenty patients has a standard error of about 0.11 near 0.7 and 0.14 near 0.6, the longest lag has only one pair of visits to average, and the fit has to separate a stable share from a decay with three noisy numbers. So 17% of pilots drawn from the fading model fit a stable share of exactly zero, reading the curve as one with no stable level, and 13.5% fit one above 0.6, reading it as nearly compound symmetry.

That is the same kind of failure the spread a pilot supplies found for a single standard deviation estimated from ten patients, and harder, because the pilot is estimating a curve from its far end, where it has the fewest pairs.

The schedule a pilot chooses

A plan that takes the pilot’s fitted curve at its word chooses the schedule that curve makes cheapest, and pays for it under the true curve.

The average extra cost of the follow-up schedule a pilot chooses, when the true stable share is 0.4, against the size of the pilot. The truth's cheapest schedule is 1 follow-up. Choosing from a pilot of ten patients costs on average 9.1% more at 3 visits, 8.0% more at 4 visits, 6.8% more at 6 visits; from a pilot of 160, 1.10%, 0.75%, 0.57%. The minimax schedule, 2 follow-ups chosen with no pilot, costs 2.9% more.
Fig. 3 The average extra cost, over the cheapest schedule, of the follow-up schedule chosen from a pilot’s fitted curve, when the truth is the fading model, against the number of patients in the pilot, for pilots measuring each patient three, four and six times. The dashed line is two follow-ups, chosen with no pilot.

When the truth fades, a pilot of twenty patients at four visits chooses the right schedule, one follow-up, 55.8% of the time, and two follow-ups 23.5%, three 13.5%, and eight — the fit having read the curve as compound symmetry with a high stable share — 3.0%. Its choice costs on average 4.14% more than the best, against the 2.9% that two follow-ups cost with no pilot at all. Ten patients cost 7.95%. The pilot draws level with the rule that ignores it only at about forty patients at four visits, 2.56% against 2.9% — within the simulation’s error of each other — and pulls clear at eighty, 1.08%. More visits per pilot patient help, because they add lags and pairs: at twenty patients, six visits cost 3.22% where four cost 4.14%.

The average extra cost of the follow-up schedule a pilot chooses, when the true stable share is 0.7, against the size of the pilot. The truth's cheapest schedule is 3 follow-ups. Choosing from a pilot of ten patients costs on average 8.2% more at 3 visits, 5.2% more at 4 visits, 2.3% more at 6 visits; from a pilot of 160, 0.18%, 0.03%, 0.00%. The minimax schedule, 2 follow-ups chosen with no pilot, costs 3.2% more.
Fig. 4 The same when the truth is compound symmetry, whose cheapest schedule is three follow-ups. The dashed line is two follow-ups with no pilot, 3.2% over the cheapest.

When the truth is compound symmetry the pilot does better, because a curve that does not fade is easier to recognise — the far lags look like the near ones, and a pilot rarely manufactures a decay out of noise as readily as it hides one. Twenty patients at four visits choose three follow-ups 71.8% of the time and cost 2.87% on average, level with the rule’s 3.2%; forty cost 0.63%. Across both truths the pattern is the same: below forty patients a pilot is little or no better a guide to the schedule than ignoring the curve, and whether it is better above that depends on which curve is true, which is what the pilot was run to find out.

What the pilot’s own visits cost

None of the costs above charges the pilot for itself. A pilot of forty patients at four visits spends 160 visits and forty recruitments; if it is being run anyway, for the spread or for feasibility, and would have measured each patient twice, the correlation curve costs it eighty extra visits.

What a pilot's extra visits save in the main trial's schedule, net of their own cost, against the size of the trial. Pilots at four visits a patient, against taking the minimax schedule with no pilot, when the true stable share is 0.4. At an effect of half a standard deviation the net is -78 visits for 40 patients, -151 visits for 80 patients, -310 visits for 160 patients. At an effect of 0.05 it is 90, 675, 663.
Fig. 5 Visits saved in the main trial by choosing the schedule from a pilot rather than taking two follow-ups, less the pilot’s extra visits, against the effect size the main trial is powered for, for pilots of forty, eighty and a hundred and sixty patients at four visits, when the truth is the fading model. The vertical scale is a signed logarithm; above zero the pilot’s visits repay themselves.

At an effect of half a standard deviation the main trial needs 33 patients an arm and costs 462 visit-units; forty pilot patients’ better choice saves 1.7 of them, against 80 extra visits spent — about fifty times more spent than saved. Eighty pilot patients save 8.6 and spend 160. The saving scales with the size of the main trial and the pilot’s cost does not, so the balance turns, and it turns late: forty pilot patients repay their extra visits only at an effect of about 0.07 standard deviations, a trial of 1,634 an arm; eighty patients at about 0.1, a trial of 801 an arm, where they save 208.8 visits and spend 160. Under compound symmetry the balance turns sooner, because the pilot’s gain there is larger — at an effect of 0.1 forty pilot patients come out 147 visits ahead — but not by an order of magnitude.

Trials of a few hundred patients or fewer, which is most trials that measure a noisy outcome repeatedly, should not run a pilot to choose their schedule. They should take the minimax schedule and spend nothing finding out what the curve is.

The number an earlier trial already reported

A pilot is not the only source of the curve, and it is the worst-placed one. The quantity that separates the models is the correlation between distant visits, and there is a correlation of exactly that kind that trials routinely report: the one between their baseline and their final visit, which is what a baseline cut in two and every covariance-adjusted analysis lean on, and what a table of baseline characteristics and outcomes lets a reader reconstruct.

Four visits apart, the fading model puts that correlation at 0.4375 and compound symmetry at 0.7. An earlier trial of two hundred patients estimates a correlation near 0.44 with a standard error of about 0.057, so the two models’ values sit 4.6 standard errors apart: one reported number from one ordinary earlier trial tells them apart decisively, where four hundred pilots of twenty patients each overlapped heavily. The earlier trial did not set out to measure the curve. It measured the curve’s far end on many patients because that is what its own analysis needed, which is precisely the pairing of lag and sample size a small pilot cannot have.

This also says where the adjacent correlation, the number planners most often carry forward, came from. It is usually a test–retest reliability study, a few days apart, which is the one lag every member of the family agrees on. How many subjects sized a trial from one number; this plan needs two, and the second is the one the reliability literature does not report and the trial literature does. And a baseline taken on the day of enrolment carries the selection that got the patient in, which lowers its correlation with later visits for a reason no curve in this family describes — so an earlier trial’s baseline-to-end correlation should be read from a baseline taken after screening, where the trial reports one.

Why the arithmetic comes out this way

Three features of the problem combine, and each is general enough to carry beyond these numbers.

The cost is flat at its minimum. The cost of a schedule is a variance that falls with each follow-up times a per-patient cost that rises with it, and at the minimum the two changes cancel. A schedule one step away from the best costs a few per cent more, not tens, and a planner who is one step wrong has lost little. That is what makes a minimax schedule possible at all.

The pilot estimates the quantity at its noisiest end. The stable share is the correlation between distant visits, and a pilot with a handful of visits has one or two pairs at its longest lag; the adjacent correlation, which it estimates well, is exactly the number every member of the family shares. Allocating on a guess found the same shape for allocation rules fed a pilot’s estimates of the quantities the experiment was run to learn: a rule that is optimal at the truth is worse than a fixed rule when fed a noisy estimate of the truth, until the estimate is good.

And the pilot pays per patient while the saving is proportional to the trial. A pilot’s cost is fixed by its size; its value is a fraction of the main trial’s cost. For a small main trial that fraction of a small number cannot pay for a fixed outlay, which is the same arithmetic that makes the chance a trial succeeds worth more attention than the precision of any one planning input.

The plan these numbers support

Choose the follow-up schedule by its worst case across the curves the adjacent correlation allows, not by one model. At an adjacent correlation of 0.7 and a patient costing five visits, two follow-ups are within 11.1% of the best under any curve and within 3.2% under either named model; the table above gives the schedule and its guarantee at other values.

Do not run a pilot to learn the correlation curve for a trial of a few hundred patients. Below forty pilot patients at four visits the pilot chooses no better on average than the minimax schedule when the correlation fades; above it, its extra visits cost more than they save unless the main trial is powered for an effect of about a tenth of a standard deviation or less.

If a large trial does run one, measure each pilot patient more often. At twenty patients, six visits a patient cut the average extra cost from 4.14% to 3.22% under the fading curve and from 2.87% to 0.83% under equal correlation, because the far lags are where the curves disagree and more visits are the only way to get pairs there.

And borrow before measuring. Earlier trials of the same outcome, or cohort studies that measured it repeatedly over years, estimate the stable share from far more pairs at long lags than any pilot will, and their correlation matrices — where they are reported — are the cheapest pilot there is.

The costs are exact sums over the family’s correlation matrices, with the adjusted analysis regressing the follow-up mean on the single baseline; the minimax search runs over seventy-one stable shares from 0 to the adjacent correlation. The pilots are four hundred simulated pilots a point, from the true curve, each fitted on a grid of hundredths in both the stable share and the decay, and the averages carry the simulation error of four hundred draws — 0.42 percentage points on the 4.14% of a twenty-patient pilot under the fading curve, and 0.30 on the 2.56% of a forty-patient one. Not measured: a curve outside the family, such as one with a learning effect at the first visit or a drift in calendar time; a schedule with more than one baseline, which the earlier essay found rarely worth buying; and a Bayesian plan that averages the cost over a prior on the stable share rather than taking its worst case, which would sit between the minimax rule and the pilot and is the natural next comparison.

Still open: a schedule the trial chooses for itself

Every pilot here is separate from the main trial. A trial can instead take its first patients at several visits, fit the curve to them, and set the follow-up schedule for the rest — an internal pilot, whose patients are counted in the analysis rather than paid for and discarded. That removes most of the pilot’s cost, which is what sank the separate pilot above.

It adds two questions this essay’s arithmetic cannot answer. The schedule for later patients would depend on data from earlier ones, and the treatment estimate would then average over patients measured differently; whether that biases it, or only changes its variance, depends on whether the correlation fitted is correlated with the effect estimated, which for a correlation estimated without looking at the arms it should not be. And the internal pilot’s first patients carry the schedule chosen before anything was known, so the schedule’s cost is a mixture of the minimax rule’s and the fitted rule’s. How large the internal pilot should be, and whether it beats the minimax schedule for trials of the size that could never afford a separate pilot, is the measurement that would settle whether the curve is ever worth learning for an ordinary trial.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Analysis of covarianceAutocorrelationExperimental designMaximin designPilot studySample sizeStatistical powerTest–retest reliabilityWorst case