What a sample-size calculation was given

A schedule the trial chooses for itself

A separate pilot cannot pay for itself in choosing a trial's follow-up schedule, because its patients are recruited, measured and discarded. A trial can instead measure its own first patients on the schedule that needs no pilot — two follow-ups — fit the correlation curve to them, and give everyone after them the schedule that curve makes cheapest. Those first patients are counted in the analysis and cost no extra visit. Forty of them cut the trial's worst-case excess cost across every curve with adjacent correlation 0.7 from 11.1% to 5.9% in a trial of two hundred an arm, and eighty cut it to 3.6%; twenty buy almost nothing. And because the curve is read from within-arm deviations, the choice cannot see the effect: the pilot's effect estimate is the same whichever schedule it chose.

Worth reading first: What a p-value does not say.

The correlation a pilot can see asked how many follow-up visits a trial measuring a noisy outcome repeatedly should schedule, when the answer depends on how the correlation between visits fades with their separation — a curve no planner knows. Two follow-ups for everyone, chosen with no knowledge of the curve, cost at most 3.2% more than the cheapest schedule under either of the two models it compared, and at most 11.1% under any curve whose adjacent visits correlate 0.7. A separate pilot could see the curve, but it needed about forty patients to draw level with that rule, and its patients’ visits repaid themselves only for main trials powered at effects near 0.1 or less. The pilot was too expensive for the information it bought.

The essay ended on the way round that. A trial can take its first patients at several visits, fit the curve to them, and set the schedule for everyone after — an internal pilot, whose patients are counted in the analysis rather than paid for and discarded. That removes the pilot’s recruitment cost. It raises two questions the separate pilot did not: what the mixture of schedules costs, and whether a schedule chosen from the trial’s own patients can bias the estimate those same patients contribute to.

What a trial's own first patients buy when they choose the follow-up schedule for the rest, powered for an effect of 0.2Internal pilots measured on two follow-ups, the schedule that needs no pilot, whose correlation curve chooses the schedule for every later patient. Extra cost over the cheapest schedule at stable shares 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7: no pilot, two follow-ups throughout, 11.1%, 9.8%, 8.1%, 5.9%, 2.9%, 0.0%, 0.0%, 3.2%; an internal pilot of 20, 10.5%, 9.8%, 9.3%, 7.9%, 5.9%, 3.7%, 4.2%, 5.2%; an internal pilot of 40, 5.9%, 5.7%, 5.6%, 4.7%, 3.4%, 2.0%, 2.4%, 2.5%; an internal pilot of 80, 3.6%, 3.4%, 3.2%, 2.8%, 1.9%, 1.0%, 1.6%, 1.6%. Worst cases: no pilot 11.1%, 20 patients 10.5%, 40 patients 5.9%, 80 patients 3.6%.0%3%6%9%12%00.1000.2000.3000.4000.5000.6000.700the stable share of the variance, s (adjacent visits correlate 0.7)extra cost of the whole trial over the cheapest scheduleno pilot: two follow-ups throughoutinternal pilot of 20 patientsinternal pilot of 40 patientsinternal pilot of 80 patients300 pilots a point; exact variancespowered for 0.2; the trial counts the pilot's patients
Fig. 1 The extra cost of the whole trial over the cheapest schedule, across every correlation curve whose adjacent visits correlate 0.7 — indexed by the stable share of the variance — for internal pilots of twenty, forty and eighty patients measured on two follow-ups, beside the rule that gives everyone two follow-ups with no pilot. A trial powered for an effect of 0.2. The slider sets the effect.

A pilot that costs nothing to recruit

The family of curves is the earlier essay’s. The correlation between two visits kk apart is ρk=s+(1−s) d k\rho_k = s + (1-s)\,d^{\,k}, with the adjacent correlation fixed at 0.7, so the stable share ss indexes the family: s=0s = 0 is a curve that fades to nothing, s=0.4s = 0.4 the fading model, s=0.7s = 0.7 compound symmetry. Under each curve the cheapest schedule — one baseline and rr follow-ups, with a patient costing five visits’ worth to recruit — is known exactly, and every cost below is measured against it.

The internal pilot measures its first mm patients at a baseline and two follow-ups. That is the schedule the minimax rule would have given them, so the pilot spends no visit it would not have spent anyway; it only looks at what the visits say. The within-arm correlations between their three visits, at lags one and two, are fitted by the same curve fit the separate pilot used, and every later patient gets the schedule the fitted curve makes cheapest.

Patients on different schedules are combined by their precision. A patient on rr follow-ups carries the reciprocal of the adjusted analysis’s variance factor V(1,r)V(1, r) in information about the effect, the trial needs a fixed total precision to reach 80% power, and the patients after the pilot supply whatever the pilot’s own did not. The cost of the whole trial — recruitment plus visits — is then compared with the cost of the trial that knew the curve and gave everyone the cheapest schedule from the start.

What forty patients buy

The hero figure shows the trade at a trial powered for an effect of 0.2, which needs about two hundred patients an arm. With no pilot, two follow-ups cost 11.1% more than the cheapest schedule when the curve fades to nothing, 2.9% at the fading model, nothing at all at stable shares of 0.5 and 0.6, where two follow-ups are the cheapest, and 3.2% under compound symmetry.

An internal pilot of forty patients flattens that curve. It costs 5.9% at a stable share of zero, 3.4% at the fading model, 2.0% and 2.4% at 0.5 and 0.6, and 2.5% under compound symmetry. Its worst case across the family is 5.9% against the no-pilot rule’s 11.1%: it nearly halves the price of not knowing the curve. Eighty patients cut the worst case to 3.6%. Twenty patients barely move it, 10.5%, because a pilot of twenty fits the curve so poorly that its choices are as often wrong as right.

The pilot is not free in the middle of the family. Where two follow-ups are already the cheapest schedule, the minimax rule pays nothing and the pilot pays for its mistakes — 2.0% at a stable share of 0.5 with forty patients — because a fitted curve sometimes says one follow-up or three when two was right. The internal pilot buys protection at the ends of the family by spending a little in the middle, which is what a minimax choice exists to avoid and an adaptive one exists to improve on. Across the family the trade is strongly in the pilot’s favour; at any one curve near the middle it is not.

Against the separate pilot

The comparison with the separate pilot is the reason to run the internal one, and it is mostly arithmetic. A separate pilot of forty patients measured at four visits costs forty recruitments and a hundred and sixty visits — 360 visit units at five a recruitment — before the main trial begins. A main trial powered for 0.2 under the fading model, on its cheapest schedule, costs about 2,800 units, so the separate pilot is an eighth of the trial’s cost again, and the schedule it chooses has to save more than that to break even. It cannot, which is why the earlier essay found it worth running only for trials powered at effects of 0.1 or below.

The internal pilot of forty spends nothing that the minimax rule would not have spent: forty patients the trial needed anyway, measured at visits they would have had anyway. All of its value is in the choice it makes for the rest, and all of its cost is the mistakes in that choice and the pilot patients’ own schedule when it turns out wrong. That changes the question from “is the curve worth paying for?” to “is the curve worth looking at?”, and the answer to the second is yes from about forty patients upwards.

Choice for choice, the internal pilot is the slightly worse guide. At the fading model the separate pilot of forty patients at four visits chose a schedule costing 2.56% more than the cheapest; the trial steered by an internal pilot of forty costs 3.4% more. The difference has two sources, and both are visible in the design: the internal pilot sees two lags rather than three, so its fitted curve is noisier, and its own forty patients carry two follow-ups where one was cheapest whatever the fit says. Neither is large. The separate pilot’s better choice came at the price of an eighth of the trial; the internal pilot’s slightly worse one came free.

It also changes what the pilot can fail at. A separate pilot that misreads the curve has wasted its cost and steered the whole trial wrong. An internal pilot that misreads it has steered only the patients after it, and the patients in it are as informative as they would have been under the rule — the worst it can do is what the rule would have done for its own patients plus a wrong schedule for the rest.

How large the pilot has to be

The separate pilot needed forty patients to draw level with the rule and repaid its own visits only in very large trials. The internal pilot’s arithmetic is different, because its patients cost nothing extra and the only question is how well they choose.

How large an internal pilot has to be to improve on the schedule that needs no pilot. The worst extra cost over the cheapest schedule across every correlation curve with adjacent correlation 0.7. No pilot, two follow-ups throughout: 11.1%. Pilot patients on two follow-ups, powered for 0.2: 10.5%, 5.9%, 3.6%. Pilot patients on three follow-ups, powered for 0.2: 9.0%, 5.8%, 5.5%. Pilot patients on two follow-ups, powered for 0.1: 10.4%, 5.4%, 2.2%. Pilot patients on two follow-ups, powered for 0.3: 10.5%, 6.6%, 6.0% — at pilots of 20, 40, 80 patients.
Fig. 2 The worst extra cost over the family of curves against the number of internal pilot patients — for pilots on two follow-ups at three trial sizes, and on three follow-ups at an effect of 0.2 — with the no-pilot rule as the dashed line.

At every trial size, twenty pilot patients leave the worst case within a point of the no-pilot rule’s 11.1%: 10.5% for trials powered at 0.3 and 0.2, 10.4% at 0.1. Forty bring it to 6.6%, 5.9% and 5.4%. Eighty bring it to 6.0%, 3.6% and 2.2%. The pilot helps more in a larger trial, because the pilot’s own patients — who carry whatever schedule they were measured on, right or wrong — are a smaller share of it, and the patients after them carry the better choice. In a trial of about ninety an arm, powered for 0.3, eighty pilot patients are nearly half the trial and buy little more than forty did.

Measuring the pilot patients a third time after baseline, so that the curve can be fitted on three lags instead of two, is not free: those patients get a visit the minimax schedule would not have given them. At forty patients it buys almost nothing — 5.8% against 5.9% — and at eighty it costs more than it gains, 5.5% against 3.6%, because eighty patients on three follow-ups is a large block of the trial on a schedule that is wrong for most of the family. The extra lag helps the fit; it does not help enough to pay for a schedule the rest of the trial would not choose.

What the pilot chooses

The follow-up schedule an internal pilot of forty patients chooses for the rest of the trial, by the true curve. Stable share 0, cheapest 1 follow-up: 1, 75%; 2, 16%; 3, 5%; 4, 0%; five or more, 4%. Stable share 0.4, cheapest 1 follow-up: 1, 56%; 2, 24%; 3, 18%; 4, 0%; five or more, 2%. Stable share 0.6, cheapest 2 follow-ups: 1, 24%; 2, 29%; 3, 45%; 4, 1%; five or more, 1%. Stable share 0.7, cheapest 3 follow-ups: 1, 7%; 2, 15%; 3, 78%; 4, 0%; five or more, 0%.
Fig. 3 Which follow-up schedule an internal pilot of forty patients chooses for the rest of the trial, under four true curves, with the cheapest schedule for each at the right.

The choices show where the pilot’s protection comes from. When the curve fades to nothing, forty pilot patients choose one follow-up — the cheapest — in 75% of trials, and two in 16%; the rule that needs no pilot would have chosen two every time. Under compound symmetry the pilot chooses three follow-ups, the cheapest, 78% of the time. At the fading model it chooses correctly 56% of the time and errs towards two or three, which is the direction that costs least. At a stable share of 0.6, where two is cheapest, it chooses two only 29% of the time and three 45%: the curve there is hard to tell from compound symmetry with three visits, and the pilot cannot tell them apart, which is the middle-of-the-family cost the hero figure showed.

Occasionally the fit reads a fading curve as one that does not fade at all and chooses five or more follow-ups — 4% of the time at a stable share of zero. That is the separate pilot’s failure mode in the earlier essay, and it survives here because two lags are very little to fit a two-parameter curve to. It costs less here, because the pilot’s patients are already in the trial and the wrong schedule applies only to those after.

Why the choice cannot see the effect

The second question is the one that makes an internal pilot different in kind from a separate one: the schedule for later patients depends on data from earlier ones, and the earlier ones are part of the treatment comparison. Whether that biases the estimate depends on whether what the schedule reads is correlated with what the estimate reads.

For a normal outcome it is not, and the reason is a theorem rather than a measurement. The pilot’s correlations are computed from each patient’s deviations from their own arm’s visit means. Within an arm, the sample covariance matrix of normal observations is independent of the sample means, and the effect estimate is a function of the arms’ means. So the fitted curve, and the schedule chosen from it, is independent of the pilot’s effect estimate, and every later patient’s data are independent of both. The estimate combining them is unbiased, and its test holds its level conditionally on the schedule chosen, since the schedule only changes how precise each later patient is.

The internal pilot's own effect estimate, by the follow-up schedule it chose for the rest of the trial. Four thousand internal pilots of forty patients, measured on two follow-ups, with a true effect of 0.5. The pilot's effect estimate averages 0.503 when it chose 1 follow-up (56.0% of pilots), 0.514 when it chose 2 follow-ups (23.6% of pilots), 0.503 when it chose 3 follow-ups (17.1% of pilots), 0.505 when it chose 4 follow-ups (1.0% of pilots), 0.456 when it chose 8 follow-ups (2.0% of pilots); its standard deviation across pilots is 0.248. The correlation between the schedule chosen and the effect estimate is −0.021.
Fig. 4 The internal pilot’s own effect estimate, averaged within each follow-up schedule it chose for the rest of the trial, over four thousand pilots of forty patients with a true effect of one half. The dashed line is the true effect.

Counted over four thousand pilots of forty patients with an effect of half a standard deviation, the pilot’s own effect estimate averages 0.503 when it chose one follow-up, 0.514 when it chose two and 0.503 when it chose three, against a true 0.5 and a pilot-to-pilot standard deviation of 0.248. The correlation between the number of follow-ups chosen and the pilot’s effect estimate is −0.021, within one and a half standard errors of zero. The choice reads the curve and nothing else.

That independence is the internal pilot’s version of the spread a pilot supplies, where a pooled variance re-estimated mid-trial was safe for the same reason: what is read is a within-arm quantity, and within-arm quantities are blind to the arms. It would fail for a skewed outcome, whose within-arm covariances and means are correlated through the third moment, and for a correlation computed with the arm labels hidden, which contains the effect — the same distinction a variance that contains the effect found for a look added on a blinded variance.

What the trial should do

Measure the first forty to eighty patients on two follow-ups, and choose the rest’s schedule from them. In a trial of two hundred an arm that nearly halves the worst-case cost of not knowing the correlation curve, or better, and spends no visit and no recruitment the minimax rule would not have spent. In a trial of eight hundred an arm, eighty pilot patients bring the worst case to 2.2%.

Do not run a pilot of twenty. It leaves the worst case within a point of the rule that needs no pilot. The curve has two parameters and two lags to fit them on, and twenty patients cannot.

Keep the effect size the trial is powered for in view. The pilot’s value grows with the trial it steers, as how many subjects would predict for anything whose benefit is a share of the sample size; in a trial of ninety an arm, forty and eighty pilot patients do about equally well, and the choice between them is not worth making.

Read the baseline from after screening if it is to be part of the curve. A baseline taken on the day of enrolment carries the selection that got the patient in, which lowers its correlation with every later visit for a reason the curve does not describe, and a pilot that reads that correlation as fading will choose too few follow-ups.

Compute the correlations within arms. That is what keeps the choice blind to the effect, and it needs no blinding beyond what the trial’s statistician already has; a correlation from lumped data would carry the effect into the schedule.

Every cost is exact given the schedule chosen, from the adjusted analysis’s variance under the true curve; each pilot’s choice is the cheapest schedule under its fitted curve; and each point is three hundred pilots. A pilot run on the schedule it chooses is checked to cost exactly what that schedule costs, and the chosen schedule is checked to be uncorrelated with the pilot’s effect estimate. The internal pilot of twenty is refused as an improvement on the rule that needs no pilot: its worst case is within a tenth of the rule’s.

The same reasoning places this design beside the other adaptations a trial can make from within-arm quantities. Choosing n after looking re-estimates the sample size from an interim variance; a visit before or a visit after chose where one extra measurement should go given the curve; this essay chooses how many measurements from a curve the trial estimates itself. All three are safe for the effect exactly as far as what they read is within-arm, and all three are worth doing exactly as far as the quantity they read was unknown.

Still open: a schedule chosen again as the trial goes

An internal pilot chooses once, after its first patients, and holds the choice. The fit improves with every patient measured, so a trial could choose again after eighty, after a hundred and sixty, each time from everyone measured so far — a sequential choice of schedule, converging on the cheapest as the correlations sharpen. The patients after each look would carry a better choice than those before them.

That design reads its own data repeatedly, and every reading is of within-arm covariances, so the independence argument above still holds at each look. What it costs is the mixture: patients measured early on a wrong schedule cannot be re-measured, and each re-choice changes how many visits the trial books. Whether re-choosing every eighty patients beats a single choice at eighty, and by how much in a trial of two hundred an arm, is computable with the same exact variances, and has not been computed.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Analysis of covarianceAutocorrelationExperimental designMaximin designPilot studySample sizeStatistical powerWorst case