What a sample-size calculation was given

A visit before or a visit after

A trial that can afford one more measurement per patient can take it before treatment, to sharpen the baseline adjustment, or after it, to average the outcome. With every pair of visits correlated 0.5, the second visit after treatment removes a quarter of the variance and the second visit before it removes a twelfth, and it is never the other way round: the baseline's contribution is ρ²(1 − ρ)/(1 + ρ), which peaks at ρ⁵ = 0.090 when ρ is the golden ratio's reciprocal. Below a correlation of one half, a single extra follow-up beats any number of baselines. When the correlation fades with the time between visits, averaging an earlier baseline into the adjustment makes the trial less precise, not more.

Worth reading first: What a p-value does not say.

A baseline cut in two and the cut that fitted best priced what a trial throws away when it adjusts for less than its baseline measurement contains — a median split keeps 2/π of it, a threshold found in the data keeps less and claims more. Both took the baseline as given: one measurement per patient, correlated with the outcome by whatever the disease and the instrument allow. That correlation is not entirely given. A trial that measures its outcome on a noisy instrument — a blood pressure, a symptom score, a walking test — can measure it more than once, and the question is where the extra measurements should go.

The intuitive answer is before treatment. The baseline is what the adjustment leans on, and a baseline that is an average of two readings is a better baseline. The arithmetic says the intuition is backwards: the visit after treatment is worth more, at every correlation, and at the correlations typical of repeated clinical measurements it is worth several times as much.

The variance of the baseline-adjusted estimate for four arrangements of visits, against the correlation between visitsAt a correlation of 0.5 between visits, the adjusted analysis has 0.750 of the unadjusted variance with one visit before treatment and one after, 0.667 with two before, 0.500 with two after and 0.417 with two of each.00.2500.5000.750100.2000.4000.6000.800correlation between any two of a patient's visitsvariance of the adjusted estimate, share of one unadjusted follow-up'sone before, one aftertwo before, one afterone before, two aftertwo before, two aftercompound symmetry, exactdashed: one follow-up, unadjusted
Fig. 1 The variance of the baseline-adjusted treatment estimate, as a share of one unadjusted follow-up’s, against the correlation between any two of a patient’s visits, for four arrangements: one visit before treatment and one after, two before, two after, and two of each. The slider switches to the change-from-baseline analysis.

The model the arithmetic is exact for

The simplest model of a repeatedly measured outcome says that every measurement of a patient is the patient’s own stable level plus independent noise, so that any two visits of one patient are correlated by the same ρ — the share of the variance that is stable. This is compound symmetry, and ρ is the instrument’s test–retest reliability. The treatment shifts every visit after randomisation by the same amount.

Under this model, with pp visits before treatment and rr after, Frison and Pocock worked out in 1992 the variance of the three usual analyses, as multiples of the variance a single post-treatment visit would give:

POST: 1+(r−1)ρr,CHANGE: 1+(r−1)ρr+1+(p−1)ρp−2ρ,ANCOVA: 1+(r−1)ρr−pρ21+(p−1)ρ.\text{POST: } \frac{1 + (r-1)\rho}{r},\qquad \text{CHANGE: } \frac{1 + (r-1)\rho}{r} + \frac{1 + (p-1)\rho}{p} - 2\rho,\qquad \text{ANCOVA: } \frac{1 + (r-1)\rho}{r} - \frac{p\rho^2}{1 + (p-1)\rho}.

POST compares the arms’ mean follow-ups, CHANGE compares the follow-up mean minus the baseline mean, and ANCOVA adjusts the follow-up mean for the baseline mean. The sample a trial needs for a given power is proportional to the variance, so every number below is also a ratio of patients.

Where the second visit goes

At a correlation of 0.5 between visits, one visit on each side of treatment gives the adjusted analysis 0.750 of the unadjusted variance — the 1−ρ21-\rho^2 the essay on the cut baseline started from. A second baseline brings it to 0.667. A second follow-up brings it to 0.500. Two of each give 0.417. In patients, at an effect of half a standard deviation and 80% power: 63 an arm unadjusted, 48 with one visit either side, 42 with a second baseline, 32 with a second follow-up, 27 with two of each.

The general statement is two lines of algebra. Going from one follow-up to two removes

1−ρ2\frac{1-\rho}{2}

of the variance, and going from one baseline to two removes

ρ2(1−ρ)1+ρ.\frac{\rho^2(1-\rho)}{1+\rho}.

Their ratio is 2ρ2/(1+ρ)2\rho^2/(1+\rho), which is below one for every correlation short of perfect. The second visit after treatment is always worth more than the second visit before it. At a correlation of 0.5 it is worth three times as much, 0.250 against 0.083; at 0.8 the two are closer, 0.100 against 0.071.

How much variance a second visit removes from the adjusted estimate, taken before treatment or after it. With ρ the correlation between visits, a second follow-up visit removes (1 − ρ)/2 of the variance and a second baseline visit removes the square of ρ times (1 − ρ)/(1 + ρ). At ρ = 0.5 those are 0.2500 and 0.0833; at 0.8, 0.1000 and 0.0711. The baseline's contribution peaks at 0.0902 near ρ = 0.62.
Fig. 2 The variance a second visit removes from the adjusted estimate, taken after treatment and taken before it, against the correlation between visits. The follow-up line falls straight from one half to zero; the baseline’s curve rises and falls, peaking at 0.090.

The baseline’s curve has a shape worth a sentence. A second baseline is useless when visits are uncorrelated, because there is nothing to adjust for, and useless when they are perfectly correlated, because one baseline already knows the patient’s level exactly. In between it peaks, and the peak is at a correlation of (5−1)/2=0.618(\sqrt5 - 1)/2 = 0.618, where 1−ρ=ρ21 - \rho = \rho^2 and 1+ρ=1/ρ1 + \rho = 1/\rho, so that the gain is exactly ρ5=\rho^5 = 0.090. That is the most a second baseline can ever do under this model: remove nine per cent of the unadjusted variance. A second follow-up removes more than that at every correlation below 0.82.

Why the visit after treatment is the better buy

The two extra visits do different jobs, and the difference explains the asymmetry.

A follow-up visit measures the thing being compared. Averaging two of them halves the part of the outcome’s variance that is visit-to-visit noise, directly, and that noise is a share 1−ρ1 - \rho of everything.

A baseline visit measures something that is only used to predict the thing being compared. Averaging two baselines makes the prediction of the patient’s stable level better, but the adjustment can only ever remove the stable part of the follow-up’s variance, a share ρ\rho, and it already removes most of that with one baseline. The second baseline improves a correction; the second follow-up improves the measurement being corrected.

There is a limit that makes this vivid. With infinitely many baselines the patient’s stable level is known exactly and the adjusted variance with one follow-up falls to 1−ρ1 - \rho — the follow-up’s own noise, which no baseline can touch. With one baseline and two follow-ups it is (1+ρ)/2−ρ2(1 + \rho)/2 - \rho^2. The two are equal at ρ = 0.5, and below that correlation one extra follow-up visit beats any number of baselines. At a reliability of 0.3 — a symptom score, a single-day activity count — a trial that doubles its follow-up visits gains more than a trial that measures its baseline every day for a month.

What the change score does with the same visits

The variance of the change-from-baseline estimate for four arrangements of visits, against the correlation between visits. At a correlation of 0.5 between visits, the change score has 1.000 of the unadjusted variance with one visit before treatment and one after, 0.750 with two before, 0.750 with two after and 0.500 with two of each.
Fig. 3 The change-from-baseline analysis under the same four arrangements. A second visit on either side of treatment helps it by the same amount, and at low correlations it is worse than ignoring the baseline altogether.

The change score treats its two means symmetrically — it subtracts one from the other with a coefficient of one — so a second visit on either side of treatment helps it by exactly the same amount, (1−ρ)/2(1 - \rho)/2. At a correlation of 0.5 one visit each side gives it 1.000 of the unadjusted variance: subtracting a baseline that carries as much noise as signal gains nothing. Two before and one after, or the reverse, give 0.750, and two of each 0.500.

So the change score is the one analysis for which extra baselines are worth as much as extra follow-ups, and it is the analysis covariance adjustment beats at every correlation. A trial planning to analyse by change has a reason to add baseline visits; a trial planning covariance adjustment, which is what the guidance on adjusted analyses recommends, mostly does not. The design and the analysis are one decision, and a protocol that plans extra baseline visits and an ANCOVA has spent the visits on the analysis it did not choose.

The table a planner reads

The same arithmetic in patients, at an effect of half a standard deviation, 80% power and a two-sided 5% test, against the 63 an arm the unadjusted single-visit analysis needs:

correlation between visits one before, one after two before, one after one before, two after two of each
0.3 58 55 36 33
0.5 48 42 32 27
0.7 33 27 23 18
0.9 12 10 9 7

Read across a row and the pattern is the same at every reliability: the second follow-up saves more patients than the second baseline, and two of each is best. Read down a column and the pattern that matters for planning appears. At a correlation of 0.3 the baseline is nearly useless — one visit each side needs 58 patients against 63 with no baseline at all — and the whole gain available is in the follow-up, where a second visit takes the trial from 58 to 36. At 0.9 the baseline is doing almost everything, the trial is already small, and every extra visit of either kind saves one or two patients. The instrument’s reliability decides not only how much repetition is worth but which side of treatment it belongs on.

The one job only a baseline visit does

There is a reason trials take more than one pre-treatment measurement that has nothing to do with precision, and it is worth separating from the arithmetic above because it is often the real reason.

Many trials enrol on the measurement itself: blood pressure above a threshold, a symptom score above a cut-off. The reading that qualified a patient was selected for being high, and part of what made it high was that day’s noise. Measured again, the same patients come down whether or not anything is done to them — the measurement that got them enrolled measured the fall at 0.702 standard deviations for the top tenth of a single screening reading, and at nothing when the change is measured from a fresh reading taken after enrolment. So a second pre-treatment visit, separate from the one used to decide eligibility, is what makes a within-arm change describe the patients rather than the selection.

That job is real and it is not a job the randomised comparison needs. The difference between arms is unaffected by regression to the mean because both arms regress by the same amount, and covariance adjustment for the screening value itself remains unbiased because selecting on a covariate does not bias a regression on it. The fresh baseline matters for the numbers a trial reports about each arm — the improvement “on treatment”, the proportion who “responded” — which are the numbers readers most often misread, and for any analysis that compares patients who were not randomised, where two analyses of one baseline disagree by exactly the amount the groups were selected apart. A protocol that takes a screening reading and a separate randomisation reading is buying honesty in its descriptive tables, not precision in its primary comparison, and it should know which it is paying for.

What the budget buys

Visits are not the whole cost. Recruiting and consenting a patient, screening them and following them for the trial’s duration costs something whatever the number of measurements, so the useful unit is the cost of reaching a fixed power: the number of patients, proportional to the variance, times the cost of each patient in visits.

With a patient costing five visits’ worth to recruit and a correlation of 0.7 between visits, the cheapest compound-symmetry design takes one baseline and three follow-ups, at 0.782 of the cost of one visit on each side — 20 patients an arm instead of 33, each measured four times instead of twice. Beyond three, each extra follow-up removes less than its own cost.

What further follow-up visits cost for the same power, when a patient costs 5 visits to recruit, adjacent visits correlated 0.7. Under compound symmetry the cheapest design takes 3 follow-up visits, at 0.782 of the cost of one; with a stable level carrying 0.4 of the variance and the rest fading between visits, it takes 1, and 8 follow-ups cost 1.421 of one.
Fig. 4 The cost of reaching the same power against the number of follow-up visits, with one baseline, when a patient costs five visits to recruit and adjacent visits correlate 0.7. Equal correlation between all visits favours three follow-ups; a correlation that fades with time favours one.

That conclusion rests entirely on the model, and the second curve shows how much. Compound symmetry says the correlation between two visits does not depend on how far apart they are, so the fourth follow-up is as informative about the patient’s level as the first. Real repeated measurements rarely behave like that. Visits close in time share more — a patient’s bad week, a season, the same assessor — and the share that is genuinely stable is smaller than the correlation between adjacent visits suggests.

When the correlation fades with time

Take a model in which a stable patient level carries 0.4 of the variance and the rest is a serial component whose correlation halves from one visit to the next, so that adjacent visits correlate 0.7, visits two apart 0.55, and visits far apart only 0.4. The adjacent correlation is the same as in the budget above. The design conclusions are not.

With one baseline and one follow-up the adjusted variance is 0.510, as before. Adding follow-ups now helps much less, because each additional one is further from the baseline and shares less with it, and at five visits’ recruitment cost one follow-up is already the cheapest design: the second raises the cost by 3%, and eight follow-ups cost 1.421 times as much as one.

Adding baselines does something stranger.

Adding earlier baseline visits when correlation fades with time: adjusting for their mean against adjusting for each. A stable patient level carries 0.4 of the variance and adjacent visits correlate 0.7. With one baseline the adjusted variance is 0.510. Adjusting for the mean of two baselines gives 0.540 and of six 0.608; adjusting for each separately gives 0.503 and 0.486.
Fig. 5 Adding earlier baseline visits when a stable level carries 0.4 of the variance and adjacent visits correlate 0.7: the adjusted variance against the number of baselines, adjusting for their mean and adjusting for each separately, with one follow-up.

The usual practice with several baseline measurements is to average them and adjust for the average. Under compound symmetry that is exactly right, because every baseline is equally informative and the average is the combination a regression would choose. When the correlation fades with time it is wrong, and not slightly: the baseline nearest to treatment is the best predictor of the follow-up, earlier ones are worse, and an equal-weight average dilutes the good one with the poor ones. Adjusting for the mean of two baselines gives 0.540 where one gave 0.510; the mean of six gives 0.608. The trial spent five extra visits per patient to become less precise.

Adjusting for each baseline as a separate covariate cannot lose, because a regression on several predictors can always put all its weight on one of them. It gives 0.503 with two baselines and 0.486 with six — a small gain, bought at a price. The earlier baselines carry information about the stable level, and in a regression they are used to strip the serial noise out of the latest one, which is a real improvement and a modest one.

What a protocol should decide, and when

Spend an extra measurement after treatment before spending it before. Under compound symmetry this is exact at every correlation, and under a fading correlation it remains true for the first extra visit. The exception is a trial analysed by change from baseline, and the repair there is the analysis, not the schedule.

If several baselines are taken, adjust for them as separate covariates, or for the latest. The mean of the baselines is the right summary only under compound symmetry, and it is the one summary that can make things worse when the correlation fades. Averaging a run-in period’s measurements into a single baseline is common and, on this arithmetic, frequently a mistake — though the loss is modest wherever the stable level dominates.

Size the trial for the analysis and the schedule together. How many subjects computed sixty-four an arm from one outcome’s standard deviation; with repeated visits, the variance in that calculation is one of the expressions above, and the choice of pp and rr changes the answer by a factor of two before a single patient is recruited. It is the same point the chance a trial succeeds made about the effect, arriving through the measurement.

Know which correlation model the plan assumes. The two curves in the budget figure start from the same adjacent correlation and disagree about whether two follow-ups are worth paying for. A plan built on compound symmetry that meets a fading correlation over-buys visits; one built on a fading correlation that meets compound symmetry under-buys them.

Why the gains add in patients, not in certainty

Everything above is about the treatment estimate’s variance, and none of it changes what the trial estimates. Averaging two follow-ups estimates the treatment’s effect on the average of two visits, which is the same as its effect on one visit under a model in which the treatment shifts every visit equally. If the effect grows or wanes between the visits, the averaged outcome estimates an average effect over the follow-up window, and the design has quietly changed the question. That is usually a defensible question — a treatment’s effect over a period rather than on a day — but it should be the protocol’s choice rather than a side effect of buying precision. The baseline visits carry no such cost: they are taken before randomisation and cannot be touched by the treatment, which is why adjusting for them never changes the estimand and only the precision.

What is exact here and what is assumed

Under compound symmetry, a second follow-up visit removes (1−ρ)/2(1-\rho)/2 of the unadjusted variance and a second baseline ρ2(1−ρ)/(1+ρ)\rho^2(1-\rho)/(1+\rho), so the follow-up is worth more at every correlation; the baseline’s gain peaks at ρ5=0.090\rho^5 = 0.090 at ρ=0.618\rho = 0.618.

Below a correlation of one half, one extra follow-up beats any number of baselines.

When a stable level carries 0.4 of the variance and adjacent visits correlate 0.7, adjusting for the mean of two baselines gives 0.540 of the unadjusted variance, against 0.510 for one; adjusting for each separately gives 0.503.

The compound-symmetry variances are Frison and Pocock’s closed forms and agree to twelve digits with the general covariance sums used for the fading model; both were checked against four thousand simulated trials of a hundred patients an arm, where the adjusted analysis with one visit each side read 0.733 against an exact 0.750 and with two of each 0.427 against 0.417.

Not claimed: that either correlation model describes any particular outcome. Real measurements have learning effects at the first visit, drift over calendar time, and treatment effects that change between visits, and each of these breaks the arithmetic in its own direction. Not claimed either that visits and patients are the only costs, or that a trial’s follow-up visits can be added without lengthening the trial.

Still open: the correlation a pilot can see

The two models used here agree about adjacent visits and disagree about whether a fourth follow-up is worth paying for. A trial choosing its schedule therefore needs to know not the correlation between two visits but how that correlation falls with the time between them, and it usually has only a pilot, or an earlier trial’s summary, to learn it from.

How many patients and how many visits a pilot needs to tell compound symmetry from a correlation that halves each visit, and what the schedule a plan chooses under a misidentified model actually costs, is the same kind of question the spread a pilot supplies answered for a standard deviation — and harder, because the quantity to be estimated is a curve rather than a number. It has not been computed here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Analysis of covarianceAutocorrelationBaseline adjustmentChange scoreExperimental designSample sizeStatistical powerTest–retest reliability