A visit before or a visit after
Worth reading first: What a p-value does not say.
A baseline cut in two and the cut that fitted best priced what a trial throws away when it adjusts for less than its baseline measurement contains — a median split keeps 2/π of it, a threshold found in the data keeps less and claims more. Both took the baseline as given: one measurement per patient, correlated with the outcome by whatever the disease and the instrument allow. That correlation is not entirely given. A trial that measures its outcome on a noisy instrument — a blood pressure, a symptom score, a walking test — can measure it more than once, and the question is where the extra measurements should go.
The intuitive answer is before treatment. The baseline is what the adjustment leans on, and a baseline that is an average of two readings is a better baseline. The arithmetic says the intuition is backwards: the visit after treatment is worth more, at every correlation, and at the correlations typical of repeated clinical measurements it is worth several times as much.
The model the arithmetic is exact for
The simplest model of a repeatedly measured outcome says that every measurement of a patient is the patient’s own stable level plus independent noise, so that any two visits of one patient are correlated by the same ρ — the share of the variance that is stable. This is compound symmetry, and ρ is the instrument’s test–retest reliability. The treatment shifts every visit after randomisation by the same amount.
Under this model, with visits before treatment and after, Frison and Pocock worked out in 1992 the variance of the three usual analyses, as multiples of the variance a single post-treatment visit would give:
POST compares the arms’ mean follow-ups, CHANGE compares the follow-up mean minus the baseline mean, and ANCOVA adjusts the follow-up mean for the baseline mean. The sample a trial needs for a given power is proportional to the variance, so every number below is also a ratio of patients.
Where the second visit goes
At a correlation of 0.5 between visits, one visit on each side of treatment gives the adjusted analysis 0.750 of the unadjusted variance — the the essay on the cut baseline started from. A second baseline brings it to 0.667. A second follow-up brings it to 0.500. Two of each give 0.417. In patients, at an effect of half a standard deviation and 80% power: 63 an arm unadjusted, 48 with one visit either side, 42 with a second baseline, 32 with a second follow-up, 27 with two of each.
The general statement is two lines of algebra. Going from one follow-up to two removes
of the variance, and going from one baseline to two removes
Their ratio is , which is below one for every correlation short of perfect. The second visit after treatment is always worth more than the second visit before it. At a correlation of 0.5 it is worth three times as much, 0.250 against 0.083; at 0.8 the two are closer, 0.100 against 0.071.
The baseline’s curve has a shape worth a sentence. A second baseline is useless when visits are uncorrelated, because there is nothing to adjust for, and useless when they are perfectly correlated, because one baseline already knows the patient’s level exactly. In between it peaks, and the peak is at a correlation of , where and , so that the gain is exactly 0.090. That is the most a second baseline can ever do under this model: remove nine per cent of the unadjusted variance. A second follow-up removes more than that at every correlation below 0.82.
Why the visit after treatment is the better buy
The two extra visits do different jobs, and the difference explains the asymmetry.
A follow-up visit measures the thing being compared. Averaging two of them halves the part of the outcome’s variance that is visit-to-visit noise, directly, and that noise is a share of everything.
A baseline visit measures something that is only used to predict the thing being compared. Averaging two baselines makes the prediction of the patient’s stable level better, but the adjustment can only ever remove the stable part of the follow-up’s variance, a share , and it already removes most of that with one baseline. The second baseline improves a correction; the second follow-up improves the measurement being corrected.
There is a limit that makes this vivid. With infinitely many baselines the patient’s stable level is known exactly and the adjusted variance with one follow-up falls to — the follow-up’s own noise, which no baseline can touch. With one baseline and two follow-ups it is . The two are equal at ρ = 0.5, and below that correlation one extra follow-up visit beats any number of baselines. At a reliability of 0.3 — a symptom score, a single-day activity count — a trial that doubles its follow-up visits gains more than a trial that measures its baseline every day for a month.
What the change score does with the same visits
The change score treats its two means symmetrically — it subtracts one from the other with a coefficient of one — so a second visit on either side of treatment helps it by exactly the same amount, . At a correlation of 0.5 one visit each side gives it 1.000 of the unadjusted variance: subtracting a baseline that carries as much noise as signal gains nothing. Two before and one after, or the reverse, give 0.750, and two of each 0.500.
So the change score is the one analysis for which extra baselines are worth as much as extra follow-ups, and it is the analysis covariance adjustment beats at every correlation. A trial planning to analyse by change has a reason to add baseline visits; a trial planning covariance adjustment, which is what the guidance on adjusted analyses recommends, mostly does not. The design and the analysis are one decision, and a protocol that plans extra baseline visits and an ANCOVA has spent the visits on the analysis it did not choose.
The table a planner reads
The same arithmetic in patients, at an effect of half a standard deviation, 80% power and a two-sided 5% test, against the 63 an arm the unadjusted single-visit analysis needs:
| correlation between visits | one before, one after | two before, one after | one before, two after | two of each |
|---|---|---|---|---|
| 0.3 | 58 | 55 | 36 | 33 |
| 0.5 | 48 | 42 | 32 | 27 |
| 0.7 | 33 | 27 | 23 | 18 |
| 0.9 | 12 | 10 | 9 | 7 |
Read across a row and the pattern is the same at every reliability: the second follow-up saves more patients than the second baseline, and two of each is best. Read down a column and the pattern that matters for planning appears. At a correlation of 0.3 the baseline is nearly useless — one visit each side needs 58 patients against 63 with no baseline at all — and the whole gain available is in the follow-up, where a second visit takes the trial from 58 to 36. At 0.9 the baseline is doing almost everything, the trial is already small, and every extra visit of either kind saves one or two patients. The instrument’s reliability decides not only how much repetition is worth but which side of treatment it belongs on.
The one job only a baseline visit does
There is a reason trials take more than one pre-treatment measurement that has nothing to do with precision, and it is worth separating from the arithmetic above because it is often the real reason.
Many trials enrol on the measurement itself: blood pressure above a threshold, a symptom score above a cut-off. The reading that qualified a patient was selected for being high, and part of what made it high was that day’s noise. Measured again, the same patients come down whether or not anything is done to them — the measurement that got them enrolled measured the fall at 0.702 standard deviations for the top tenth of a single screening reading, and at nothing when the change is measured from a fresh reading taken after enrolment. So a second pre-treatment visit, separate from the one used to decide eligibility, is what makes a within-arm change describe the patients rather than the selection.
That job is real and it is not a job the randomised comparison needs. The difference between arms is unaffected by regression to the mean because both arms regress by the same amount, and covariance adjustment for the screening value itself remains unbiased because selecting on a covariate does not bias a regression on it. The fresh baseline matters for the numbers a trial reports about each arm — the improvement “on treatment”, the proportion who “responded” — which are the numbers readers most often misread, and for any analysis that compares patients who were not randomised, where two analyses of one baseline disagree by exactly the amount the groups were selected apart. A protocol that takes a screening reading and a separate randomisation reading is buying honesty in its descriptive tables, not precision in its primary comparison, and it should know which it is paying for.
What the budget buys
Visits are not the whole cost. Recruiting and consenting a patient, screening them and following them for the trial’s duration costs something whatever the number of measurements, so the useful unit is the cost of reaching a fixed power: the number of patients, proportional to the variance, times the cost of each patient in visits.
With a patient costing five visits’ worth to recruit and a correlation of 0.7 between visits, the cheapest compound-symmetry design takes one baseline and three follow-ups, at 0.782 of the cost of one visit on each side — 20 patients an arm instead of 33, each measured four times instead of twice. Beyond three, each extra follow-up removes less than its own cost.
That conclusion rests entirely on the model, and the second curve shows how much. Compound symmetry says the correlation between two visits does not depend on how far apart they are, so the fourth follow-up is as informative about the patient’s level as the first. Real repeated measurements rarely behave like that. Visits close in time share more — a patient’s bad week, a season, the same assessor — and the share that is genuinely stable is smaller than the correlation between adjacent visits suggests.
When the correlation fades with time
Take a model in which a stable patient level carries 0.4 of the variance and the rest is a serial component whose correlation halves from one visit to the next, so that adjacent visits correlate 0.7, visits two apart 0.55, and visits far apart only 0.4. The adjacent correlation is the same as in the budget above. The design conclusions are not.
With one baseline and one follow-up the adjusted variance is 0.510, as before. Adding follow-ups now helps much less, because each additional one is further from the baseline and shares less with it, and at five visits’ recruitment cost one follow-up is already the cheapest design: the second raises the cost by 3%, and eight follow-ups cost 1.421 times as much as one.
Adding baselines does something stranger.
The usual practice with several baseline measurements is to average them and adjust for the average. Under compound symmetry that is exactly right, because every baseline is equally informative and the average is the combination a regression would choose. When the correlation fades with time it is wrong, and not slightly: the baseline nearest to treatment is the best predictor of the follow-up, earlier ones are worse, and an equal-weight average dilutes the good one with the poor ones. Adjusting for the mean of two baselines gives 0.540 where one gave 0.510; the mean of six gives 0.608. The trial spent five extra visits per patient to become less precise.
Adjusting for each baseline as a separate covariate cannot lose, because a regression on several predictors can always put all its weight on one of them. It gives 0.503 with two baselines and 0.486 with six — a small gain, bought at a price. The earlier baselines carry information about the stable level, and in a regression they are used to strip the serial noise out of the latest one, which is a real improvement and a modest one.
What a protocol should decide, and when
Spend an extra measurement after treatment before spending it before. Under compound symmetry this is exact at every correlation, and under a fading correlation it remains true for the first extra visit. The exception is a trial analysed by change from baseline, and the repair there is the analysis, not the schedule.
If several baselines are taken, adjust for them as separate covariates, or for the latest. The mean of the baselines is the right summary only under compound symmetry, and it is the one summary that can make things worse when the correlation fades. Averaging a run-in period’s measurements into a single baseline is common and, on this arithmetic, frequently a mistake — though the loss is modest wherever the stable level dominates.
Size the trial for the analysis and the schedule together. How many subjects computed sixty-four an arm from one outcome’s standard deviation; with repeated visits, the variance in that calculation is one of the expressions above, and the choice of and changes the answer by a factor of two before a single patient is recruited. It is the same point the chance a trial succeeds made about the effect, arriving through the measurement.
Know which correlation model the plan assumes. The two curves in the budget figure start from the same adjacent correlation and disagree about whether two follow-ups are worth paying for. A plan built on compound symmetry that meets a fading correlation over-buys visits; one built on a fading correlation that meets compound symmetry under-buys them.
Why the gains add in patients, not in certainty
Everything above is about the treatment estimate’s variance, and none of it changes what the trial estimates. Averaging two follow-ups estimates the treatment’s effect on the average of two visits, which is the same as its effect on one visit under a model in which the treatment shifts every visit equally. If the effect grows or wanes between the visits, the averaged outcome estimates an average effect over the follow-up window, and the design has quietly changed the question. That is usually a defensible question — a treatment’s effect over a period rather than on a day — but it should be the protocol’s choice rather than a side effect of buying precision. The baseline visits carry no such cost: they are taken before randomisation and cannot be touched by the treatment, which is why adjusting for them never changes the estimand and only the precision.
What is exact here and what is assumed
Under compound symmetry, a second follow-up visit removes of the unadjusted variance and a second baseline , so the follow-up is worth more at every correlation; the baseline’s gain peaks at at .
Below a correlation of one half, one extra follow-up beats any number of baselines.
When a stable level carries 0.4 of the variance and adjacent visits correlate 0.7, adjusting for the mean of two baselines gives 0.540 of the unadjusted variance, against 0.510 for one; adjusting for each separately gives 0.503.
The compound-symmetry variances are Frison and Pocock’s closed forms and agree to twelve digits with the general covariance sums used for the fading model; both were checked against four thousand simulated trials of a hundred patients an arm, where the adjusted analysis with one visit each side read 0.733 against an exact 0.750 and with two of each 0.427 against 0.417.
Not claimed: that either correlation model describes any particular outcome. Real measurements have learning effects at the first visit, drift over calendar time, and treatment effects that change between visits, and each of these breaks the arithmetic in its own direction. Not claimed either that visits and patients are the only costs, or that a trial’s follow-up visits can be added without lengthening the trial.
Still open: the correlation a pilot can see
The two models used here agree about adjacent visits and disagree about whether a fourth follow-up is worth paying for. A trial choosing its schedule therefore needs to know not the correlation between two visits but how that correlation falls with the time between them, and it usually has only a pilot, or an earlier trial’s summary, to learn it from.
How many patients and how many visits a pilot needs to tell compound symmetry from a correlation that halves each visit, and what the schedule a plan chooses under a misidentified model actually costs, is the same kind of question the spread a pilot supplies answered for a standard deviation — and harder, because the quantity to be estimated is a curve rather than a number. It has not been computed here.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Not half and half — both name experimental design, sample size, statistical power
- The check before the standard error — both name autocorrelation, sample size, statistical power
- A boundary for giving up — both name sample size, statistical power
- A coverage table with its own error — both name sample size, statistical power
- A detector built for the ordering — both name autocorrelation, statistical power
- Allocating on a guess — both name experimental design, sample size
Named objects
A flat tag is an object no other essay names yet.
Analysis of covarianceAutocorrelationBaseline adjustmentChange scoreExperimental designSample sizeStatistical powerTest–retest reliability