What a sample-size calculation was given

A baseline cut in two

Adjusting a trial's result for a baseline measurement that predicts the outcome shrinks the sample it needs to 1 − ρ² of the unadjusted one. Adjusting for whether that measurement was above its median shrinks it only to 1 − (2/π)ρ² — the cut keeps 63.7% of what the covariate could remove, exactly the fraction a cut outcome keeps. In patients the loss grows with the covariate's strength: a cut costs 12% more patients at a correlation of 0.5, 35% at 0.7 and 2.55 times as many at 0.9, where a simple change from baseline would have done better than the cut adjustment.

Worth reading first: What a p-value does not say.

An outcome cut in two priced the replacement of a measured outcome by whether it crossed a threshold: a cut at the mean keeps 63.7% of the information, 2/π, and a trial that needs 63 patients an arm on the measured outcome needs 102 on the cut one. It ended on the other place trials cut measurements. Baseline variables are dichotomised too — to define subgroups, to stratify the randomisation, and to adjust the analysis — and a baseline cut should weaken the adjustment it was meant to provide. Whether the loss is the same 36% the outcome loses, and what it costs in patients, was not computed.

The fraction turns out to be the same. The cost in patients does not.

The sample a trial needs for the same power, as a share of the unadjusted analysis's, by how strongly a baseline covariate predicts the outcome. Adjusting for the covariate as measured needs one minus the squared correlation of the unadjusted sample; cut at the median, one minus 2/π of the squared correlation. At a correlation of 0.7 those are 0.510 and 0.688; cut in three, 0.611; the change from baseline, 0.600. The change score needs fewer patients than the median-cut adjustment once the correlation passes 0.624.
Fig. 1 The sample a randomised trial needs for the same power, as a share of the unadjusted analysis’s, against the correlation between a baseline covariate and the outcome: adjusted for the covariate as measured, for it cut into five, three and two groups, and analysed as the change from baseline. The dashed line is the unadjusted analysis.

What adjustment buys

A randomised trial’s treatment effect can be estimated by comparing the arms’ mean outcomes. If a baseline measurement predicts the outcome — a blood pressure before treatment predicting the blood pressure after it, a score at entry predicting the score at the end — the comparison can be adjusted for it by analysis of covariance, which removes the part of each patient’s outcome that their baseline predicts. Randomisation makes the adjustment unbiased; its purpose is precision.

With the covariate and outcome jointly normal and correlated ρ\rho, adjustment reduces the outcome’s residual variance from 1 to 1−ρ21 - \rho^2, and the sample a trial needs for a given power is proportional to that variance. At ρ=0.5\rho = 0.5 adjustment cuts the sample to three quarters; at 0.7, to 0.51; at 0.9, to 0.19. The sixty-three or sixty-four patients an arm an unadjusted trial needs for 80% power at half a standard deviation become 33 when a baseline correlated 0.7 with the outcome is adjusted for. It is the largest, cheapest gain in precision available to most trials, and it is the one a cut throws away.

What the cut keeps

Adjusting for whether the covariate was above its median replaces the covariate by a binary indicator. The outcome’s variance explained by that indicator is its squared correlation with the outcome, which for a normal covariate cut at its median is ρ2⋅φ(0)2/(12⋅12)=(2/π) ρ2\rho^2 \cdot \varphi(0)^2/(\tfrac12 \cdot \tfrac12) = (2/\pi)\,\rho^2. So the cut-adjusted analysis’s residual variance is

1−2πρ2,1 - \tfrac{2}{\pi}\rho^2,

and the cut keeps exactly 63.7% of the variance the covariate as measured would have removed — the same 2/π, to every digit, that a cut outcome keeps of its information. That is not a coincidence. In both cases a normal quantity is replaced by the side of its median it fell on, and 2/π is the squared correlation between a standard normal and its own sign.

More categories keep more. Cut in three equal groups the covariate keeps 79.3% of its variance reduction; in four, 86.1%; in five, 89.7%. A cut anywhere but the median keeps less: a single cut at the top tenth keeps 34.2%, the same share a cut outcome keeps at the top tenth.

How much of a baseline covariate's predictive value one cut keeps, by where the cut is made. A cut at the median keeps 2/π, 63.7%, of the variance the covariate would remove; one at the top tenth keeps 34.3%. At a correlation of 0.7 that is 1.349 and 1.631 times the patients the covariate as measured needs.
Fig. 2 The share of a baseline covariate’s predictive variance kept by a single cut, against where the cut is made. The median keeps 2/π; a cut at a clinical threshold far from the middle keeps much less.

What the cut costs in patients

The fraction is the same as for an outcome. The consequence is not, and the difference is the most useful thing here.

For a cut outcome, the sample multiple is the reciprocal of the information kept: 1.57 at a median cut, whatever the effect. For a cut covariate, the sample is proportional to the residual variance, and what matters is the ratio of the two analyses’ residual variances:

1−2πρ21−ρ2,\frac{1 - \tfrac{2}{\pi}\rho^2}{1 - \rho^2},

which depends on how strong the covariate was. A weak covariate removes little variance, so losing a third of that little costs little. A strong one removes most of the variance, so the third it loses is a large share of what is left.

The extra patients a cut baseline covariate costs, against how strongly it predicts the outcome. Adjusting for a median cut rather than the covariate as measured needs 1.036 times the patients at a correlation of 0.3, 1.121 at 0.5, 1.349 at 0.7 and 2.549 at 0.9. Cut in three: 1.199 at 0.7 and 1.881 at 0.9.
Fig. 3 The patients a trial adjusted for a cut covariate needs, as a multiple of the trial adjusted for the covariate as measured, against the covariate’s correlation with the outcome, for a median cut and for three and five categories.
correlation patients, median cut ÷ as measured cut in three ÷ as measured
0.3 1.036 1.020
0.5 1.121 1.069
0.7 1.349 1.199
0.9 2.549 1.881

At a correlation of 0.3 the median cut costs 4% more patients; at 0.5, 12%; at 0.7, 35%; at 0.9, 2.55 times as many. The trial of 63 patients an arm unadjusted, adjusted for a baseline correlated 0.7 with its outcome, needs 33 patients an arm adjusted for the baseline as measured and 44 adjusted for its median split. Three categories recover the same share of the excess at every correlation — 43.1%, which is (79.3% − 63.7%)/(100% − 63.7%) — so at a correlation of 0.9 even a tertile cut needs 1.88 times the patients; five categories recover 71.6%.

The high-correlation case is not exotic. It is the ordinary case of a baseline measurement of the outcome itself — the same blood pressure, the same symptom score, the same biomarker taken before treatment — where correlations of 0.7 to 0.9 are common, and it is exactly where cutting the baseline into “high” and “low” to define strata or subgroups is most tempting.

One trial, four analyses

Power against patients per arm for one trial analysed four ways, baseline correlated 0.7 with the outcomeFor 80% power at an effect of half a standard deviation the unadjusted analysis needs 63 patients an arm, the change from baseline 38, adjustment for the median cut 44 and adjustment for the covariate as measured 33.00.2500.5000.750120406080patients per armpower, effect of half a standard deviationunadjustedchange from baselineadjusted for the median cutadjusted, as measurednormal approximation to each analysis's powerthe dashed line marks 80%
Fig. 4 Power against patients per arm for one trial with a baseline correlated 0.7 with its outcome and an effect of half a standard deviation, analysed four ways. The dashed line marks 80%. The slider sets the correlation.

The four analyses are four readings of the same patients, and the power curves put the choice in the currency a trial is planned in. At a baseline correlated 0.7 with the outcome and an effect of half a standard deviation, 80% power takes 63 patients an arm unadjusted, 38 analysed as the change from baseline, 44 adjusted for the median cut and 33 adjusted for the baseline as measured. The curves are drawn from the same patients; the gap between the two adjusted curves is paid entirely at the analysis, after every patient has been recruited.

At a correlation of 0.9 the picture sharpens. Adjusting for the baseline as measured needs 12 patients an arm; the change score, 13; a tertile adjustment, 23; the median cut, 31. The cut analysis needs more than twice the patients of an analysis anyone could have run on the same data, and the loss would not show in the trial’s report — which states the analysis it ran, not the one it could have.

The outcome’s cut, beside it

How much of a normal outcome's information survives cutting it into two, by where the cut is. For a small shift, a cut at the mean keeps 63.7% of the information, so the trial needs 1.57 times the sample. A cut at the top tenth keeps 34.2% and needs 2.92 times; a cut two standard deviations out keeps 13.1%.
Fig. 5 The information a cut outcome keeps against where the cut is made — the curve the baseline’s kept share follows exactly. At the median both keep 2/π.

The two curves are the same function. A cut at cc standard deviations keeps φ(c)2/[Φ(c)(1−Φ(c))]\varphi(c)^2/[\Phi(c)(1 - \Phi(c))] of what the continuous quantity would give, whether the quantity is the outcome or a covariate, because in both cases the question is how well a normal variable is predicted by the side of cc it falls on. So everything the essay on the cut outcome found about where to cut carries over unchanged: the median is best, a cut a standard deviation out keeps 43.9%, and a cut two standard deviations out keeps 13.1%.

What does not carry over is the step from the kept share to patients, which the section below explains. A trial that cuts both its outcome and its baseline — a responder analysis adjusted for baseline severity category, which is a common combination — pays twice, once at a fixed rate and once at a rate set by how strongly the baseline predicts. Multiplying the two factors at a baseline correlated 0.7 gives 1.57 × 1.35 = 2.12, a little more than twice the patients the measured outcome adjusted for the measured baseline would need — as a rough figure only, since adjusting a binary outcome does not follow the normal arithmetic exactly and the combined loss has not been computed here.

Why the fraction is the same and the cost is not

The two essays on cutting reach the same number, 2/π, from opposite ends of a regression, and the reason they diverge after it is worth spelling out.

A cut outcome loses information about the treatment effect directly: the effect is a shift in the outcome’s distribution, and the side of the median a patient lands on carries 2/π of the information a measured value would. The loss is a fixed proportion of what the trial learns, whatever else is true, so the sample multiple is a constant, 1/(2/π) = 1.57.

A cut covariate loses information about something else — the part of each patient’s outcome that is predictable in advance — and that part is not the thing being estimated. It matters only through the variance it removes from the comparison. When the covariate is weak, the variance it can remove is small and losing a third of it barely changes the residual; when it is strong, the residual after adjustment is small, and a third of the covariate’s removal added back to a small residual multiplies it. The same 2/π that costs a constant factor on the outcome costs a factor that explodes as the covariate’s correlation approaches one.

The cut in the subgroup table

A cut baseline also appears where nobody calls it an adjustment. A trial’s subgroup table — the effect among patients with high and low baseline severity — is the cut covariate twice over: it splits the covariate into two categories and then reports the treatment effect within each, which is the cut adjustment’s two stratum estimates printed separately. Twenty variables nobody stratified on counted how often such a table reverses on some variable by chance, and a subgroup inside its own trial priced reading its verdicts; the arithmetic here adds that the table is also a lossy summary of the precision the baseline could have given the whole trial, before either subgroup is read.

Where the change score overtakes the cut

When the covariate is a baseline measurement of the outcome on the same scale, there is a third analysis: the change from baseline, outcome minus baseline, compared between arms. Its variance is 2(1−ρ)2(1 - \rho) in the same units. It is worse than analysis of covariance at every correlation — which is well known and is why covariance adjustment is recommended — but it does not waste the covariate, and against a cut adjustment it can win.

The change score needs fewer patients than the median-cut adjustment once the correlation exceeds 0.624. At 0.7 the change score needs 0.600 of the unadjusted sample against the cut adjustment’s 0.688; at 0.9, 0.200 against 0.484. So a trial with a strongly predictive baseline that adjusts for “above or below the median” at the analysis is using less of its baseline than a trial that does nothing more sophisticated than subtract it. The simplest analysis of all, and the one analysis of covariance is usually recommended over, beats the cut adjustment across the range where baseline measurements of the outcome usually sit.

Why trials cut their baselines

The cut rarely enters an analysis because anyone chose it for the analysis. It enters through the design.

Stratified randomisation needs categories. A trial that balances its arms on baseline severity randomises within “mild” and “severe”, and a common habit is then to adjust the analysis for the stratification variable as it was used — the cut version — rather than for the measurement it was made from. The stratification itself costs nothing; the choice of analysis does. Adjusting for the covariate as measured is compatible with having stratified on its cut, keeps the balance the stratification bought, and recovers the precision the cut would lose.

Subgroups need categories. A trial that reports its effect in “high” and “low” baseline groups has, for that table, cut its covariate, and a reader who takes the subgroup table as the adjusted analysis inherits the loss. The verdicts in a subgroup table were already a poor guide to a difference; the cut adds a poor adjustment to them.

A threshold is clinically meaningful. A cut at a treatment threshold — a blood pressure of 140, a score of 10 — is rarely at the median, and a cut far from the median keeps far less: at the top tenth, a third of the covariate’s value, which at a correlation of 0.7 is 1.63 times the patients of the covariate as measured. The threshold’s clinical meaning is a reason to report results by it, not a reason to adjust by it.

What the numbers say to a planner

Adjust for the baseline as measured, not as categorised. Covariance adjustment for a pre-specified baseline is standard practice and is what regulatory guidance on adjusted analyses has in mind; the categorised version usually arrives by default, carried over from the randomisation or the subgroup table, rather than by anyone’s choice. At the correlations baseline measurements of the outcome usually have, the difference is between a trial of 33 patients an arm and one of 44, or at a correlation of 0.9 between one of 12 and one of 31, and it is paid for by a decision made when the analysis plan was written.

If a category must be used, use more than two. Three categories recover 43.1% of what the median cut loses and five recover 71.6%, at every correlation — so the number of categories matters most exactly where the covariate is strongest and the loss is largest.

Plan the sample for the analysis that will be run. The chance a trial succeeds depends on the variance the analysis will actually leave, and a trial sized for covariance adjustment and analysed with a cut covariate is underpowered by the ratio in the table — at a correlation of 0.7, a trial planned for 80% power has 67.4%.

The correlation a planner has to guess

Every number above depends on the correlation between baseline and outcome, and a planner rarely knows it well. It is estimated from a pilot or an earlier trial, with the same imprecision the spread a pilot supplies found in a pilot’s standard deviation.

The cost curve makes that imprecision matter unevenly. Near a correlation of 0.3 the cut costs almost nothing whatever the true value, and a planner can be casual. Near 0.8 or 0.9 the cost of the cut changes by a large factor over a range of correlations a small pilot cannot distinguish — 1.65 at 0.8 and 2.55 at 0.9 — so a trial planned for covariance adjustment and analysed with a cut is underpowered by an amount that is itself uncertain. The safe plan is the one whose cost does not depend on the guess, and that is adjusting for the baseline as measured: its sample is proportional to 1−ρ21 - \rho^2, which is uncertain too, but it is the smallest of the options at every correlation, so a mistaken guess never makes it the wrong choice.

The same two over pi, and the patients it costs

Adjusting for a normal baseline covariate cut at its median keeps exactly 2/π of the variance it would remove as measured, the same fraction a median-cut outcome keeps; three categories keep 79.3% and five 89.7%.

The cut costs 1.121 times the patients at a correlation of 0.5, 1.349 at 0.7 and 2.549 at 0.9, relative to adjusting for the covariate as measured.

A change-from-baseline analysis needs fewer patients than the median-cut adjustment once the correlation exceeds 0.624.

Every ratio is exact under bivariate normality, and the variances of all four estimators were checked against four thousand simulated trials of two hundred at a correlation of 0.7: 0.514, 0.691 and 0.598 of the unadjusted variance, against exact values of 0.510, 0.688 and 0.600.

Not claimed: that baseline covariates are normal, or that their relation to the outcome is linear. A skewed covariate cut at its median keeps a different share, and a relation that is itself a step would make the cut the right model rather than a loss. Not claimed either that the change score is the right analysis; it is beaten by covariance adjustment at every correlation, and the comparison here is only with the cut.

Still open: the cut chosen from the data

Every cut here is fixed in advance at a quantile. A cut chosen after looking — the threshold at which the baseline best separates the outcomes, found in the trial’s own data — would appear to keep more of the covariate’s value than a fixed median does, since it was chosen to. A cut chosen after looking at the outcome was found to inflate the effect it reported, and a baseline cut chosen to maximise the adjusted precision would understate the adjusted variance in the same way.

How much a data-chosen baseline cut overstates its own precision, and whether it then reports an adjusted interval narrower than covariance adjustment honestly gives, is a calculation over the choice of threshold that has not been made here.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Analysis of covarianceBaseline adjustmentChange scoreCorrelationDichotomisationRandomisationSample sizeStatistical power