What a sample-size calculation was given

The cut that fitted best

A trial that adjusts for its baseline cut at whichever of the sample's nine deciles fits the outcomes best — without ever looking at the treatment difference — reports a variance 0.671 of the unadjusted analysis's where its estimates actually have 0.736, which is worse than the 0.700 a median fixed in advance delivers. The sample says the chosen cut kept 72.4% of what the baseline could remove; the population says it kept less than a median. A true null is rejected 5.87% of the time at fifty patients an arm and 9.31% at ten, and a cut chosen for its p-value rejects 14.22%.

Worth reading first: What a p-value does not say.

A baseline cut in two priced a covariate that a randomised trial adjusts for after cutting it at a quantile chosen in advance. Cut at its median, a normal baseline keeps exactly 2/π of the variance it would remove as measured, and at a correlation of 0.7 with the outcome that costs 35% more patients. Every cut there was fixed before the data arrived. The essay ended on the cut that is not: a threshold found in the trial’s own data, the value at which the baseline best separates the outcomes. Such a cut should appear to keep more than 2/π, because it was chosen to. Whether it actually does, and whether the precision it reports is precision it has, was left uncomputed.

It appears to keep more. It has less. And it says so to nobody, because the only numbers a trial reports are the ones computed from the sample that chose the cut.

What four adjusted analyses report about their own precision and what they have, 50 patients an arm, baseline correlated 0.7Reported against actual variance of the treatment estimate, as shares of the unadjusted analysis's, over 20,000 trials of 50 patients an arm. Adjusted for the baseline as measured: 0.529 reported, 0.518 actual. Median cut fixed in advance: 0.716 and 0.700. The decile cut with the best fit: 0.671 and 0.736. The decile cut with the smallest p-value: 0.762 and 1.615.00.50011.50variance of the treatment estimate, share of the unadjusted analysis'sadjusted, as measuredmedian cut, fixed in advancecut chosen for the best fitcut chosen for the smallest popen: reported · filled: actual20,000 trials, cuts at the sample's nine decilesa chosen cut claims more than it has
Fig. 1 Four adjusted analyses of the same twenty thousand trials of fifty patients an arm, with a baseline correlated 0.7 with the outcome. For each, the open circle is the variance of the treatment estimate the analysis reports and the filled one is the variance its estimates actually have, both as shares of the unadjusted analysis’s. The slider sets the patients per arm.

Two ways to let the data choose a threshold

A cut chosen from the data can be chosen in two ways that feel very different and are worth keeping apart.

The first looks innocent. The analyst never compares the arms while choosing. Among a set of candidate thresholds — here the nine deciles of the pooled baseline — the chosen one is the cut that leaves the smallest residual variance in the adjusted model, which is the same as the cut at which “high” and “low” baseline patients differ most in their outcomes. It is the natural way to find a clinically meaningful break — the same move that sets a diagnostic threshold where the data separate best — it can be done before unblinding, and it is the choice the essay on the cut baseline named as the obvious next step: the cut that maximises the adjusted precision. Here it is called the best-fit cut.

The second is the one nobody defends. The analyst runs the adjusted analysis at every candidate cut and reports the one with the smallest p-value for the treatment. That is the forking path in its plainest form, and the interest here is only in how much worse it is than the first, since the first is what honest analysts actually do.

Both are compared with two analyses that choose nothing: the median cut, written into the protocol, and adjustment for the baseline as measured. All four are fitted by least squares with the arm and the covariate as predictors, and each reports the standard error that least squares gives it — which is exactly the problem, because least squares does not know which of its columns was chosen by looking.

What the chosen cut reports and what it has

At fifty patients an arm and a correlation of 0.7, the best-fit cut reports a variance for the treatment estimate of 0.671 of the unadjusted analysis’s. Across twenty thousand trials its estimates actually vary by 0.736. It claims 10% more precision than it has — its reported variance is 0.911 of its actual — where the median fixed in advance reports 0.716 and has 0.700, the small excess that any estimated standard error carries.

The surprising half is the second number. The best-fit cut’s actual variance, 0.736, is larger than the fixed median’s, 0.700. The cut chosen because it fitted best delivers a less precise treatment estimate than the cut chosen because it was in the protocol. It reports better precision than the median and has worse, and the reversal is not small: on the reported numbers the chosen cut looks 6% more precise than the median, and on the actual ones it is 5% less.

The baseline as measured is untouched by any of this. It reports 0.529 and has 0.518, which is 1−ρ2=0.511-\rho^2 = 0.51 plus the price of estimating one coefficient in a hundred patients. Every cut, fixed or chosen, is compared against that number, and the chosen one is further from it than the fixed one.

Why choosing among nine cuts is worse than using one

The reason the chosen cut does worse is that in the population there is nothing to choose. For a normal baseline and a linear relation with the outcome, the cut that removes the most variance is the median, and it removes 2/π of what the covariate as measured would. Any other cut removes less — at the top decile, about a third as much. So every departure from the median is a loss, and the best-fit rule departs from it most of the time.

Where the best-fitting cut of a baseline correlated 0.7 with the outcome lands, 50 patients an arm. Among the sample's nine deciles, the cut that leaves the smallest residual variance is the median on 21.8% of trials and one of the two outermost deciles on 1.5%; the five middle deciles take 88.1%.
Fig. 2 The decile at which the best-fitting cut sits, over twenty thousand trials of fifty patients an arm with a baseline correlated 0.7 with the outcome. The median is chosen on about a fifth of trials; the rest of the time noise moves the cut away from it.

It departs because the residual variance at each candidate cut is itself an estimate. With a hundred patients, the difference in fit between the median and the fourth or sixth decile is smaller than the noise in either, so which one wins is decided mostly by where the sample’s outcomes happened to fall. The median is the winner on 21.8% of trials. The five middle deciles take 88.1% between them, and the two outermost deciles 1.5%. The cut wanders, and each trial’s analysis is then run at a threshold that is worse in the population than the one it could have fixed in advance.

The wandering is also what makes the reported variance too small. The chosen cut is the one whose residuals happened to be smallest, and the smallest of nine correlated residual variances is below the typical one. The standard error is built from that residual variance, so it inherits the flattery. Nothing in the fitted model records that nine models were tried; it reports three estimated coefficients and charges three degrees of freedom for them, when the search spent more. R2R^2 is not a measure of fit measured the same optimism in a different setting — useless predictors raising R2R^2 by construction — and a searched threshold is a useless predictor chosen to look useful.

The stronger the baseline, the less the cut wanders, because the difference in fit between the median and its neighbours grows with the correlation while the noise does not. At a correlation of 0.9 the middle three deciles take 79.0% of trials between them; at 0.3 the choice is spread across all nine, the median taking 12.8% and each outermost decile 8.0%, so a weak baseline is cut at the tenth or the ninetieth percentile about one trial in six.

The share the sample says it kept

The flattery can be put in the unit the essay on the cut baseline used. In each trial, the share of the covariate’s variance reduction a cut keeps can be measured from the trial’s own residual variances: how much the cut model removed, as a share of how much the model with the baseline as measured removed. For the fixed median this reads 0.631 on average at fifty patients an arm, a little under 2/π because the sample is finite. For the best-fit cut it reads 0.724.

How much of the baseline's value a chosen cut appears to keep, against the fixed median, 50 patients an arm. In its own sample the best-fitting decile cut appears to keep 1.044 of the variance the baseline as measured removes at a correlation of 0.3, 0.724 at 0.7 and 0.665 at 0.9; the median fixed in advance appears to keep 0.625, 0.631 and 0.632, against 2/π, 0.637, in the population.
Fig. 3 The share of the baseline’s variance reduction the best-fitting cut appears to keep, measured in each trial’s own sample, against the baseline’s correlation with the outcome; the fixed median beside it, sitting on 2/π. Fifty patients an arm, ten thousand trials a point.

So a trial that searched for its threshold would conclude that its cut kept nearly three quarters of the baseline’s value, and would be wrong about its cut in the direction that matters: the cut kept less than a median would have.

At weaker correlations the flattery grows, because the true reduction the cut can offer is smaller and the noise it fits is not. At a correlation of 0.3 the chosen cut appears to keep 1.051 of what the baseline as measured removes — more than the covariate itself. That answers the question the cut baseline left in its sharpest form. At fifty patients an arm and a correlation of 0.3, the best-fit cut reports a variance of 0.939 of the unadjusted analysis’s, and adjustment for the baseline as measured reports 0.944: the chosen cut prints a narrower interval than honest covariance adjustment. Its estimates actually vary by 0.976, where covariance adjustment’s vary by 0.925.

In a small trial the numbers turn from embarrassing to perverse. At twenty patients an arm and a correlation of 0.3 the best-fit cut reports 0.914 and has 1.010 — its estimates are more variable than the unadjusted comparison’s, so the adjustment added noise, while its printed interval is the narrowest of the four analyses.

The rejections a true null collects

A standard error that is too small produces rejections a true null should not produce. At fifty patients an arm and a correlation of 0.7 the best-fit cut rejects a true null 5.87% of the time at a nominal 5%, where the fixed median rejects 4.69% and the baseline as measured 4.87%.

How often a trial with no treatment effect is declared significant, by how the baseline cut was chosen, baseline correlated 0.7. At ten patients an arm the best-fitting cut rejects a true null 9.31% of the time and the smallest-p cut 16.46%; at two hundred, 5.03% and 14.27%. The median fixed in advance holds 4.78% and 4.56%.
Fig. 4 The share of trials with no treatment effect declared significant at 5%, against patients per arm, for the median cut fixed in advance, the best-fitting cut and the cut with the smallest p-value. Baseline correlated 0.7 with the outcome; ten thousand null trials a point.

The excess depends on the trial’s size, and in the direction a reader would hope: at ten patients an arm the best-fit cut rejects 9.31% of null trials, and at two hundred, 5.03%. As the sample grows, the fit at each candidate cut is estimated better, the search spends less on noise and the chosen cut settles on the middle. A large trial that searches blindly for its threshold loses little either way; the search is a small-trial problem, which is unfortunate, because small trials are where a threshold that “makes the analysis work” is most often sought.

It also depends on how many cuts were tried. Among three candidates — the quartiles — the best-fit cut rejects 5.35% of null trials at fifty an arm; among nine, 5.87%; among nineteen, at every twentieth of the sample, 6.39%. The share it appears to keep rises from 0.678 to 0.724 to 0.746 over the same three searches, while its actual variance moves from 0.725 to 0.736 to 0.739 — more search, more flattery, slightly less precision.

The cut chosen for its p-value

Choosing the cut by the treatment’s p-value is a different order of problem. Among the nine deciles at fifty patients an arm it rejects a true null 14.22% of the time, nearly three times the nominal rate, and its estimates vary by 1.615 of the unadjusted analysis’s — more than half as much again as not adjusting at all — while it reports 0.762. Among three candidate cuts it rejects 9.13%, and among nineteen, 17.22%.

How often a cut point chosen after looking makes two identical arms differ significantly. Sixty-four per arm, no difference between the arms, candidate cuts spread between one standard deviation below and above the mean, and the most significant comparison reported. With one cut fixed in advance the false-positive rate is 4.2%; choosing among five it is 17.7%, and among seventeen 27.4%.
Fig. 5 The same search on the outcome rather than the baseline: two arms with no difference, sixty-four patients in each, and the most significant comparison of proportions among candidate cuts reported. One fixed cut holds the rate near 5%; five chosen afterwards reach 17.7%.

The comparison with the outcome’s cut chosen afterwards is instructive. Searching five thresholds of the outcome takes the error rate to 17.7%; searching nine thresholds of a baseline takes it to 14.22%. The baseline search is cheaper because the treatment estimate barely moves between adjustments: every candidate cut is the same comparison of arms, re-weighted slightly, so the nine estimates are highly correlated and the maximum of nine correlated tests is not much larger than one. The outcome search re-defines the comparison itself at every threshold. Cheaper is not cheap, and a p-value reported after either search is one of several analyses counted as one.

The power it appears to add

At an effect of half a standard deviation, fifty patients an arm and a correlation of 0.7, the four analyses reach significance 93.24% of the time adjusted for the baseline as measured, 84.51% with the fixed median and 85.44% with the best-fit cut. So the chosen cut appears to buy nearly a point of power over the median, and the appearance is exact accounting: the extra point is the same too-small standard error that produced the extra point of false positives. A test whose rejection rate rises equally with and without an effect has not become more powerful, it has become less honest.

The cut chosen for its p-value reaches 94.72%, apparently better than covariance adjustment itself. But its estimates average 0.595, where the true effect is 0.500: the search reports the threshold at which the effect looked largest, so the estimate it prints is inflated by a fifth — the winner’s curse arriving through a choice of covariate rather than a choice of study.

What four adjusted analyses report about their own precision and what they have, 20 patients an arm, baseline correlated 0.7. Reported against actual variance of the treatment estimate, as shares of the unadjusted analysis's, over 20,000 trials of 20 patients an arm. Adjusted for the baseline as measured: 0.530 reported, 0.517 actual. Median cut fixed in advance: 0.722 and 0.700. The decile cut with the best fit: 0.630 and 0.773. The decile cut with the smallest p-value: 0.740 and 1.624.
Fig. 6 The same four analyses at twenty patients an arm. The best-fitting cut’s reported variance and its actual one move further apart, and its actual variance moves further above the fixed median’s.

The small trial is where the cost is largest and most tempting. At twenty patients an arm the best-fit cut reports 0.630 of the unadjusted variance and has 0.773, where the fixed median has 0.700; it is the analysis with the most attractive printed interval and the second-worst real one. A small trial is also the one most likely to be analysed by someone looking for a threshold that “makes sense”, because in a small trial nothing is significant without help.

Why a search that never looked at the arms still costs

The best-fit cut never compares the arms, so it cannot select for a treatment difference, and in a randomised trial its estimate is unbiased: at an effect of 0.5 its estimates average 0.500. What it selects for is a small residual variance, and a residual variance appears twice in the analysis — once in the estimate’s actual precision, through which cut is used, and once in the reported standard error, through the sum of squares that survived the search. The search makes the first worse, because it moves away from the population’s best cut, and the second smaller, because it keeps the smallest of nine. Both effects push the reported interval away from the truth in the same direction.

That is the general shape of choosing a nuisance model from the data: the choice does not bias the parameter of interest, and it does bias every statement about that parameter’s precision. It is the same shape a break that was looked for measured for a structural break found by searching a series — a statistic maximised over candidate locations and then read as though the location had been known — arriving here in the part of the model that nobody reports.

What a protocol can fix

The repair is not to search more cleverly. It is to take the decision before the data exist, which in this case costs nothing.

Adjust for the baseline as measured. Covariance adjustment for a pre-specified continuous baseline is the standard recommendation for adjusted analyses of randomised trials because it cannot be flattered by a search: it has one column and nothing to choose. Its reported variance at fifty patients an arm is 0.529 and its actual one 0.518, and no cut, fixed or chosen, comes near it.

If a category must be used, fix it in the protocol. The median, or tertiles, chosen before unblinding and before anyone has looked at the pooled outcomes, keep a known share of the baseline’s value and report their precision honestly.

A threshold with clinical meaning is still a threshold chosen in advance. A blood-pressure cut at 140 or a score of 10 is not searched; it is worse than the median for precision, as the cut baseline measured, but its reported standard error means what it says. The objection here is only to thresholds found in the trial’s data.

If a search was done, report it. A trial that chose its threshold by fit should say how many candidates it considered, and its interval should be widened by a factor calibrated to that search — by re-randomising the whole procedure, for example, which re-runs the search on every relabelling the design allowed and so charges for it. The shortest honest summary is that the chosen cut’s interval, at fifty patients an arm and nine candidates, is about 4.5% too narrow in width, since the square root of 0.911 is 0.955.

What is shown here and what is not

A cut chosen among the sample’s nine deciles for the best fit reports 0.911 of its actual variance and has more variance than the fixed median, 0.736 against 0.700 of the unadjusted analysis’s, at fifty patients an arm and a baseline correlated 0.7 with the outcome. In its own sample it appears to keep 0.724 of the baseline’s variance reduction, against 0.631 for the median.

It rejects a true null 5.87% of the time at fifty patients an arm and 9.31% at ten, and a cut chosen for the smallest p-value rejects 14.22%.

At a correlation of 0.3 the chosen cut reports a narrower interval than covariance adjustment for the baseline as measured, 0.939 against 0.944 of the unadjusted variance, while having a wider one, 0.976 against 0.925.

Every number is counted over twenty thousand simulated trials of a normal baseline with a linear relation to a normal outcome, except the curves over trial size and correlation, which use ten thousand a point. The fixed median’s share, 2/π, and the covariate’s, 1−ρ21-\rho^2, are exact, and the counted analyses agree with them to within their finite-sample corrections.

Not shown: that the search is harmless when the relation between baseline and outcome is genuinely a step. If the outcome really does jump at a threshold, a search that finds it recovers a real feature and the comparison with the median changes; the cost measured here is the cost of searching for a feature the data do not have. Not shown either for non-normal baselines, where the population’s best single cut is not the median and a search has something real to find.

Still open: which covariate rather than which cut

The search here chooses a threshold for one pre-specified covariate. The more common search chooses the covariates themselves. A trial records a dozen baseline variables, and the adjusted analysis includes those that predict the outcome best in the trial’s own data — by stepwise selection, or by keeping whichever are “significant”. That is the best-fit rule applied to a different menu, and it should produce the same pair of effects: an unbiased treatment estimate whose reported precision is flattered by the search, and an actual precision below that of the covariates a protocol would have named.

What is not measured here is how the two effects scale with the menu. A search among a dozen candidate covariates, of which two genuinely predict, might cost more than a search among nine thresholds of one covariate because the candidates are less correlated with each other, or less, because most of them predict nothing and are rarely kept. Which of those holds, and whether the covariates a search keeps are the ones that were worth adjusting for, is a calculation over the selection rule that has not been made.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Analysis of covarianceBaseline adjustmentData snoopingDichotomisationError rateOptimismSelection effectStatistical power