The cut that fitted best
Worth reading first: What a p-value does not say.
A baseline cut in two priced a covariate that a randomised trial adjusts for after cutting it at a quantile chosen in advance. Cut at its median, a normal baseline keeps exactly 2/π of the variance it would remove as measured, and at a correlation of 0.7 with the outcome that costs 35% more patients. Every cut there was fixed before the data arrived. The essay ended on the cut that is not: a threshold found in the trial’s own data, the value at which the baseline best separates the outcomes. Such a cut should appear to keep more than 2/π, because it was chosen to. Whether it actually does, and whether the precision it reports is precision it has, was left uncomputed.
It appears to keep more. It has less. And it says so to nobody, because the only numbers a trial reports are the ones computed from the sample that chose the cut.
Two ways to let the data choose a threshold
A cut chosen from the data can be chosen in two ways that feel very different and are worth keeping apart.
The first looks innocent. The analyst never compares the arms while choosing. Among a set of candidate thresholds — here the nine deciles of the pooled baseline — the chosen one is the cut that leaves the smallest residual variance in the adjusted model, which is the same as the cut at which “high” and “low” baseline patients differ most in their outcomes. It is the natural way to find a clinically meaningful break — the same move that sets a diagnostic threshold where the data separate best — it can be done before unblinding, and it is the choice the essay on the cut baseline named as the obvious next step: the cut that maximises the adjusted precision. Here it is called the best-fit cut.
The second is the one nobody defends. The analyst runs the adjusted analysis at every candidate cut and reports the one with the smallest p-value for the treatment. That is the forking path in its plainest form, and the interest here is only in how much worse it is than the first, since the first is what honest analysts actually do.
Both are compared with two analyses that choose nothing: the median cut, written into the protocol, and adjustment for the baseline as measured. All four are fitted by least squares with the arm and the covariate as predictors, and each reports the standard error that least squares gives it — which is exactly the problem, because least squares does not know which of its columns was chosen by looking.
What the chosen cut reports and what it has
At fifty patients an arm and a correlation of 0.7, the best-fit cut reports a variance for the treatment estimate of 0.671 of the unadjusted analysis’s. Across twenty thousand trials its estimates actually vary by 0.736. It claims 10% more precision than it has — its reported variance is 0.911 of its actual — where the median fixed in advance reports 0.716 and has 0.700, the small excess that any estimated standard error carries.
The surprising half is the second number. The best-fit cut’s actual variance, 0.736, is larger than the fixed median’s, 0.700. The cut chosen because it fitted best delivers a less precise treatment estimate than the cut chosen because it was in the protocol. It reports better precision than the median and has worse, and the reversal is not small: on the reported numbers the chosen cut looks 6% more precise than the median, and on the actual ones it is 5% less.
The baseline as measured is untouched by any of this. It reports 0.529 and has 0.518, which is plus the price of estimating one coefficient in a hundred patients. Every cut, fixed or chosen, is compared against that number, and the chosen one is further from it than the fixed one.
Why choosing among nine cuts is worse than using one
The reason the chosen cut does worse is that in the population there is nothing to choose. For a normal baseline and a linear relation with the outcome, the cut that removes the most variance is the median, and it removes 2/π of what the covariate as measured would. Any other cut removes less — at the top decile, about a third as much. So every departure from the median is a loss, and the best-fit rule departs from it most of the time.
It departs because the residual variance at each candidate cut is itself an estimate. With a hundred patients, the difference in fit between the median and the fourth or sixth decile is smaller than the noise in either, so which one wins is decided mostly by where the sample’s outcomes happened to fall. The median is the winner on 21.8% of trials. The five middle deciles take 88.1% between them, and the two outermost deciles 1.5%. The cut wanders, and each trial’s analysis is then run at a threshold that is worse in the population than the one it could have fixed in advance.
The wandering is also what makes the reported variance too small. The chosen cut is the one whose residuals happened to be smallest, and the smallest of nine correlated residual variances is below the typical one. The standard error is built from that residual variance, so it inherits the flattery. Nothing in the fitted model records that nine models were tried; it reports three estimated coefficients and charges three degrees of freedom for them, when the search spent more. is not a measure of fit measured the same optimism in a different setting — useless predictors raising by construction — and a searched threshold is a useless predictor chosen to look useful.
The stronger the baseline, the less the cut wanders, because the difference in fit between the median and its neighbours grows with the correlation while the noise does not. At a correlation of 0.9 the middle three deciles take 79.0% of trials between them; at 0.3 the choice is spread across all nine, the median taking 12.8% and each outermost decile 8.0%, so a weak baseline is cut at the tenth or the ninetieth percentile about one trial in six.
The share the sample says it kept
The flattery can be put in the unit the essay on the cut baseline used. In each trial, the share of the covariate’s variance reduction a cut keeps can be measured from the trial’s own residual variances: how much the cut model removed, as a share of how much the model with the baseline as measured removed. For the fixed median this reads 0.631 on average at fifty patients an arm, a little under 2/π because the sample is finite. For the best-fit cut it reads 0.724.
So a trial that searched for its threshold would conclude that its cut kept nearly three quarters of the baseline’s value, and would be wrong about its cut in the direction that matters: the cut kept less than a median would have.
At weaker correlations the flattery grows, because the true reduction the cut can offer is smaller and the noise it fits is not. At a correlation of 0.3 the chosen cut appears to keep 1.051 of what the baseline as measured removes — more than the covariate itself. That answers the question the cut baseline left in its sharpest form. At fifty patients an arm and a correlation of 0.3, the best-fit cut reports a variance of 0.939 of the unadjusted analysis’s, and adjustment for the baseline as measured reports 0.944: the chosen cut prints a narrower interval than honest covariance adjustment. Its estimates actually vary by 0.976, where covariance adjustment’s vary by 0.925.
In a small trial the numbers turn from embarrassing to perverse. At twenty patients an arm and a correlation of 0.3 the best-fit cut reports 0.914 and has 1.010 — its estimates are more variable than the unadjusted comparison’s, so the adjustment added noise, while its printed interval is the narrowest of the four analyses.
The rejections a true null collects
A standard error that is too small produces rejections a true null should not produce. At fifty patients an arm and a correlation of 0.7 the best-fit cut rejects a true null 5.87% of the time at a nominal 5%, where the fixed median rejects 4.69% and the baseline as measured 4.87%.
The excess depends on the trial’s size, and in the direction a reader would hope: at ten patients an arm the best-fit cut rejects 9.31% of null trials, and at two hundred, 5.03%. As the sample grows, the fit at each candidate cut is estimated better, the search spends less on noise and the chosen cut settles on the middle. A large trial that searches blindly for its threshold loses little either way; the search is a small-trial problem, which is unfortunate, because small trials are where a threshold that “makes the analysis work” is most often sought.
It also depends on how many cuts were tried. Among three candidates — the quartiles — the best-fit cut rejects 5.35% of null trials at fifty an arm; among nine, 5.87%; among nineteen, at every twentieth of the sample, 6.39%. The share it appears to keep rises from 0.678 to 0.724 to 0.746 over the same three searches, while its actual variance moves from 0.725 to 0.736 to 0.739 — more search, more flattery, slightly less precision.
The cut chosen for its p-value
Choosing the cut by the treatment’s p-value is a different order of problem. Among the nine deciles at fifty patients an arm it rejects a true null 14.22% of the time, nearly three times the nominal rate, and its estimates vary by 1.615 of the unadjusted analysis’s — more than half as much again as not adjusting at all — while it reports 0.762. Among three candidate cuts it rejects 9.13%, and among nineteen, 17.22%.
The comparison with the outcome’s cut chosen afterwards is instructive. Searching five thresholds of the outcome takes the error rate to 17.7%; searching nine thresholds of a baseline takes it to 14.22%. The baseline search is cheaper because the treatment estimate barely moves between adjustments: every candidate cut is the same comparison of arms, re-weighted slightly, so the nine estimates are highly correlated and the maximum of nine correlated tests is not much larger than one. The outcome search re-defines the comparison itself at every threshold. Cheaper is not cheap, and a p-value reported after either search is one of several analyses counted as one.
The power it appears to add
At an effect of half a standard deviation, fifty patients an arm and a correlation of 0.7, the four analyses reach significance 93.24% of the time adjusted for the baseline as measured, 84.51% with the fixed median and 85.44% with the best-fit cut. So the chosen cut appears to buy nearly a point of power over the median, and the appearance is exact accounting: the extra point is the same too-small standard error that produced the extra point of false positives. A test whose rejection rate rises equally with and without an effect has not become more powerful, it has become less honest.
The cut chosen for its p-value reaches 94.72%, apparently better than covariance adjustment itself. But its estimates average 0.595, where the true effect is 0.500: the search reports the threshold at which the effect looked largest, so the estimate it prints is inflated by a fifth — the winner’s curse arriving through a choice of covariate rather than a choice of study.
The small trial is where the cost is largest and most tempting. At twenty patients an arm the best-fit cut reports 0.630 of the unadjusted variance and has 0.773, where the fixed median has 0.700; it is the analysis with the most attractive printed interval and the second-worst real one. A small trial is also the one most likely to be analysed by someone looking for a threshold that “makes sense”, because in a small trial nothing is significant without help.
Why a search that never looked at the arms still costs
The best-fit cut never compares the arms, so it cannot select for a treatment difference, and in a randomised trial its estimate is unbiased: at an effect of 0.5 its estimates average 0.500. What it selects for is a small residual variance, and a residual variance appears twice in the analysis — once in the estimate’s actual precision, through which cut is used, and once in the reported standard error, through the sum of squares that survived the search. The search makes the first worse, because it moves away from the population’s best cut, and the second smaller, because it keeps the smallest of nine. Both effects push the reported interval away from the truth in the same direction.
That is the general shape of choosing a nuisance model from the data: the choice does not bias the parameter of interest, and it does bias every statement about that parameter’s precision. It is the same shape a break that was looked for measured for a structural break found by searching a series — a statistic maximised over candidate locations and then read as though the location had been known — arriving here in the part of the model that nobody reports.
What a protocol can fix
The repair is not to search more cleverly. It is to take the decision before the data exist, which in this case costs nothing.
Adjust for the baseline as measured. Covariance adjustment for a pre-specified continuous baseline is the standard recommendation for adjusted analyses of randomised trials because it cannot be flattered by a search: it has one column and nothing to choose. Its reported variance at fifty patients an arm is 0.529 and its actual one 0.518, and no cut, fixed or chosen, comes near it.
If a category must be used, fix it in the protocol. The median, or tertiles, chosen before unblinding and before anyone has looked at the pooled outcomes, keep a known share of the baseline’s value and report their precision honestly.
A threshold with clinical meaning is still a threshold chosen in advance. A blood-pressure cut at 140 or a score of 10 is not searched; it is worse than the median for precision, as the cut baseline measured, but its reported standard error means what it says. The objection here is only to thresholds found in the trial’s data.
If a search was done, report it. A trial that chose its threshold by fit should say how many candidates it considered, and its interval should be widened by a factor calibrated to that search — by re-randomising the whole procedure, for example, which re-runs the search on every relabelling the design allowed and so charges for it. The shortest honest summary is that the chosen cut’s interval, at fifty patients an arm and nine candidates, is about 4.5% too narrow in width, since the square root of 0.911 is 0.955.
What is shown here and what is not
A cut chosen among the sample’s nine deciles for the best fit reports 0.911 of its actual variance and has more variance than the fixed median, 0.736 against 0.700 of the unadjusted analysis’s, at fifty patients an arm and a baseline correlated 0.7 with the outcome. In its own sample it appears to keep 0.724 of the baseline’s variance reduction, against 0.631 for the median.
It rejects a true null 5.87% of the time at fifty patients an arm and 9.31% at ten, and a cut chosen for the smallest p-value rejects 14.22%.
At a correlation of 0.3 the chosen cut reports a narrower interval than covariance adjustment for the baseline as measured, 0.939 against 0.944 of the unadjusted variance, while having a wider one, 0.976 against 0.925.
Every number is counted over twenty thousand simulated trials of a normal baseline with a linear relation to a normal outcome, except the curves over trial size and correlation, which use ten thousand a point. The fixed median’s share, 2/π, and the covariate’s, , are exact, and the counted analyses agree with them to within their finite-sample corrections.
Not shown: that the search is harmless when the relation between baseline and outcome is genuinely a step. If the outcome really does jump at a threshold, a search that finds it recovers a real feature and the comparison with the median changes; the cost measured here is the cost of searching for a feature the data do not have. Not shown either for non-normal baselines, where the population’s best single cut is not the median and a search has something real to find.
Still open: which covariate rather than which cut
The search here chooses a threshold for one pre-specified covariate. The more common search chooses the covariates themselves. A trial records a dozen baseline variables, and the adjusted analysis includes those that predict the outcome best in the trial’s own data — by stepwise selection, or by keeping whichever are “significant”. That is the best-fit rule applied to a different menu, and it should produce the same pair of effects: an unbiased treatment estimate whose reported precision is flattered by the search, and an actual precision below that of the covariates a protocol would have named.
What is not measured here is how the two effects scale with the menu. A search among a dozen candidate covariates, of which two genuinely predict, might cost more than a search among nine thresholds of one covariate because the candidates are less correlated with each other, or less, because most of them predict nothing and are rarely kept. Which of those holds, and whether the covariates a search keeps are the ones that were worth adjusting for, is a calculation over the selection rule that has not been made.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A window for every candidate — both name data snooping, optimism, selection effect
- Choosing whether to break — both name error rate, optimism, selection effect
- Eight forecasters and one benchmark — both name data snooping, error rate, selection effect
- The models that were never in the running — both name data snooping, error rate, statistical power
- A boundary for giving up — both name error rate, statistical power
- A charge that reads the draw — both name optimism, selection effect
Named objects
A flat tag is an object no other essay names yet.
Analysis of covarianceBaseline adjustmentData snoopingDichotomisationError rateOptimismSelection effectStatistical power