What a sample-size calculation was given

An outcome cut in two

Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.

Worth reading first: What a p-value does not say.

The spread a pilot supplies and the chance a trial succeeds took the two inputs of a sample-size calculation — the standard deviation and the effect — and priced the uncertainty in each. Both assumed the outcome itself was fixed. It rarely is. A blood pressure becomes “controlled or not”, a depression score becomes “responder or not”, a time becomes “within target or not”, and the choice is usually made for reasons of interpretation rather than of power.

It is also a choice about power, and a large one.

How much of a normal outcome's information survives cutting it into two, by where the cut is. For a small shift, a cut at the mean keeps 63.7% of the information, so the trial needs 1.57 times the sample. A cut at the top tenth keeps 34.2% and needs 2.92 times; a cut two standard deviations out keeps 13.1%.
Fig. 1 The share of a normal outcome’s information that survives replacing each measurement with whether it lies above a threshold, against where the threshold is placed, in standard deviations from the mean. Three conventional cuts are marked.

The information a cut keeps

A measured outcome carries its position on a scale; a cut outcome carries one bit, above or below. For detecting a small shift in a normal outcome, the efficiency of the cut relative to the measurement has a closed form. Cutting at cc standard deviations from the mean keeps

φ(c)2Φ(c)(1Φ(c))\frac{\varphi(c)^2}{\Phi(c)\,\bigl(1 - \Phi(c)\bigr)}

of the information, where φ\varphi and Φ\Phi are the normal density and distribution function.

At the mean that is 2/π2/\pi, 63.7%. It is the most any single cut can keep, and it already throws away more than a third of the data. Away from the mean it falls quickly:

where the cut is information kept sample needed, as a multiple
at the mean 63.7% 1.57
a quarter standard deviation out 62.2% 1.61
one standard deviation out 43.9% 2.28
at the top tenth 34.2% 2.92
one and a half standard deviations out 26.9% 3.72
two standard deviations out 13.1% 7.63

A responder threshold set at the top tenth of the control distribution needs almost three times the patients the measured outcome needs, and one set two standard deviations out needs more than seven times. Thresholds of clinical interest are usually far from the middle of the distribution, because they are defined by what matters rather than by where the data sit, and those are exactly the thresholds that cost most.

In patients

The efficiency is a small-shift limit. For a trial designed to detect half a standard deviation, the powers can be computed directly for the measured outcome — a two-sample comparison of means — and for the cut one, a comparison of two proportions.

Power for the outcome as measured and cut at the control meanThe outcome as measured reaches 80% power at 63 per arm; cut at the control mean, at 102, 1.62 times as many.00.2500.5000.7501100200300400observations per armpower at half a standard deviationas measuredcut into twotwo-sample z against two-proportion z1.62 times the sample
Fig. 2 Power to detect half a standard deviation, against observations per arm, for the outcome as measured and cut at the control mean. The vertical rules mark where each reaches 80%. The slider moves the cut.
cut, above the control mean per arm for 80% power multiple of the measured outcome’s 63
none — the outcome as measured 63 1.00
at the mean 102 1.62
one standard deviation 124 1.97
the top tenth 152 2.41
one and a half standard deviations 185 2.94
two standard deviations 345 5.48

The multiples at a finite effect are close to the small-shift limits and not identical to them: half a standard deviation is not small, and a cut above the control mean sits nearer the middle of the pooled distribution, which is kinder to it. At sixty-four per arm, the trial that has 80% power on the measured outcome has 59.9% power cut at the mean and 37.4% cut at one and a half standard deviations. The closed forms agree with counted trials: at the one-and-a-half cut, twenty thousand simulated pairs of arms give 38.6%.

Power for the outcome as measured and cut 1.5 standard deviations above it. The outcome as measured reaches 80% power at 63 per arm; cut 1.5 standard deviations above it, at 185, 2.94 times as many.
Fig. 3 The same comparison with the cut one and a half standard deviations above the control mean. The cut outcome needs nearly three times the sample to reach 80%.

Why the loss is so large

The measured outcome’s comparison uses every patient’s distance from every other. The cut outcome uses only whether each patient is on one side of a line, so a patient just above the threshold and a patient far above it count the same, and so do a patient just below and a patient far below. Near the threshold that loses little; everywhere else it loses almost everything the patient’s value said.

A cut at the mean keeps the most because it maximises the share of patients whose side of the line is genuinely uncertain — half of them are near it. A cut far out puts most patients unambiguously below it, so they all contribute the same zero, and only the few near the threshold carry information about the shift. That is why the curve falls so steeply: two standard deviations out, nineteen patients in twenty are below the line in either arm, and the comparison rests on the remaining one.

How many subjects lists the ways to raise power other than by recruiting more patients, and one of them is to measure the outcome better. Keeping the outcome as measured is the cheapest instance of that: nothing extra is collected, and the analysis simply declines to throw away what was.

What a responder rate says about patients

The cut changes more than the power. It changes what the result seems to say.

Every patient shifted by half a standard deviation, read as responders above 1 standard deviation. The treated arm is the control arm moved half a standard deviation, for every patient. Above the threshold lie 15.9% of control patients and 30.9% of treated ones — a difference of 15.0% that reads as a share of patients who respond, when every patient responded by the same amount.
Fig. 4 Two normal outcomes, the treated arm the control arm moved half a standard deviation for every patient, with the share of each arm above a responder threshold shaded.

Suppose the treatment moves every patient by exactly half a standard deviation — no patient more, none less. Cut at one standard deviation above the control mean, the “responder” rates are 15.9% in the control arm and 30.9% in the treated arm. The natural reading of that difference is that the treatment makes an extra 15.0% of patients respond, and that the other 85% get nothing from it.

Both halves of the reading are false in the example. Every patient benefited, equally; the 15.0% is the share whose uniform improvement happened to carry them across a line. Moved to the mean the same shift gives 50.0% against 69.1%, a difference of 19.1%; moved to one and a half standard deviations, 6.7% against 15.9%, a difference of 9.2%. The same treatment effect is a 19, 15 or 9 point difference in responders depending only on where the line was drawn, and none of those numbers is the share of patients who benefit.

A responder analysis can be the right analysis when the threshold has a meaning independent of the trial — a value above which a complication occurs, a level a regulator acts on. What it cannot do is reveal heterogeneity in who responds. A difference in the share above a line is equally consistent with every patient improving a little and with a few improving a lot, and separating those needs the measured outcome — which the cut has discarded. The same confusion between a shift and a subgroup is at the heart of reading one significant result beside another.

The relative measures move with the line too

A difference in responder rates is the absolute measure. Trials and meta-analyses more often report a relative one, the risk ratio or the odds ratio, and it is tempting to think those are properties of the treatment rather than of the threshold. For the same uniform shift of half a standard deviation:

threshold above the control mean difference in responders risk ratio odds ratio
one standard deviation below 9.2 points 1.11 2.63
at the mean 19.1 points 1.38 2.24
one standard deviation above 15.0 points 1.94 2.37
one and a half above 9.2 points 2.37 2.63
two above 4.4 points 2.94 3.08

The risk ratio runs from 1.11 to 2.94 and the odds ratio from 2.24 to 3.08, for a treatment that did one thing to every patient. The difference in responders peaks near the middle, the risk ratio climbs steadily as the line moves out, and the odds ratio is smallest near the middle and grows in both directions. Each summary picks up a different feature of where two normal curves cross a line, and none of them is constant, because none of them is the scale the effect actually acts on.

That matters most when results are combined. Two trials of one treatment that define response at different thresholds report different risk ratios with no difference in the treatment, and a meta-analysis pooling them sees heterogeneity that the choice of threshold created. The forest-plot reading of two intervals would then mistake a difference in definitions for a difference in effects.

Every patient shifted by half a standard deviation, read as responders above 0 standard deviations. The treated arm is the control arm moved half a standard deviation, for every patient. Above the threshold lie 50.0% of control patients and 69.1% of treated ones — a difference of 19.1% that reads as a share of patients who respond, when every patient responded by the same amount.
Fig. 5 The same half-standard-deviation shift read at the control mean. The difference in responder rates is at its largest here, and the risk ratio near its smallest.

The same shift at the mean gives the largest absolute difference, 19.1 points, and a risk ratio of only 1.38 — the reading that looks most impressive in one summary looks least impressive in another, with nothing changed but the summary.

A threshold that sits near the data

The costs above are for thresholds placed by their meaning, which usually puts them in a tail. A threshold that happens to sit near the middle of the distribution being studied costs much less — but never less than a third.

The information kept is largest when the cut divides the patients roughly in half, and 2/π2/\pi is its ceiling: no placement of a single line recovers more than 63.7% of a normal outcome’s information. A responder threshold that splits a trial’s population evenly is therefore the best case, and even it needs 57% more patients than the measured outcome. When the population a trial recruits is itself selected near a threshold — patients enrolled because they are just above a diagnostic cut-off, say — the cut can sit near the middle of their distribution by construction, and the loss approaches that floor.

The reverse is also worth seeing. Because the loss depends on where the line falls within the trial’s distribution, the same clinical threshold costs different amounts in different trials. A cut-off that divides a severe population evenly is a tail cut in a mild one, and a sample size borrowed from a trial in the first population will leave a trial in the second badly underpowered.

A cut chosen after looking

A cut that is fixed in the protocol costs power. A cut chosen after seeing the data costs something worse.

How often a cut point chosen after looking makes two identical arms differ significantly. Sixty-four per arm, no difference between the arms, candidate cuts spread between one standard deviation below and above the mean, and the most significant comparison reported. With one cut fixed in advance the false-positive rate is 4.2%; choosing among five it is 17.7%, and among seventeen 27.4%.
Fig. 6 Two arms with no difference at all, sixty-four patients in each. The share of trials in which the most significant comparison of proportions, among several candidate cuts between one standard deviation below the mean and one above, reaches p < 0.05.

If the threshold is chosen among several candidates — the median, a round number, a published cut-off, the value that “looked like a natural break” — and the one giving the clearest difference is reported, the test is one of a family:

cuts considered false-positive rate with no true difference
one, fixed in advance 4.23%
two 9.33%
three 12.39%
five 17.72%
nine 22.92%
seventeen 27.40%

Choosing the best of five cuts turns a 5% test into an 18% one. The comparisons at neighbouring cuts share most of their patients, so the rate grows more slowly than five separate tests would make it, and it keeps growing: seventeen closely spaced candidates give 27%. The fixed cut reads 4.23% rather than 5% because a comparison of counts is discrete and slightly conservative, the same discreteness that a coin in the interval removes.

This is the garden of forking paths in its most innocent form. Nobody runs seventeen analyses on purpose; a threshold is simply picked after the distribution has been seen, and the distribution that was seen is the one whose chance fluctuations the threshold now fits.

What the lost power does next

A trial analysed on a cut outcome at its original size is an underpowered trial, and an underpowered trial does more than miss effects. When it finds them it overstates them, because a significant result from a low-powered comparison has been selected for landing high.

At sixty-four per arm with the outcome cut one and a half standard deviations out, the comparison has 37.4% power. Among the trials that reach significance, the estimated difference in responder rates is inflated by the same arithmetic that inflates any significant estimate at that power — to 1.62 times the true difference on average — and the risk ratio with it. So the cut costs twice: once in the trials that fail, which are the majority, and once in the trials that succeed, whose reported effect is larger than the treatment’s.

In patients, the cost of keeping the power is concrete. A trial designed at 63 per arm on the measured outcome needs 124 per arm with a responder threshold one standard deviation out: 122 extra patients enrolled, followed and exposed to an experimental treatment, to answer the same question with the same power after discarding most of what each of them contributed. That is a cost an ethics committee would weigh if it were stated, and it is rarely stated.

The share above a line, without cutting

A responder rate is sometimes exactly the quantity a decision needs, and the measurements do not argue against reporting it. They argue against estimating it by cutting the data.

For a normal outcome, the share of an arm above a threshold cc can be read off the fitted distribution — one minus the normal distribution function at cc, using the arm’s own mean and standard deviation — rather than counted. Both estimates are unbiased for large samples, and their variances differ by a closed-form factor: the fitted estimate’s variance is φ(c)2(1+c2/2)\varphi(c)^2(1 + c^2/2) against the counted proportion’s Φ(c)(1Φ(c))\Phi(c)(1 - \Phi(c)), per observation.

At the mean the fitted share has 0.637 of the counted share’s variance; one standard deviation out, 0.658; one and a half out, 0.572; two out, 0.393. So the share above a threshold two standard deviations out is estimated from the measured outcome with the precision the counted proportion would need two and a half times the patients to reach. The responder rate, the risk ratio and the odds ratio can all be reported from the fitted model, with their uncertainty, at no cost in patients, and a protocol can name them as secondary summaries of the primary analysis rather than as analyses of their own.

The one caution is the one that runs through every figure here: the fitted estimate relies on the outcome being normal in the region of the threshold. Far in a tail that is an assumption about exactly the part of the distribution the data say least about, and the tail converges last is the reason to check it rather than trust it. A threshold near the middle is safe; a threshold three standard deviations out is a statement about the model.

What to measure and what to report

The measurements point in one direction, and they leave room for the reasons a cut is chosen.

Analyse the outcome as measured. It needs fewer patients, it uses every patient’s value, and it estimates the size of the effect rather than a share that depends on an arbitrary line.

If a threshold matters, report it as well, from the same analysis. A fitted distribution for each arm gives the share above any threshold, with its uncertainty, without the loss of cutting the data first. The responder rates in the example above are exactly those the two fitted normals imply.

If a cut must be the primary analysis, fix it before the data and size the trial for it. At a threshold one standard deviation out that means planning for roughly twice the patients, which is a cost worth stating in the protocol rather than discovering in the result.

What the formulas establish, and what they do not

A median split keeps exactly 2/π2/\pi of the information for a small shift, and the closed-form powers of the continuous and split analyses match a count at sixty-four per arm. A cut further from the mean keeps less, monotonically, out to two and a half standard deviations.

The cut outcome is never more powerful than the outcome as measured, at any sample size or threshold drawn.

One cut fixed in advance holds its level, and every extra candidate cut raises the false-positive rate, counted over twenty thousand null trials at each number of candidates.

The responder measures move with the threshold for a single uniform shift — the risk ratio from 1.11 to 2.94 and the odds ratio from 2.24 to 3.08 across thresholds from one standard deviation below the control mean to two above — which is closed-form arithmetic on two normal tails and needs no simulation. The difference in responder rates is the shifted normal’s extra mass above the threshold, exactly, at every threshold drawn.

What does not survive is a responder rate difference read as the share of patients who benefit. With every patient improved by the same half standard deviation, it is 19.1, 15.0 or 9.2 points depending only on the threshold.

Not claimed: anything for outcomes that are not normal. A heavily skewed outcome, or one with a floor, changes the efficiency of a cut — sometimes in the cut’s favour, when a few extreme values would otherwise dominate a comparison of means — and the right comparison there is with an analysis suited to the skew, not with the t test. The candidate cuts in the forking table are evenly spaced; cuts chosen from a wider or a lumpier range would give different, and not smaller, rates.

Still open: the cut applied to a baseline

Everything here cuts the outcome. Trials also cut baseline measurements — to define subgroups, to stratify, to adjust — and a cut covariate loses information in the same way, which weakens the adjustment it was meant to provide — and a cut that is not a quantile shows that how the cut is defined changes what a balancing rule guarantees about it. How much of the precision gained by adjusting for a baseline measurement survives adjusting for whether it was above its median instead, and whether that loss is the same 36% this essay found for an outcome, is a calculation of the same kind that has not been made.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

DichotomisationEffect sizeThe garden of forking pathsInformationResponder analysisSample sizeStatistical powerTwo-sample test