What a sample-size calculation was given

The spread a pilot supplies

A trial sized for 80% power from a pilot's standard deviation is sized from an estimate that is too small more often than not. With a pilot of ten, 55.9% of the trials it sizes have less than 80% power and 11.1% less than 50%, although the planned sample is right on average. Sizing from the pilot's 80% upper confidence limit instead leaves 19.8% short, at 1.65 times the sample; from its 90% limit, 10.0% short at 2.12 times.

Worth reading first: What a p-value does not say.

How many subjects computes the sixty-four per arm that give 80% power to detect half a standard deviation, and gets the same answer from a closed form and from simulated experiments. Both routes are exact about the question they answer, which takes the effect and the standard deviation as given. That essay then lists where the effect comes from in practice and warns that a pilot’s estimate of it is inflated by selection.

The standard deviation has a quieter problem. Nobody selects pilots for having a small spread, the estimate is unbiased for the variance, and the usual advice is to take it from a pilot and move on. The measurement below is of what that advice delivers.

The power trials actually have when sized for 80% from a pilot of 10Four thousand pilots of 10 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 55.9% of the trials have less than 80% power and 11.1% less than 50%; the median trial has 76.8%.0%20%40%60%80%100%power the trial actually has55.9% fall short of 80%exact t power for each trial's own sizea pilot's spread is too small more often than not
Fig. 1 Four thousand pilots of ten observations, each used to size a trial for 80% power at a true effect of half a standard deviation. The histogram is of the power each trial actually has, computed exactly for the sample size its own pilot chose; the rule marks 80%. The slider is the pilot’s size.

The calculation, with one input estimated

The per-arm sample size for a two-sample comparison is, near enough,

n  =  2(z0.975+z0.80)2σ2Δ2  +  z0.97524n \;=\; \frac{2\,(z_{0.975} + z_{0.80})^2\,\sigma^2}{\Delta^2} \;+\; \frac{z_{0.975}^2}{4}

where Δ\Delta is the effect worth detecting and the second term is the usual allowance for estimating the spread in the final analysis. With σ\sigma known and Δ/σ=0.5\Delta/\sigma = 0.5 it gives 64, which is exactly the smallest sample whose exact power reaches 80%. So a trial planned from a perfect pilot has its planned power, and any shortfall below comes from the pilot and nothing else.

A pilot of n0n_0 observations replaces σ2\sigma^2 with s2s^2, which is distributed as σ2χν2/ν\sigma^2\chi^2_\nu/\nu with ν=n01\nu = n_0 - 1. Every quantity in the calculation then follows from that one distribution: the planned nn is proportional to s2s^2, and the power of the trial is the exact non-central t power at whatever nn the pilot produced, evaluated at the true effect.

Why a shortfall costs more than an excess saves

The second ingredient is visible in the power curve itself, before any pilot is involved.

Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.
Fig. 2 The exact power of a two-sample t test at an effect of half a standard deviation, against the number of observations per arm on a doubling scale. The curve rises steeply below 80% and flattens above it.

At half a standard deviation the exact power is 80.15% at sixty-four per arm. A quarter fewer units, forty-eight, gives 67.88%: twelve points lost. A quarter more, eighty, gives 88.16%: eight points gained. Halving to thirty-two gives 50.36%, thirty points lost; doubling to a hundred and twenty-eight gives 97.86%, eighteen gained.

The curve is concave around the target, so errors in the planned size do not cancel in power even when they cancel in units. A plan that is too small by some factor loses more than a plan too large by the same factor gains, and a pilot that is right about the sample size on average is wrong about the power on average — below it.

Right on average, short more often than not

Over four thousand pilots of ten, the planned trials’ average size is 0.99 of the known-spread 64 — the planning is right on average, as an unbiased variance estimate promises. And:

pilot size trials under 80% power under 50% median power tenth percentile
6 59.1% 21.6% 73.8% 33.8%
10 55.9% 11.1% 76.8% 47.8%
20 52.6% 3.0% 78.9% 59.8%
30 52.2% 0.9% 79.5% 64.0%
50 50.0% 0.1% 79.5% 68.8%

More than half the trials sized from a pilot of ten have less than the power they were designed for, and one in nine has less than half. The planners’ chance of reaching their own target is worse than a coin toss.

Two asymmetries produce this, and both are worth separating.

The first is in the pilot’s estimate. A χ2\chi^2 is skewed to the right, so its median is below its mean, and s2s^2 is below σ2\sigma^2 more often than above it — on 56.3% of pilots of ten, and still on 52.7% of pilots of fifty. An estimate that is unbiased on average is too small in the typical case, and the typical pilot therefore plans a trial that is too small.

The second is in the power curve. Power is concave in nn near 80%: an undersized trial loses more power than an equally oversized trial gains, because the curve flattens as it approaches one. So even a symmetric error in the planned size would average to less than 80%, and the pilot’s error is not symmetric.

The table’s last column is where the practical damage is. One trial in ten planned from a pilot of ten has 47.8% power or less — a study that will usually fail to detect the effect it was built to detect, with nothing about its planning looking wrong.

What an underpowered trial then reports

A trial with less power than planned does not merely fail more often. When it succeeds, it reports an effect that is too large, because a significant result from a low-powered study has been selected for landing high. The one trial in ten that a pilot of ten leaves at under 48% power is, when it does reach significance, a trial whose estimate is on average 1.44 times the true effect — the inflation at that power in the winner’s-curse arithmetic.

So the pilot’s error propagates in two directions at once. The planned trial is too small more often than not, which lowers its chance of finding the effect; and among the trials that find it, the too-small ones report it too large. A literature built from pilot-planned trials inherits both: fewer detections than planned, and larger effects among the detections than the truth. The next trial planned from that literature’s effect is then too small for a different reason, which is the loop how many subjects describes for the effect and this essay adds the spread to.

A bigger pilot helps less than expected

The right-hand columns improve steadily with the pilot’s size: the chance of a disastrously underpowered trial falls from over a fifth at six observations to one in a thousand at fifty. The first column barely moves. At fifty observations exactly half of trials are still short of 80%.

That is the concavity again. With a large pilot the planned sizes cluster tightly around 64, half a little below and half a little above, and the half below have a little less than 80% power. The miss is small — the median trial has 79.5% — but it is a miss half the time, because the target sits exactly at the median of what the pilot delivers. A calculation that aims at 80% from an estimated spread hits 80% only as often as its estimate is at least as large as the truth, and for any size of pilot that is about half the time.

The power trials actually have when sized for 80% from a pilot of 50. Four thousand pilots of 50 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 50.0% of the trials have less than 80% power and 0.1% less than 50%; the median trial has 79.5%.
Fig. 3 The same histogram for pilots of fifty observations. Almost no trial is badly underpowered, and the distribution sits tightly around 80%, with half of it just below.

So the question “how large a pilot is enough” has two answers. For avoiding a badly underpowered trial, a pilot of twenty or thirty is enough. For having the planned power, no pilot is enough, because the plan is aimed at the median of a distribution rather than at a point it can be sure of reaching.

Sizing from a confidence limit

The repair that follows is to plan from a spread the pilot is unlikely to have understated: an upper confidence limit for σ\sigma rather than its estimate. A one-sided 1γ1 - \gamma upper limit is sν/χν,γ2s\sqrt{\nu/\chi^2_{\nu,\gamma}}, and sizing the trial from it is Browne’s recommendation.

Sizing a trial from a pilot of 10: its standard deviation, or an upper confidence limit for it. From the pilot's own standard deviation, 55.9% of trials fall short of 80% power at an average size ×0.99 of the known-spread size. From its 80% upper limit, 19.8% fall short at ×1.65; from its 90% limit, 10.0% at ×2.12.
Fig. 4 For pilots of ten, the share of trials still short of 80% power against the average trial size, as the confidence level of the upper limit used for the spread rises from the plain estimate to 95%.
sized from trials under 80% power average size, as a multiple of 64
the pilot’s standard deviation 55.9% ×0.99
its 50% upper limit 50.2% ×1.07
its 80% upper limit 19.8% ×1.65
its 90% upper limit 10.0% ×2.12

From a pilot of ten, the 80% upper limit leaves one trial in five short of 80% power, and costs 65% more units on average. The level of the limit is close to the probability of reaching the target — 80% limit, 80.2% of trials reach 80% power — which is not a coincidence: the planned power is reached exactly when the planning spread is at least the true one, and a one-sided 80% limit is at least the true spread 80% of the time.

The price falls with the pilot. At twenty observations the 80% limit costs 38% more units for the same one-in-five protection, and at thirty 29%. A larger pilot does not change how often the plan is met — the confidence level sets that — but it shrinks how much extra sample the protection needs, because a larger pilot’s upper limit sits closer to its estimate.

Sizing a trial from a pilot of 30: its standard deviation, or an upper confidence limit for it. From the pilot's own standard deviation, 52.2% of trials fall short of 80% power at an average size ×1.00 of the known-spread size. From its 80% upper limit, 19.4% fall short at ×1.29; from its 90% limit, 9.5% at ×1.46.
Fig. 5 The same trade for pilots of thirty. Each confidence level buys about the same protection as before, and every level costs less extra sample than it did with a pilot of ten.

What the pilot itself could have said

None of this requires a simulation to anticipate. A pilot of ten carries its own statement of how well it knows the spread: the 95% confidence interval for σ\sigma runs from 0.688 to 1.826 times the pilot’s standard deviation.

Squared, those are the range of sample sizes consistent with the pilot: from 0.473 to 3.333 times the size its estimate gives. A trial planned at sixty-four per arm from a pilot of ten is a trial whose honest size is somewhere between about thirty and about two hundred and thirteen, and the pilot said so. At twenty observations the range narrows to 0.578 to 2.133 times, and at thirty to 0.634 to 1.807 times.

Reporting the planned sample size as a single number discards exactly the part of the pilot that the power problem is made of. A plan that said “sixty-four per arm from the pilot’s estimate; between thirty and two hundred and thirteen across its 95% interval” would already have told its readers that 80% was a point on a wide curve rather than a property of the design.

Choosing between the two plans

The two sizing rules answer different questions, and neither is wrong.

Sizing from the estimate gives the right trial size on average across many planned trials. A funder who commissions hundreds of studies and cares about the total number of discoveries per unit spent is served by it: the underpowered trials are balanced, in resources, by the overpowered ones.

Sizing from an upper limit gives each trial a stated probability of having its planned power. An investigator running one trial, for whom a 47%-powered study is a wasted study, is served by it, and should choose the level by how much a failed trial costs against how much sample the protection costs.

The levels between the table’s rows fill in the trade smoothly. From a pilot of ten, the 60% upper limit leaves 40.1% of trials short at 1.21 times the sample, and the 70% limit 30.3% short at 1.39 times; at the other extreme the 95% limit leaves 5.0% short and costs 2.65 times. Each ten points of confidence bought near the middle costs about a fifth more sample, and near the top it costs much more, so the last few points of protection are the expensive ones, and a planner who wants most of the protection for a modest price will find it around the 80% limit.

What neither should do is report the plan from the estimate as “80% power”. The accurate description of a trial sized from a pilot of ten is that its power is uncertain, 76.8% at the median, and below 80% with probability 56%. The effect it is powered to detect is a second uncertain input, and the two compound.

Where a better spread comes from

The measurements point to the obvious improvement, which is to estimate the spread from something bigger than a pilot. Three sources are usually available and are usually larger.

Previous studies of the same outcome. A spread pooled across several published trials of the same measurement carries far more degrees of freedom than any pilot, and its confidence interval is correspondingly narrow. The caution is that published trials were run in selected populations, and a spread from a narrow population understates the spread in a broad one.

Routine data. Registries, audits and baseline measurements in the target population estimate the spread of the outcome where the trial will actually run. They are observational and often noisy in their own ways, but noise in a spread estimated from thousands of records is not the problem measured here.

The trial’s own blinded data, early. Re-estimating the spread from the first stage of the trial itself uses the right population by construction, and the upper-limit repair applies to it exactly as to a pilot.

What should not happen is that a ten-person pilot, run to test procedures, supplies the number the trial’s size depends on. The spread is the one input a pilot of that size is least equipped to estimate, and the p-value the eventual trial produces is only as meaningful as the power it was run with.

The same problem inside a trial

The spread is sometimes estimated from the trial’s own first stage rather than a separate pilot, and the same arithmetic applies there. Choosing n after looking prices one version: re-estimating the sample size from a blinded interim, where the pooled spread includes the effect and the design overshoots. The estimate from an interim of a given size has exactly the χ2\chi^2 skew measured here, so a re-estimation that takes the interim spread at face value is short of its target more often than not for the same reason, and the upper-limit repair transfers directly.

It is also the same object that makes the t interval wider than the normal one. The t correction accounts for the spread being estimated in the analysis. Nothing in the standard sample-size calculation accounts for the spread being estimated in the plan, and the z2/4z^2/4 allowance in the formula is the first of those corrections, not the second.

A related trap waits in allocating on a guess, where a pilot’s estimates of two arms’ spreads decide how units are split between them. There the plug-in rule can lose to equal allocation, for the same reason as here: an estimate that is right on average is used as though it were right, and the decision it feeds is not linear in it.

What the counts establish, and what they do not

A pilot of ten understates the variance more often than not, and the trial it sizes has under 80% power more often than not — 56.3% and 55.9%. The first is an exact χ2\chi^2 probability; the second is counted over four thousand pilots with each trial’s power computed exactly for its own size, so the only sampling error is in which pilots were drawn.

Sizing from the pilot’s 80% upper confidence limit leaves under a quarter of trials short of 80%, at a larger average sample — 19.8% short at 1.65 times.

Power is concave around the target, exactly: a quarter fewer units than sixty-four costs 12.3 points of exact power and a quarter more gains 8.0. A pilot of ten’s own 95% interval for the spread already spans planned sizes from 0.473 to 3.333 times its estimate’s, from the χ2\chi^2 quantiles alone. Neither of those needs the simulation; together they are why its result could have been predicted, and the simulation is the check that the two combine the way the argument says.

What does not survive is a trial sized from a pilot’s standard deviation read as having 80% power. From a pilot of ten, 55.9% of such trials fall short.

Not claimed: anything about pilots whose outcome is not normal. A skewed or heavy-tailed outcome makes the sample variance more variable than the χ2\chi^2 describes, which widens every distribution in the tables, and the upper limit computed from χ2\chi^2 then covers less than its level. Nor does this essay treat the effect as uncertain; every trial here is sized for the true effect of half a standard deviation, which is the favourable case.

Still open: the effect is a guess too

Every number above assumes the effect worth detecting is known exactly. It never is. A sample size computed at “half a standard deviation” is computed at a point, and the probability that a trial of that size succeeds depends on where the true effect really is — which the planners do not know and usually have a range of beliefs about.

Averaging the power over that range gives a quantity called assurance, and it behaves in a way power does not: it has a ceiling that no sample size can pass, set by how sure the planners are that the effect exists at all. How far below power it falls at the conventional sample size, how many units an 80% chance of success really takes, and when no number of units is enough, are measured in the chance a trial succeeds.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Chi squaredEstimated varianceThe non-central tPilot studySample sizeStandard deviationStatistical powerUpper confidence limit