What a sample-size calculation was given

The chance a trial succeeds

A trial of sixty-four per arm has 80% power at an effect of half a standard deviation. If the effect is only believed to be about half a standard deviation, give or take a quarter, the chance the trial reaches significance is 69.2%; give or take a half, 61.4%. Reaching 80% then takes 113 per arm, or 1,268 — and when the belief is uncertain by three quarters of a standard deviation no number of patients reaches 80%, because the chance can never exceed the 74.8% probability that the effect is positive at all.

Worth reading first: What a p-value does not say.

The spread a pilot supplies took one input to a sample-size calculation — the standard deviation — and showed that estimating it leaves the planned trial short of its power more often than not. The other input is the effect. How many subjects is clear that it is chosen rather than known: from a pilot, from what would be worth acting on, or backwards from a budget.

Whichever way it is chosen, the calculation then treats it as a fact. Power is a function of the true effect, and it is evaluated at one value. The question a funder, an ethics committee or a patient actually asks is different: given what is believed about the effect, how likely is this trial to succeed?

The chance a trial succeeds against its size, when the expected effect of 0.5 is uncertain by four amountsWith the effect known, 80% is reached at 63 per arm. With the effect uncertain by 0.25 standard deviations it takes 113; by 0.5, 1268; by 0.75, no sample size at all, because the chance can never exceed the 74.8% prior probability that the effect is positive.101001,00010,00000.2500.5000.7501observations per armchance of a significant resulteffect knownuncertain by 0.25uncertain by 0.5uncertain by 0.75normal prior on the effect, closed formpower at a point is not the chance of success
Fig. 1 The chance a trial reaches significance, against its size per arm on a logarithmic axis, when the effect is expected to be half a standard deviation and is uncertain by 0, 0.25, 0.5 or 0.75 standard deviations. The horizontal rule is 80%; the short dashed lines are the ceilings the uncertain curves approach. The slider changes the expected effect.

Power averaged over the effect

Suppose the effect, in units of the outcome’s standard deviation, is believed to be around δ0\delta_0 with an uncertainty that can be described as a normal distribution of spread s0s_0. A trial of nn per arm produces a test statistic whose mean is δn/2\delta\sqrt{n/2} for a given true effect δ\delta. Averaged over the belief about δ\delta, the statistic is normal with mean δ0n/2\delta_0\sqrt{n/2} and variance 1+ns02/21 + n s_0^2/2 — the trial’s own noise plus the planners’ uncertainty, scaled by the sample. The probability that it clears the significance threshold is

assurance  =  Φ ⁣(δ0n/2    1.961+ns02/2)\text{assurance} \;=\; \Phi\!\left(\frac{\delta_0\sqrt{n/2} \;-\; 1.96}{\sqrt{1 + n\,s_0^2/2}}\right)

With s0=0s_0 = 0 this is the ordinary power at δ0\delta_0. With any uncertainty, the denominator grows with nn, and it changes the behaviour of the whole curve.

The formula is an average of power over a belief, and it can be checked as one. Drawing two hundred thousand effects from a normal belief centred at 0.5 with spread 0.25 and running a test at sixty-four per arm on each gives a success rate of 69.29%; the closed form gives 69.20%.

At the conventional sample

Sixty-four per arm is the trial with 80% power at half a standard deviation. Its chance of success, as the belief about the effect widens:

uncertainty about the effect chance of success at sixty-four per arm
none 80.7%
0.1 standard deviations 77.5%
0.25 69.2%
0.5 61.4%
0.75 57.9%

(The first row is the normal approximation’s power at sixty-four; the exact t power is 80.15%.)

A trial designed for 80% power is a trial with about a two-in-three chance of success when the planners are honestly unsure of the effect by a quarter of a standard deviation. That is not a pessimistic uncertainty. An effect “of about half a standard deviation” believed to lie, two times in three, between a quarter and three quarters of one is a modest claim about most treatments before they are tested.

The source of the shortfall is the same concavity that shortened pilot-planned trials. Power rises steeply below the design effect and flattens above it, so averaging over effects on both sides pulls the average below the value at the centre. Uncertainty about the effect costs success in a way that no symmetry of the belief can undo.

Why averaging over the effect lowers the chance

The shortfall has a picture, and it is the ordinary power curve read sideways.

Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.
Fig. 2 Exact power at an effect of half a standard deviation, against the observations per arm. At the conventional sixty-four per arm the curve has reached 80% and is bending over.

At sixty-four per arm the exact power is 28.93% if the true effect is a quarter of a standard deviation, 80.15% at half, and 98.78% at three quarters. A belief that puts equal weight on a quarter and three quarters has a centre of one half, and its chance of success is the average of the two ends: 63.9%, sixteen points below the power at the centre.

That is the entire mechanism in three numbers. A trial sized to reach 80% at the expected effect sits on the steep part of the power curve for smaller effects and the flat part for larger ones, so an effect a little smaller than expected costs far more than an effect a little larger returns. The more spread out the belief, the more weight falls on the steep side, and the further the average falls below the power at the centre.

The same shape made the pilot-planned trial fall short in the spread a pilot supplies. There the uncertain input moved the sample size along the curve; here it moves the effect. Either way, a calculation that evaluates a concave curve at its expected input overstates the expected output. The gap is not a rounding matter: at the settings above it is the difference between a trial that meets its target four times in five and one that meets it two times in three, and it is invisible in any calculation that reports a single number.

The sample an 80% chance takes

If the goal is an 80% chance of success rather than 80% power at one effect, the formula can be solved for nn.

The sample a trial needs for an 80% chance of success, as the effect it expects becomes uncertain. For an expected effect of half a standard deviation: 63 per arm with the effect known, 113 at an uncertainty of 0.25, 1268 at 0.5, and no finite number past 0.594, where the prior probability of a positive effect falls to 80%.
Fig. 3 The observations per arm needed for an 80% chance of success, on a logarithmic axis, as the uncertainty about an expected effect of half a standard deviation grows. The vertical rule marks the uncertainty beyond which no sample size is enough.
uncertainty about an effect of 0.5 per arm for 80% success ceiling
none 63 100%
0.1 69 100%
0.25 113 97.7%
0.5 1,268 84.1%
0.75 none 74.8%

A quarter of a standard deviation of uncertainty raises the trial from 63 per arm to 113. Half a standard deviation raises it to 1,268 — twenty times the conventional size. And at three quarters there is no answer.

The ceiling

As the sample grows, the numerator and the denominator of the formula both grow like n\sqrt{n}, and assurance converges to

Φ(δ0/s0)\Phi(\delta_0 / s_0)

which is the probability, under the belief, that the effect is positive. No trial of any size can be more likely to succeed than the planners believe the treatment is to work. An infinitely large trial detects any positive effect however small, and detects nothing if the effect is zero or negative, so its chance of success is exactly the chance the effect is positive.

That makes the ceiling an honest limit rather than a mathematical curiosity. A belief centred at half a standard deviation with an uncertainty of three quarters puts a quarter of its weight below zero, and with that belief the chance of success tops out at 74.8% whatever is spent. At an uncertainty of 0.594 standard deviations the ceiling is exactly 80%, and every uncertainty beyond it makes an 80% target unreachable.

The practical meaning is sharp. A sample-size calculation that returns 64 per arm for 80% power, for a treatment its own planners think has a one-in-four chance of doing nothing or harm, is describing a trial whose realistic chance of success is 57.9% and whose best possible chance, with unlimited patients, is 74.8%. The 80% in the protocol is a statement about a conditional world the planners do not believe they are in.

A smaller expected effect

The ceiling bites harder when the expected effect is small relative to the uncertainty about it.

The chance a trial succeeds against its size, when the expected effect of 0.3 is uncertain by four amounts. With the effect known, 80% is reached at 175 per arm. With the effect uncertain by 0.25 standard deviations it takes 1031; by 0.5, no sample; by 0.75, no sample size at all, because the chance can never exceed the 65.5% prior probability that the effect is positive.
Fig. 4 The same curves for an expected effect of 0.3 standard deviations. The uncertain curves flatten at lower ceilings, and two of them never reach 80%.

At an expected effect of 0.3 standard deviations the conventional trial is 175 per arm. With an uncertainty of 0.1 an 80% chance needs 226; with 0.25 it needs 1,031, against a ceiling of 88.5%; with 0.5 the ceiling is 72.6% and no trial reaches 80%. A modest expected effect that the planners are not confident of — the typical position for a new indication — is exactly where power at the expected effect is most misleading.

A belief from a pilot

The uncertainty about the effect need not be a matter of opinion. A pilot estimates the effect with a standard error, and that standard error is a natural choice for s0s_0 — the same uncertainty the pilot’s own confidence interval expresses.

A pilot of forty per arm estimates a standardised effect with a standard error of 2/40\sqrt{2/40}, 0.224. If it estimated half a standard deviation, a trial of sixty-four per arm planned from that estimate has a chance of success of 70.5%, and an 80% chance needs 100 per arm — a trial more than half as large again as the calculation from the point estimate.

That figure is before the selection that inflates pilot effects is accounted for. A pilot that went ahead because its estimate looked promising overstates the effect as well as being uncertain about it, and both push the realistic chance of success down. Planning from the pilot’s interval rather than its estimate handles the second and not the first.

Both inputs uncertain at once

The spread and the effect are usually both uncertain, and they are usually both taken from the same small pilot. Their costs do not simply add.

The power trials actually have when sized for 80% from a pilot of 10. Four thousand pilots of 10 observations, each sizing a trial for 80% power at half a standard deviation from its own standard deviation. 55.9% of the trials have less than 80% power and 11.1% less than 50%; the median trial has 76.8%.
Fig. 5 The power a trial actually has when a pilot of ten supplies its standard deviation, with the effect known. Adding uncertainty about the effect moves every trial in this histogram down the same concave curve again.

A hundred thousand simulated plans, each sizing a trial from a pilot of ten observations for 80% at an expected effect of half a standard deviation, with the true effect then drawn from a belief of spread a quarter:

what is uncertain chance of success
nothing 80.7%
the spread only 74.8%
the effect only 69.2%
both 65.7%

(All four use the normal approximation to the test, so that the rows differ only in what is uncertain.)

With both inputs uncertain, a trial planned for 80% succeeds about two times in three. The spread costs six points and the effect eleven on their own; together they cost fifteen, less than the sum, because a trial that the pilot happened to make large is also a trial better placed to withstand an effect smaller than expected. With the belief about the effect widened to half a standard deviation, the combination falls to 59.1%.

A belief with a chance of nothing

A normal belief is a convenient description and a flattering one. Treatments are often believed to have some chance of doing nothing at all — the mechanism may simply not work in people — and a spread of possible effects if they do.

A belief that puts 30% on no effect and the rest on a normal around half a standard deviation with spread a quarter has, at sixty-four per arm, a chance of success of 49.2%: the 69.2% of the working case, weighted by 70%, plus the 2.5% chance of a false positive in the right direction, weighted by 30%. Its ceiling is 69.2%, however large the trial.

The ceiling here is almost exactly the probability the planners gave the treatment of working, which is what it should be. A trial is a way of finding out whether a treatment works, and a trial of a treatment believed 70% likely to work can at best succeed about 70% of the time; the rest of the time the correct answer is the one it gives. Presenting such a trial as having 80% power is a category error, since the largest risk to its success is not its size.

The prior the data estimates and a prior that is confident and wrong both treat a prior as something the data can overrule. At the planning stage there are no data yet, and the prior is the only description of the treatment there is. Writing it down with a mass at zero is the plainest way to see what a trial can and cannot deliver.

The replication is a special case

The p-value a replication gets found that an exact replication of a result at p=0.05p = 0.05 is significant again exactly half the time, under the plain model of what the original says about the effect. That number is an assurance.

Plan a replication of the same size as an original whose test statistic was zz, and take the belief about the effect from the original itself: centred at its estimate, with its standard error as the uncertainty. The uncertainty term in the formula then equals one exactly, and assurance collapses to

Φ ⁣(z1.962)\Phi\!\left(\frac{z - 1.96}{\sqrt 2}\right)

which is 50.0% for an original that just reached significance, 72.4% for one at z=2.8z = 2.8 — an original with an 80%-powered-looking result — and 82.7% for one at p=0.001p = 0.001.

So the replication probability that surprises readers is not a separate phenomenon. It is the chance of success of a trial planned from an honest statement of what the previous trial found, and it is low for the same reason every assurance above is below power: the effect is uncertain, the power curve is concave, and a study sized as though the estimate were the truth sits on the steep side for the half of the belief below the estimate.

It also shows what planning from a single estimate implies. A team that repeats a just-significant finding at the same size, expecting 80% power because the original’s effect would give it, is planning a coin toss. Re-estimating the sample size partway through can recover some of this, at the costs that essay measures; planning from the original’s interval in the first place recovers it before any patient is enrolled.

What the calculation should report

Power at an assumed effect is a correct answer to a conditional question: if the effect is this, how often will the trial detect it. It remains useful, because it describes what the design does in a stated world. What it is not is the chance of success, and a protocol that reports it as if it were has substituted the conditional for the unconditional.

The repair costs one more line. State the belief about the effect as a centre and an uncertainty; report power at the centre and the chance of success averaged over the uncertainty; and report the ceiling, which is the planners’ own probability that the treatment works. A trial whose chance of success is 69% and whose ceiling is 98% is a sound investment that is somewhat undersized. A trial whose chance of success is 58% and whose ceiling is 75% is a trial whose largest risk is the treatment rather than the sample, and no increase in size addresses it.

A prior is worth a stated number of observations when it is used to analyse data. The same prior, used to plan a trial, sets the most that trial can hope to achieve, and naming it at the planning stage is the cheapest place to learn that.

What the formulas establish, and what they do not

The closed-form chance of success matches drawing effects from the belief and running the test — 69.20% against 69.29% counted, at sixty-four per arm with an uncertainty of a quarter of a standard deviation.

At ten million per arm it has reached, and not passed, the prior probability that the effect is positive, so with an uncertainty of 0.75 around an effect of 0.5 no sample size reaches 80%.

With both the spread and the effect uncertain the chance of success is lower than with either alone, and higher than their separate costs added would suggest — 65.7% against 74.8% and 69.2%, counted over a hundred thousand plans. A replication planned from the original’s own estimate and standard error succeeds exactly half the time after a result at p = 0.05, which is the assurance formula with its uncertainty term equal to one.

What does not survive is power at the expected effect read as the probability that the trial succeeds. At sixty-four per arm with an uncertainty of a quarter of a standard deviation, the two differ by more than eleven points.

Not claimed: that a normal belief is the right description of uncertainty about an effect. Beliefs about treatments are often a mixture — some chance of no effect at all, and a spread of effects if there is one — and a mixture with a mass at zero lowers the ceiling to one minus that mass directly. The formula also uses the normal approximation to the test, which is what the conventional calculation uses; the exact t power at sixty-four per arm is 80.15% against the approximation’s 80.7%, a difference too small to change any comparison above. And success here means significance in the right direction, not a correct estimate of the effect.

Still open: the outcome the trial measures

Both the spread and the effect are properties of the outcome, and the outcome is itself a choice. Trials often replace a measured quantity — a blood pressure, a symptom score, a time — with whether it crossed a threshold: responder or not, controlled or not, above target or below.

That choice changes the effect, the spread and the information in every observation at once. How much of the power it spends, how the loss depends on where the threshold is placed, and how many more patients the same trial then needs, are measured in an outcome cut in two.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AssuranceEffect sizeNormal approximationPilot studyPriorPrior predictiveSample sizeStatistical power