The chance a trial succeeds
Worth reading first: What a p-value does not say.
The spread a pilot supplies took one input to a sample-size calculation — the standard deviation — and showed that estimating it leaves the planned trial short of its power more often than not. The other input is the effect. How many subjects is clear that it is chosen rather than known: from a pilot, from what would be worth acting on, or backwards from a budget.
Whichever way it is chosen, the calculation then treats it as a fact. Power is a function of the true effect, and it is evaluated at one value. The question a funder, an ethics committee or a patient actually asks is different: given what is believed about the effect, how likely is this trial to succeed?
Power averaged over the effect
Suppose the effect, in units of the outcome’s standard deviation, is believed to be around with an uncertainty that can be described as a normal distribution of spread . A trial of per arm produces a test statistic whose mean is for a given true effect . Averaged over the belief about , the statistic is normal with mean and variance — the trial’s own noise plus the planners’ uncertainty, scaled by the sample. The probability that it clears the significance threshold is
With this is the ordinary power at . With any uncertainty, the denominator grows with , and it changes the behaviour of the whole curve.
The formula is an average of power over a belief, and it can be checked as one. Drawing two hundred thousand effects from a normal belief centred at 0.5 with spread 0.25 and running a test at sixty-four per arm on each gives a success rate of 69.29%; the closed form gives 69.20%.
At the conventional sample
Sixty-four per arm is the trial with 80% power at half a standard deviation. Its chance of success, as the belief about the effect widens:
| uncertainty about the effect | chance of success at sixty-four per arm |
|---|---|
| none | 80.7% |
| 0.1 standard deviations | 77.5% |
| 0.25 | 69.2% |
| 0.5 | 61.4% |
| 0.75 | 57.9% |
(The first row is the normal approximation’s power at sixty-four; the exact t power is 80.15%.)
A trial designed for 80% power is a trial with about a two-in-three chance of success when the planners are honestly unsure of the effect by a quarter of a standard deviation. That is not a pessimistic uncertainty. An effect “of about half a standard deviation” believed to lie, two times in three, between a quarter and three quarters of one is a modest claim about most treatments before they are tested.
The source of the shortfall is the same concavity that shortened pilot-planned trials. Power rises steeply below the design effect and flattens above it, so averaging over effects on both sides pulls the average below the value at the centre. Uncertainty about the effect costs success in a way that no symmetry of the belief can undo.
Why averaging over the effect lowers the chance
The shortfall has a picture, and it is the ordinary power curve read sideways.
At sixty-four per arm the exact power is 28.93% if the true effect is a quarter of a standard deviation, 80.15% at half, and 98.78% at three quarters. A belief that puts equal weight on a quarter and three quarters has a centre of one half, and its chance of success is the average of the two ends: 63.9%, sixteen points below the power at the centre.
That is the entire mechanism in three numbers. A trial sized to reach 80% at the expected effect sits on the steep part of the power curve for smaller effects and the flat part for larger ones, so an effect a little smaller than expected costs far more than an effect a little larger returns. The more spread out the belief, the more weight falls on the steep side, and the further the average falls below the power at the centre.
The same shape made the pilot-planned trial fall short in the spread a pilot supplies. There the uncertain input moved the sample size along the curve; here it moves the effect. Either way, a calculation that evaluates a concave curve at its expected input overstates the expected output. The gap is not a rounding matter: at the settings above it is the difference between a trial that meets its target four times in five and one that meets it two times in three, and it is invisible in any calculation that reports a single number.
The sample an 80% chance takes
If the goal is an 80% chance of success rather than 80% power at one effect, the formula can be solved for .
| uncertainty about an effect of 0.5 | per arm for 80% success | ceiling |
|---|---|---|
| none | 63 | 100% |
| 0.1 | 69 | 100% |
| 0.25 | 113 | 97.7% |
| 0.5 | 1,268 | 84.1% |
| 0.75 | none | 74.8% |
A quarter of a standard deviation of uncertainty raises the trial from 63 per arm to 113. Half a standard deviation raises it to 1,268 — twenty times the conventional size. And at three quarters there is no answer.
The ceiling
As the sample grows, the numerator and the denominator of the formula both grow like , and assurance converges to
which is the probability, under the belief, that the effect is positive. No trial of any size can be more likely to succeed than the planners believe the treatment is to work. An infinitely large trial detects any positive effect however small, and detects nothing if the effect is zero or negative, so its chance of success is exactly the chance the effect is positive.
That makes the ceiling an honest limit rather than a mathematical curiosity. A belief centred at half a standard deviation with an uncertainty of three quarters puts a quarter of its weight below zero, and with that belief the chance of success tops out at 74.8% whatever is spent. At an uncertainty of 0.594 standard deviations the ceiling is exactly 80%, and every uncertainty beyond it makes an 80% target unreachable.
The practical meaning is sharp. A sample-size calculation that returns 64 per arm for 80% power, for a treatment its own planners think has a one-in-four chance of doing nothing or harm, is describing a trial whose realistic chance of success is 57.9% and whose best possible chance, with unlimited patients, is 74.8%. The 80% in the protocol is a statement about a conditional world the planners do not believe they are in.
A smaller expected effect
The ceiling bites harder when the expected effect is small relative to the uncertainty about it.
At an expected effect of 0.3 standard deviations the conventional trial is 175 per arm. With an uncertainty of 0.1 an 80% chance needs 226; with 0.25 it needs 1,031, against a ceiling of 88.5%; with 0.5 the ceiling is 72.6% and no trial reaches 80%. A modest expected effect that the planners are not confident of — the typical position for a new indication — is exactly where power at the expected effect is most misleading.
A belief from a pilot
The uncertainty about the effect need not be a matter of opinion. A pilot estimates the effect with a standard error, and that standard error is a natural choice for — the same uncertainty the pilot’s own confidence interval expresses.
A pilot of forty per arm estimates a standardised effect with a standard error of , 0.224. If it estimated half a standard deviation, a trial of sixty-four per arm planned from that estimate has a chance of success of 70.5%, and an 80% chance needs 100 per arm — a trial more than half as large again as the calculation from the point estimate.
That figure is before the selection that inflates pilot effects is accounted for. A pilot that went ahead because its estimate looked promising overstates the effect as well as being uncertain about it, and both push the realistic chance of success down. Planning from the pilot’s interval rather than its estimate handles the second and not the first.
Both inputs uncertain at once
The spread and the effect are usually both uncertain, and they are usually both taken from the same small pilot. Their costs do not simply add.
A hundred thousand simulated plans, each sizing a trial from a pilot of ten observations for 80% at an expected effect of half a standard deviation, with the true effect then drawn from a belief of spread a quarter:
| what is uncertain | chance of success |
|---|---|
| nothing | 80.7% |
| the spread only | 74.8% |
| the effect only | 69.2% |
| both | 65.7% |
(All four use the normal approximation to the test, so that the rows differ only in what is uncertain.)
With both inputs uncertain, a trial planned for 80% succeeds about two times in three. The spread costs six points and the effect eleven on their own; together they cost fifteen, less than the sum, because a trial that the pilot happened to make large is also a trial better placed to withstand an effect smaller than expected. With the belief about the effect widened to half a standard deviation, the combination falls to 59.1%.
A belief with a chance of nothing
A normal belief is a convenient description and a flattering one. Treatments are often believed to have some chance of doing nothing at all — the mechanism may simply not work in people — and a spread of possible effects if they do.
A belief that puts 30% on no effect and the rest on a normal around half a standard deviation with spread a quarter has, at sixty-four per arm, a chance of success of 49.2%: the 69.2% of the working case, weighted by 70%, plus the 2.5% chance of a false positive in the right direction, weighted by 30%. Its ceiling is 69.2%, however large the trial.
The ceiling here is almost exactly the probability the planners gave the treatment of working, which is what it should be. A trial is a way of finding out whether a treatment works, and a trial of a treatment believed 70% likely to work can at best succeed about 70% of the time; the rest of the time the correct answer is the one it gives. Presenting such a trial as having 80% power is a category error, since the largest risk to its success is not its size.
The prior the data estimates and a prior that is confident and wrong both treat a prior as something the data can overrule. At the planning stage there are no data yet, and the prior is the only description of the treatment there is. Writing it down with a mass at zero is the plainest way to see what a trial can and cannot deliver.
The replication is a special case
The p-value a replication gets found that an exact replication of a result at is significant again exactly half the time, under the plain model of what the original says about the effect. That number is an assurance.
Plan a replication of the same size as an original whose test statistic was , and take the belief about the effect from the original itself: centred at its estimate, with its standard error as the uncertainty. The uncertainty term in the formula then equals one exactly, and assurance collapses to
which is 50.0% for an original that just reached significance, 72.4% for one at — an original with an 80%-powered-looking result — and 82.7% for one at .
So the replication probability that surprises readers is not a separate phenomenon. It is the chance of success of a trial planned from an honest statement of what the previous trial found, and it is low for the same reason every assurance above is below power: the effect is uncertain, the power curve is concave, and a study sized as though the estimate were the truth sits on the steep side for the half of the belief below the estimate.
It also shows what planning from a single estimate implies. A team that repeats a just-significant finding at the same size, expecting 80% power because the original’s effect would give it, is planning a coin toss. Re-estimating the sample size partway through can recover some of this, at the costs that essay measures; planning from the original’s interval in the first place recovers it before any patient is enrolled.
What the calculation should report
Power at an assumed effect is a correct answer to a conditional question: if the effect is this, how often will the trial detect it. It remains useful, because it describes what the design does in a stated world. What it is not is the chance of success, and a protocol that reports it as if it were has substituted the conditional for the unconditional.
The repair costs one more line. State the belief about the effect as a centre and an uncertainty; report power at the centre and the chance of success averaged over the uncertainty; and report the ceiling, which is the planners’ own probability that the treatment works. A trial whose chance of success is 69% and whose ceiling is 98% is a sound investment that is somewhat undersized. A trial whose chance of success is 58% and whose ceiling is 75% is a trial whose largest risk is the treatment rather than the sample, and no increase in size addresses it.
A prior is worth a stated number of observations when it is used to analyse data. The same prior, used to plan a trial, sets the most that trial can hope to achieve, and naming it at the planning stage is the cheapest place to learn that.
What the formulas establish, and what they do not
The closed-form chance of success matches drawing effects from the belief and running the test — 69.20% against 69.29% counted, at sixty-four per arm with an uncertainty of a quarter of a standard deviation.
At ten million per arm it has reached, and not passed, the prior probability that the effect is positive, so with an uncertainty of 0.75 around an effect of 0.5 no sample size reaches 80%.
With both the spread and the effect uncertain the chance of success is lower than with either alone, and higher than their separate costs added would suggest — 65.7% against 74.8% and 69.2%, counted over a hundred thousand plans. A replication planned from the original’s own estimate and standard error succeeds exactly half the time after a result at p = 0.05, which is the assurance formula with its uncertainty term equal to one.
What does not survive is power at the expected effect read as the probability that the trial succeeds. At sixty-four per arm with an uncertainty of a quarter of a standard deviation, the two differ by more than eleven points.
Not claimed: that a normal belief is the right description of uncertainty about an effect. Beliefs about treatments are often a mixture — some chance of no effect at all, and a spread of effects if there is one — and a mixture with a mass at zero lowers the ceiling to one minus that mass directly. The formula also uses the normal approximation to the test, which is what the conventional calculation uses; the exact t power at sixty-four per arm is 80.15% against the approximation’s 80.7%, a difference too small to change any comparison above. And success here means significance in the right direction, not a correct estimate of the effect.
Still open: the outcome the trial measures
Both the spread and the effect are properties of the outcome, and the outcome is itself a choice. Trials often replace a measured quantity — a blood pressure, a symptom score, a time — with whether it crossed a threshold: responder or not, controlled or not, above target or below.
That choice changes the effect, the spread and the information in every observation at once. How much of the power it spends, how the loss depends on where the threshold is placed, and how many more patients the same trial then needs, are measured in an outcome cut in two.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The price of control — both name effect size, sample size, statistical power
- A boundary for giving up — both name sample size, statistical power
- A coverage table with its own error — both name sample size, statistical power
- Allocating on a guess — both name pilot study, sample size
- Not half and half — both name sample size, statistical power
- Pooling a proportion — both name normal approximation, prior
Named objects
A flat tag is an object no other essay names yet.
AssuranceEffect sizeNormal approximationPilot studyPriorPrior predictiveSample sizeStatistical power