How many subjects
Worth reading first: What a p-value does not say · The variance removed before the data.
Every design decision in this field eventually converts into one number: how many units the experiment needs. It is the calculation everybody has heard of, it is the one most often done wrong, and until this essay it was the one quantity on this site computed by simulation with no closed form beside it.
The quantity being computed
Power is the probability that the test rejects the null when the null is false — a probability over repeated experiments, like every other quantity on this site, and therefore countable.
It depends on four things and no others: the size of the effect, the variability of the outcome, the number of units, and the level the test is run at. Three of those are usually known or chosen. The first is not, which is where every difficulty in this subject lives.
The convention is to state the effect in units of the outcome’s own standard deviation — d = 0.5 is “half a standard deviation” — which makes the calculation independent of what is being measured, at the cost of hiding a decision inside a ratio. More on that below, because it is where sample size calculations go wrong in practice rather than in arithmetic.
Sixty-four per arm
At d = 0.5, a two-sided test at the 5% level, and 80% power, the answer is 64 units per arm.
It is worth memorising, because it anchors everything else. Power scales with d²n, so:
- at d = 0.2 — a small effect, and the size most social and behavioural effects turn out to be — the requirement is 394 per arm;
- at d = 0.8 — a large effect — it is 26 per arm.
A factor of four in the effect size is a factor of fifteen in the units required, which is the single most consequential fact about designing studies and the reason underpowered research is so persistent: halving the effect being chased quadruples the study.
What the closed form makes cheap to ask
Once the answer is per arm, three questions a designer always asks can be answered without another calculation.
How the sample scales with the effect. The constant at 5% and 80% power is , so per arm: 393 at d = 0.2, 174 at 0.3, 63 at 0.5, 25 at 0.8 and 16 at 1.0. Halving the effect quadruples the study, and the range a designer might plausibly guess over — say 0.3 to 0.5 — is a factor of nearly three in cost.
What more power costs. Raising the target from 80% to 90% replaces 0.842 with 1.282, giving — 34% more subjects. Ninety-five per cent power costs 66% more than eighty.
What a stricter level costs. Moving from α = 0.05 to α = 0.01 replaces 1.96 with 2.576 and gives — 49% more.
So the three dials are not equally expensive. A tenth of power costs a third of the study, a fivefold tightening of the level costs a half, and a third off the effect size costs a doubling. The one nobody gets to choose is the one that dominates, which is why every honest sample size calculation is a sensitivity analysis over d rather than a number.
The approximation is worth one line of its own: at d = 0.5 it gives 62.8 where the non-central t gives 64. The difference is what having to estimate the variance costs, and it is a little over one subject per arm.
Two routes, and the gap this closes
The statistic in a two-sample t test does not follow a t distribution when the null is false. Its numerator is off-centre, and the distribution it does follow is the non-central t, with a non-centrality parameter δ = d√(n/2) on 2n − 2 degrees of freedom.
Power is then the probability that this non-central t exceeds the critical value of the ordinary one, which is a closed form and needs no simulation at all.
The absence of that closed form has been recorded on this site since its first essays. Every power number here — the error rates spent across interim looks, the winner’s curse inflation at low power, the two different promises a test and an interval make — was a Monte Carlo count with nothing independent to disagree with. Where a count is the only route to a number, a bug in the simulation and a fact about the world look identical.
The two routes now agree at every sample size on the curve, within four standard errors of the simulation’s own noise, on a check that runs on every build. The dots in these figures are not decoration: they are the second route, drawn on top of the first.
What the normal approximation overstates
The calculation most people actually do replaces the non-central t with a normal, which amounts to pretending the standard deviation was known rather than estimated from the same small sample.
At n = 10 per arm and d = 1, the normal approximation gives 60.9% power where the exact answer is 56.2% — nearly five percentage points, always in the same direction. The approximation says the study is better than it is.
By n = 400 the two agree to within a twentieth of a percentage point, which is why the approximation survives: it is wrong exactly where studies are small, which is exactly where it is used, and right where nobody needed the extra precision. Five points of overstated power is not a catastrophe on its own; it becomes one in combination with everything else in this essay that also points the same way.
The number that is guessed
Everything above treats d as given, and it is not. It is the ratio of the effect worth detecting to the outcome’s standard deviation, and both parts are usually estimates.
There are three ways it is chosen and they are not equally defensible.
From a pilot study. The most common and the worst. A pilot is small, so its estimate of the effect is noisy, and it is usually run because somebody thought there might be something there — which selects for pilots that overstate. That is the winner’s curse exactly: an effect estimated conditional on having looked promising is inflated, and a sample size computed from it is too small, so the full study is underpowered, and its own estimate is inflated in turn.
From what would be worth acting on. The defensible version. The smallest effect that would change a decision is a judgement about the subject rather than about the data, it can be stated in advance, and it does not depend on any previous estimate. It usually produces a larger study than anyone wanted, which is the honest answer rather than a problem with the method.
Backwards from the affordable sample. Common and rarely admitted: fix n at what the budget allows, solve for the d that gives 80% power, and describe that as the effect the study is designed to detect. It is not dishonest if reported as what it is — but the resulting d is often far larger than anything plausible, and reporting it makes that visible, which is why it usually is not reported.
What an underpowered study actually produces
The usual account of low power is that the study probably misses a real effect, which is true and is the smaller half of the problem.
The larger half is what happens when it does not miss. A study with 20% power that reports a significant result has, by construction, observed an unusually large estimate — because at that power only unusually large estimates clear the threshold. So the published effect from an underpowered study is inflated, systematically, and the inflation is measurable: it is the winner’s curse, and this site counts it at over twice the true effect where power is low.
Three consequences follow, and they compound.
The literature is biased upwards. Not by anyone’s dishonesty — by the arithmetic of a threshold applied to noisy estimates.
Replications fail. A replication powered from the published, inflated estimate is powered for an effect larger than the real one, so it is underpowered for the real one, and it fails. That is read as a contradiction between studies rather than as a property of how the first was sized.
And each failure feeds the next calculation. A pilot inflates, the study is undersized, its significant result inflates further, and the next study is sized from that.
Power is therefore not only about whether one study will find something. It is about whether a body of work converges on the truth, and at low power it does not converge on it at all.
Power after the fact is not a quantity
One practice worth naming because it appears in published papers and in referee reports.
Observed power — power computed by plugging the effect the study actually estimated into the power formula — is not a diagnostic. It is a deterministic function of the p-value: a study with p just above 0.05 always has observed power near 50%, whatever the data was. It carries no information beyond the p-value it was computed from, and it cannot support the sentence it is usually deployed to support, which is that a non-significant result was adequately powered.
The useful post-hoc quantity is the confidence interval, which says which effects the data is compatible with. A non-significant study whose interval spans −0.1 to +0.9 has not ruled out a large effect; one whose interval spans −0.05 to +0.05 has. That distinction is what “was the study big enough” means after the fact, and the interval answers it directly.
This is the same argument as the second number throughout this site: a p-value alone is uninterpretable, and the quantity that interprets it is the one describing how large the effect could plausibly be.
What raises power other than n
Sample size is one of four inputs and the most expensive to change. The other three are worth listing because the rest of this field is about them.
Reduce the variability. Everything in the blocking essay is a way of doing this without collecting a single extra unit: removing a nuisance factor by design cut the variance to a fifth, which is a fivefold reduction in the units needed.
Measure the outcome better. A less noisy measurement is a smaller σ, which is a larger d for the same real effect. Repeated measurement per unit, a more precise instrument, a better-defined endpoint — each of these buys power in the same currency and often more cheaply than recruiting.
Choose what to test. A factorial design estimates every main effect from every run, so it reaches the same precision on half the runs a one-at-a-time programme needs. Power is a property of the design, not only of its size.
And raise α, deliberately, where the trade justifies it. A screening experiment whose false positives are cheap to follow up has no reason to run at 0.05. Testing at 0.10 raises power substantially at every sample size, and it is a defensible choice when stated in advance — the opposite of the same choice made after seeing the result.
The two arms need not be equal, and usually should be
One arithmetic detail with a large practical consequence, since it is the first thing a constrained study should reach for.
The variance of the estimated difference is σ²(1/n₁ + 1/n₂), which for a fixed total is smallest when the arms are equal. Equal allocation is optimal, and the optimum is flat: a two-to-one split carries about 12% more variance than an equal one, which is a 12% increase in the units needed rather than a doubling.
That flatness is what makes unequal allocation worth considering when the arms cost different amounts. If one treatment is expensive and the control is nearly free, more controls per treated unit buys precision cheaply, and the optimum shifts in proportion to the square root of the cost ratio. Three controls per treated unit recovers most of what four would, so the useful range is narrow.
The same flatness is a caution in the other direction. A trial that ended three-to-one because recruitment ran unevenly has lost little; one that planned three-to-one to make the treated group look larger has spent real precision for presentation.
Power is a curve, and one point of it is a poor summary
“80% power” fixes one point on a curve and the rest of it is where the interesting reading is.
An experiment sized for d = 0.5 at 80% power has, at those same 64 units per arm, about 50% power at d = 0.35 and about 20% at d = 0.2. Those are not failures of the design — they are what the design says about effects other than the one it was sized for, and they are computable in advance from the same closed form, before a single unit is recruited.
Two of those three numbers are worth carrying together, because the drop is steeper than intuition suggests: cutting the effect by a third takes the study from four-in-five to a coin flip, and cutting it by half takes it to one-in-five. Nothing about the experiment changed.
Reporting the curve rather than the point makes two questions answerable that the point does not. What is the smallest effect this study can reliably see? and how much would the study have to grow to see something half that size? — to which the answer is always the same shape: four times as many units, from the d²n scaling.
For a reader assessing a published study, the curve is also the honest way to read a null result. A study that had 80% power at d = 0.5 and reported nothing has said something firm about effects near 0.5 and almost nothing about effects near 0.2, and the interval it reported will show exactly that.
What a sample size calculation should report
Three lines, all of which follow from the above.
The effect it is powered for, and where that came from. “80% power to detect d = 0.5” is incomplete without whether 0.5 is a pilot estimate, a decision threshold, or the result of dividing the affordable sample back out.
The test it assumes. The 64 above is for a two-sample t test at the 5% level, two-sided. A different test, a different level, or a one-sided alternative gives a different number, and the difference between one- and two-sided at the same nominal level is a fifth of the sample.
Whether the design was accounted for. A blocked design, a factorial arrangement, or repeated measures all change the variance the calculation should use, and a sample size computed from the unblocked variance for a blocked study is the mismatch the blocking essay describes, in advance rather than in retrospect.
What this field established
Four decisions taken before any data exists, and each one has a number attached.
Blocking removes exactly the share of the variance the blocks carry — forty units doing the work of a hundred and eighty-nine.
Randomisation does not deliver balance, leaving a spread of exactly 2/√n in every covariate, and delivers instead a reference set that makes a test exact with no distributional assumption in it.
Factorial arrangement estimates every main effect from every run, at (k+1)/2 times the precision per run, and answers the question about combinations that one-factor-at-a-time answers wrong 100% of the time.
And sample size is the arithmetic all three feed into: 64 per arm at half a standard deviation, 394 at a fifth, 26 at four fifths — with a closed form that this site did not have until now, and a simulation beside it that has to agree.
The four share a property worth stating as the field’s own claim. Every one of them is a decision whose consequences are computable before any data exists, and every one of them is therefore auditable by a reader who never sees the data: the blocking factor and its variance share, the allocation rule, the arrangement of the factors, and the effect the study was sized for. A methods section containing those four things describes an experiment whose properties are known. One missing any of them describes an experiment whose properties will have to be guessed at afterwards, which is the work every other field on this site is doing.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Randomisation is not balance — both name allocation, blocking, p-value, standard deviation
- The tail converges last — both name normal approximation, p-value, sample size, standard deviation
- Allocating on a guess — both name allocation, sample size, standard deviation
- Randomising towards the winner — both name allocation, blocking, statistical power
- The reversal a coin cannot prevent — both name allocation, blocking, sample size
- A block size that changes — both name blocking, sample size
Named objects
A flat tag is an object no other essay names yet.
AllocationBlockingEffect sizeThe non-central tNormal approximationp-valueSample sizeStandard deviationStatistical power