What naming it in advance costs
Worth reading first: Twenty analyses of nothing.
The repair for twenty analyses of nothing is to say which one is the analysis before the data arrive. The 5% comes back, the multiplicity disappears, and nothing appears to have been given up.
Something has been given up, and it is measurable.
Two strategies, three numbers
An effect of a given size sits in exactly one of twenty analyses, which are correlated at 0.6 because they share a dataset.
Prespecify. Name one analysis in advance and test it at an uncorrected 5%. If it is the right one, the power is the ordinary power of that test: 51.5% against an effect of two standard errors. If it is the wrong one, the analysis is being run against a null, and it “detects” the effect 4.7% of the time — which is not a detection, it is the test’s own false-positive rate landing on the wrong analysis.
Explore. Run all twenty and apply the exact family-wise threshold, which for this family is 2.848 standard errors. The power is the chance that the largest of the twenty clears it: 22.5%, and that number does not depend on which analysis holds the effect.
The comparison is between a strategy with a cliff and a strategy without one. Prespecifying is more than twice as powerful when it is right and worthless when it is wrong; correcting is uniformly mediocre.
The break-even
Let q be the probability that the prespecified analysis is the one holding the effect. The prespecified strategy’s expected power is
and the exploratory strategy’s is 22.5% regardless. Setting them equal gives q = 38%.
Below 38% confidence in the guess, correcting for all twenty is the better bet. Above it, prespecifying is.
The 4.7% in that expression deserves a second look, because it is doing something the arithmetic hides. It is not a benefit; it is the wrong analysis returning a significant result when there is nothing in it to find. Counting it as a “detection” is generous to the prespecified strategy — every one of those is a false positive that will be written up as a finding — and removing it raises the break-even from 38% to 44%. The table below keeps it in, because taking it out would make the comparison one of power against something that is not power, and the honest version of that objection is that the prespecified strategy’s cliff is worse than the numbers show rather than better.
That is a low bar and it is the reason preregistration is usually the right choice: an investigator who has a specific hypothesis is usually more than 38% sure which measure it will show up in. But the number is not a constant, and the direction it moves in is the surprise.
The bar rises with the size of the effect
| effect, in standard errors | prespecified, and right | all twenty, corrected | break-even certainty |
|---|---|---|---|
| 1.5 | 32.1% | 12.1% | 27% |
| 2.0 | 51.5% | 22.5% | 38% |
| 2.5 | 70.8% | 38.6% | 51% |
| 3.0 | 85.3% | 58.5% | 67% |
| 3.5 | 94.0% | 77.0% | 81% |
| 4.0 | 98.1% | 89.4% | 91% |
Against a large effect, prespecification needs to be more certain to be worth it. At four standard errors the break-even is 91%: a study looking for a large effect and 80% sure where it lives does better to look everywhere and correct.
The mechanism is a ceiling. A large effect clears almost any threshold, so the correction costs almost nothing — 98.1% against 89.4%, a loss of nine points. What prespecification risks is the whole detection, and against a large effect the thing risked is worth nearly 100% while the thing saved is worth nine points. The ratio is what the break-even measures, and it rises.
Against a small effect the situation reverses. The correction costs twenty points out of thirty-two, which is most of what there was, so almost any confidence in the guess is worth having.
So the argument for preregistration is strongest exactly where the argument for research is weakest — when the effect sought is small, the study is underpowered, and neither strategy works very well. That is an uncomfortable finding and it is what the arithmetic says.
And it falls as the analyses become more alike
The correlation between the analyses is the other parameter, and it moves the break-even the way the effective-count measurement predicts.
At a correlation of zero the break-even at two standard errors is 30%; at 0.6 it is 38%; at 0.85 it is 55%. Analyses that share more of their data are cheaper to correct for, so the exploratory strategy improves and prespecification has to be more confident to beat it.
That gives a rule that can be applied without any of the arithmetic. Prespecify when the twenty candidate analyses are genuinely different questions; explore and correct when they are twenty versions of one question. Twenty distinct outcome measures on distinct constructs are the first case. Twenty exclusion rules, twenty covariate sets, twenty ways of handling the same missing values are the second, and correcting for those costs very little because they were never twenty analyses.
The power the design chose, and the power the analysis chose
The break-even is a statement about the analysis plan, and it sits next to a quantity the design already fixed, so the two are worth putting on one scale.
Sixty-four subjects an arm is what 80% power at half a standard deviation costs. Running twenty analyses instead of one and correcting for them exactly takes a study from 51.5% power to 22.5% at two standard errors — which, in subjects, is the difference between the study that was designed and a study less than half its size.
That conversion is the useful one, because it puts multiplicity in the currency a study is planned in. Correcting for twenty correlated analyses costs about as much power as halving the sample, at the effect sizes most studies are powered for. A study that has paid for four hundred subjects and then reports twenty outcomes has spent the money and thrown the design away, and the loss is in the analysis plan rather than in anything visible in the methods section.
The corollary is the reason the trade is worth taking seriously rather than dismissing. Nobody would accept a proposal to halve a study’s sample size to make the write-up more flexible. That is exactly what running twenty outcomes and correcting properly does, and it does it silently, because power is computed at the design stage against one outcome and the multiplicity arrives later.
The strategy of naming the analysis after seeing part of the data
A fourth arrangement deserves naming because it is honest, it is rarer than it should be, and it sits outside the table.
Split the data. Choose the analysis on the first half, run it on the second, and report the second at an uncorrected 5%. The threshold is legitimate: the second half has seen no selection, so there is one analysis and no family.
Its power is the power of the right analysis at half the sample — and choosing the right analysis from the first half is itself imperfect, so the strategy also pays for the chance of choosing wrongly. Against an effect of two standard errors in the whole sample, each half carries an effect of about 1.41 standard errors, the chance of the right analysis having the largest statistic in the first half is modest, and the second half then tests at that reduced size.
The virtue is not power. It is that the strategy converts an unquantifiable problem — how many analyses would this author have run? — into a quantifiable one, at a stated cost. That is the trade the exact correction cannot make, because it needs the family enumerated and splitting needs only the discipline to not look.
Where it is genuinely good is when the family is large and unknowable. A screen over hundreds of candidate analyses has no enumerable family, so a family-wise correction cannot be computed honestly; splitting requires no enumeration at all, and the cost is a known factor rather than an unknown one.
What the comparison leaves out, and it is in favour of prespecifying
Three things are missing from the table and all three make prespecification look better than it does above.
The exploratory strategy’s estimate is worse. Both strategies produce an effect size along with their p-value, and the exploratory one has selected a maximum. Its reported effect is inflated by an amount the correction increases rather than reduces, and a detection that reports an effect 1.7 times its true size is a weaker result than a detection that reports it correctly.
The family has to be complete. The exploratory strategy’s threshold is computed over the analyses that were run, and it controls the error rate only if those are all the analyses that would have been run. Prespecification makes that condition checkable — the document says what the family is — and the exploratory strategy relies on the author having enumerated it honestly, which is the thing the enterprise was supposed to stop relying on.
A prespecified analysis can be wrong and still be informative. Naming the wrong outcome does not only fail to detect; it produces a reported null for a specific question, which is publishable and interpretable. The “4.7%” above counts only the chance of a false detection and gives no credit for the correct null. The exploratory strategy that fails produces twenty nulls, which is a weaker statement about each of them.
None of those is in the table because none of them is a power. They are the reason the break-even figures are a floor on the case for prespecification rather than a settlement of it.
There is a fourth, and it runs the other way. The exploratory strategy finds effects nobody was looking for. A prespecified study that names the wrong outcome has produced a null and stopped; an exploratory one has twenty statistics and can see that something is happening in the eleventh, even when the corrected threshold is not cleared. That is not a detection and should not be reported as one, but it is a hypothesis, and a hypothesis is what a confirming study needs before it can be run. The table above is a comparison of two ways of ending an investigation, and only one of them is also a way of starting one.
The third strategy, which is the one people actually use
Neither column of the table describes common practice. What is common is to run all twenty, report the best, and describe it as though it had been prespecified.
That strategy’s power against an effect of two standard errors is the chance that the largest of twenty clears an uncorrected threshold, which is high — and its error rate when there is no effect at all is 35.5% rather than 5%, which is the measurement twenty honest analyses produce.
So it is not a third point on the same trade-off curve. It is off the curve entirely: it buys its power by spending an error rate it does not report, and the two honest strategies are the ones that pay for their power out of something they declare.
The useful way to hold the comparison is that both honest strategies cost about the same, and the choice between them is a judgement about how well the question is understood, not about rigour. A field that treats preregistration as the rigorous option and exploration as the lax one has mistaken a power trade-off for an ethical one — and, at large effect sizes, has the trade-off backwards.
What is claimed here, and what is not
Two statements carry the figure and they constrain it from opposite directions.
The prespecified analysis has more power against an effect that is in it, at every effect size drawn. That is not interesting on its own — it must be true, since one analysis at 5% is a looser test than twenty at a corrected level — and it is exactly why it is worth stating: if it ever came out the other way, the exploratory arm would be testing at a threshold below the uncorrected one, which is the plausible slip and would invert the whole page.
The certainty prespecification requires rises with the size of the effect. That is the finding, and it holds as a monotone relationship between the two ends of the sweep. It would not if the exploratory power were computed against a fixed threshold instead of the family’s own, since a fixed threshold would make both columns rise together and leave the ratio flat.
The reading that does not survive is the strategy in the section above, taken for the second column. The exploratory arm’s threshold is the measured family-wise one; were it the uncorrected 5%, the “explore” curve would sit above the “prespecify” curve everywhere and the break-even would be negative, and the page would read as an argument against preregistration. That the break-even is a probability between 0 and 1 at every setting is what says the two arms are being compared at the same error rate.
What a registered report changes about all of this
The comparison has treated preregistration as a statement about which analysis will be run. The stronger form changes the quantity being traded rather than its price.
A registered report commits a journal to publishing the result before the data exist. That does nothing to the power in the table — a prespecified analysis has the same 51.5% either way — and it removes the selection that operates after the analysis, which is the filter the winner’s curse depends on and which no correction described here touches.
The two devices are therefore not alternatives and not the same thing. Prespecification fixes the family the p-value is computed over. Guaranteed publication fixes the population of results that reaches a reader. A literature can have the first without the second — every study preregistered, only the positive ones published — and the reported effects would still be inflated by the same selection as before, with every individual p-value correct.
Which of the two matters more depends on where the filtering is. If authors run twenty analyses and report one, the first device is the repair. If authors run one analysis and report it only when it works, the second is. Both are usually true at once, and the arithmetic for their combination is not on this page.
Still open: naming the analysis loosely
Preregistration in practice is not the binary this page models. A document that names an outcome measure but not the exclusion rule, or a family of three outcomes rather than one, sits between the two columns — and the space between them is where almost all real preregistrations live.
The arithmetic for that middle is available and is not done here: a document naming m of the twenty analyses has power somewhere between the two, a break-even that depends on m, and an error rate that is the family-wise rate over m rather than over twenty. What is not obvious in advance is whether the trade-off is close to linear in m or whether most of the benefit arrives in the first step from twenty to a handful.
The practical question that would answer is a real one. If naming three analyses captures most of what naming one captures, then the strictness that makes preregistration unpopular — commit to one outcome, one rule, one model — is buying very little, and a looser document would be adopted more widely at almost no cost.
The shape of the answer can be guessed from the effective-count arithmetic and not from here. The correction’s severity depends on the count logarithmically, so going from twenty to three moves the threshold much further than going from three to one — which argues that most of the benefit does arrive in the first step, and that a document naming a handful is nearly as good as one naming a single analysis. What that guess omits is the second column of the trade: naming three raises the chance that the effect is in one of them, and by more than it raises the threshold. Both effects push the same way, which is unusual and is why the question is worth doing properly rather than arguing about.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- An order that spends the error rate — both name false positive, familywise error rate, multiple comparisons, statistical power
- False discoveries that arrive together — both name false positive, familywise error rate, multiple comparisons, statistical power
- The price of control — both name effect size, false positive, multiple comparisons, statistical power
- An outcome cut in two — both name effect size, the garden of forking paths, statistical power
- One control, many arms — both name familywise error rate, multiple comparisons, statistical power
- The models that were never in the running — both name familywise error rate, multiple comparisons, statistical power
Named objects
A flat tag is an object no other essay names yet.
Effect sizeFalse positiveFamilywise error rateThe garden of forking pathsMultiple comparisonsPreregistrationResearcher degrees of freedomStatistical power