Naming a handful in advance
Worth reading first: Twenty analyses of nothing.
What naming it in advance costs compared two plans against an effect that lives in one of twenty analyses. Name one analysis and test it at an uncorrected 5%, which has 51.5% power at two standard errors if the guess is right and nothing if it is wrong; or run all twenty and correct for them exactly, which has 22.5% power wherever the effect is. The two are worth the same when the analyst is 38% sure of the guess.
Real preregistrations are rarely either plan. A protocol names a primary outcome and two secondary ones; a registered report commits to “the effect on any of three measures of anxiety”; an analysis plan names four covariate sets and promises to correct for them. Every one of those is a middle plan, naming some number of the candidate analyses and correcting for those . The question left open was how the middle behaves — whether naming three captures most of what naming one captures, as the logarithmic cost of a correction suggests, or whether the middle is a compromise that gets the worst of both.
It is neither in general, and which it is in particular turns on one quantity the analyst has and never writes down.
What each extra name costs
Every number here is exact. The twenty analyses are correlated at 0.6, as they are when they share a dataset, and under the one-factor model that produces that correlation the chance that none of statistics clears a threshold is a single integral over the shared component. So the family-wise threshold for any can be solved for exactly rather than read off a simulation, and so can the power — which matters, because the differences between neighbouring values of are a few points and a simulation’s own noise would be the same size.
The threshold rises with each name, and it rises fastest at the start:
| analyses named | exact 5% family-wise threshold | power if the effect is among them | certainty needed to beat exploring |
|---|---|---|---|
| 1 | 1.960 | 51.6% | 37.9% |
| 2 | 2.199 | 43.6% | 45.8% |
| 3 | 2.328 | 39.1% | 51.9% |
| 5 | 2.480 | 33.8% | 61.3% |
| 10 | 2.671 | 27.7% | 77.8% |
| 20 | 2.846 | 22.7% | 100.0% |
The last column is the break-even of the essay on naming one, generalised: the probability that the named set contains the effect at which naming that many is worth the same as exploring all twenty. It counts a false positive from a set that missed the effect as a detection, as the earlier comparison did; leaving those out raises every entry by four to six points and changes no conclusion.
The shape the logarithmic argument predicted is there: the second name costs eight points of power, the third four and a half, the tenth a fraction of one. Most of the cost of a correction arrives with the first addition. But the certainty column shows what that concave cost does not imply. Naming three rather than one lowers the power by twelve and a half points and raises the certainty the plan needs from 37.9% to 51.9%, so a three-analysis plan beats exploring only if the analyst is more than half sure the effect is somewhere in the three. Whether that is easier to be than 37.9% sure of one is the whole question, and the arithmetic of the threshold cannot answer it.
The answer lives in the belief
What decides it is how an analyst’s belief is spread over their own ranking of the twenty. Two analysts can both be 50% sure of their first choice and differ completely about the rest.
An analyst whose belief halves down the list — 50% on the first choice, 25% on the second, 12.5% on the third — has put 87.5% of their belief into three analyses. For that analyst naming three is the best plan: 34.8% expected power, against 28.3% for naming one and 22.7% for exploring everything. The looser document is better than the strict one, by more than six points.
An analyst who is 50% sure of the first choice and has no idea about the rest spreads the other half evenly, 2.6% on each of the nineteen. For that analyst the second and third names add almost no belief and cost the full price in threshold, and naming three gives 23.8% — worse than naming one, and barely better than exploring all twenty. Every middle plan is worse than both ends, and the worst is around ten names, at 21.7%: less than exploring, because ten names pay most of the threshold of twenty and cover barely more belief than one.
The figure makes the pattern general. With the rest of the belief spread evenly, no middle plan is best for any analyst: the three who are at least 38% sure of their first choice should name it alone, and the one who is 20% sure should explore everything. The analyst at 38% is the one the essay on naming one found indifferent between its two plans, and here the indifference extends to every plan in between being worse than either. A list longer than one is worth writing only if the second and later entries are themselves real guesses — which is a statement about the analyst’s knowledge, not about statistics, and is the reason the question could not be settled by the arithmetic of the threshold alone.
It also explains why twenty analyses of nothing and a long preregistered list can look so alike in their results. An analyst with a vague idea who lists ten outcomes to seem careful has run, in effect, the uniform analyst’s ten-name plan: most of the correction’s cost has been paid, little of the belief has been concentrated, and the significant result that emerges, if one does, has been selected nearly as hard as the smallest p-value from an unplanned search. The winner’s curse applies to it in full, and the preregistration has protected the error rate without protecting the estimate.
The hero figure draws four such analysts, with belief falling by factors of 0.2, 0.5, 0.7 and 0.9 a step. Their best numbers to name are 1, 3, 5 and 15, and their best sets hold 80%, 88%, 83% and 90% of their belief respectively. The number of analyses is not what is being chosen. What is being chosen is a share of belief, and at this effect size and correlation the right share is somewhere around four fifths to nine tenths — whatever number of analyses it takes to reach it.
That is a cleaner answer than the question expected. “Should a preregistration name one analysis or a handful?” has no general answer, and “name the smallest set of analyses that holds most of the analyst’s belief” does. An analyst sure of one measure names one. An analyst who knows it is one of three anxiety scales and cannot say which names three. An analyst who knows only that it is “some outcome” has nothing to name, and correcting for everything is the honest plan.
The rule for one more
The belief is the analyst’s and cannot be computed. What can be computed is the price it is being compared with, and that turns out to need no belief at all.
Going from named analyses to raises the expected power exactly when the added analysis’s probability of holding the effect, as a share of the probability the first already hold, exceeds
the relative power lost to the stricter threshold. The ratio depends only on the family and the effect size — the threshold for and for , and the power at each — so it is a table an analyst can hold their own belief against.
At twenty analyses correlated at 0.6 and an effect of two standard errors:
- a second analysis is worth naming if it is at least 20.8% as likely to hold the effect as the first;
- a third, if it is at least 13.2% as likely as the first two together;
- a fourth, 9.7%; a fifth, 7.7%; a sixth, 6.4%.
Against a larger effect the bar is much lower — 6.8%, 4.9% and 3.8% for the second, third and fourth at three standard errors — because a larger effect clears a stricter threshold easily and the price of a name is almost nothing.
The four geometric analysts show what the lower bar does to a whole plan. Against an effect of three standard errors their best lists are 2, 5, 8 and 20 names long rather than 1, 3, 5 and 15, and those lists hold between 94% and all of the analysts’ belief rather than four fifths to nine tenths. When the effect is large, missing it is the expensive mistake and a stricter threshold is cheap, so the right list is the one that almost certainly contains the effect. When it is small, the threshold is expensive and the right list is shorter than the analyst’s uncertainty would suggest. That is the practical form of a trade-off every study planner already makes about sample size: the effect the study expects decides not only how many participants it needs but how many analyses it can afford to name — and the second decision is usually made without the first in view, although the chance a trial succeeds depends on both. That is the same direction the break-even moved in the essay on naming one, where a large effect made exploring attractive; here it makes a longer list attractive, which is the same fact seen from the middle.
The correlation moves the bar the way the effective count would predict. Independent analyses make each name expensive — the second must carry 25.4% of the first’s belief — and analyses correlated at 0.85 make it cheap, 12.7%, because twenty versions of one question are closer to one question and a correction for them costs little.
What the rule says about real documents
The rule turns a debate about strictness into a question an analyst can answer honestly.
The single primary outcome is right more often than it is fashionable to say. An analyst who would put a second outcome at less than a fifth of the first’s likelihood gains nothing by naming it, and a regulatory insistence on one primary endpoint is, on these numbers, the correct plan for any sponsor who knows which endpoint the drug should move. The strictness is not buying rigour at a cost in power; for a confident analyst it is also the most powerful plan.
A short list of near-equivalent measures should be named together. Three scales of one construct, each about as likely as the others to show the effect, are the case the geometric analyst describes, and the rule says to name all three: each carries about half the belief of the set before it, well above the 20.8% and 13.2% bars. Choosing one of them at random for the sake of a single primary outcome throws away belief for no gain in rigour, since the correction for three correlated measures is small.
A long list named for the sake of completeness is the worst document there is. Ten outcomes listed because they might matter, with belief concentrated on one or two, pay nine tenths of the threshold of exploring everything and cover little more belief than the first name. The uniform analyst’s ten-name plan, at 21.7%, is below the power of simply exploring all twenty with the exact correction — so the long list is dominated by honestly admitting the analysis is exploratory. It has also, in practice, the reputation of a preregistration, which the exploratory plan does not.
And the belief should be written down. The rule needs the analyst’s belief about the list and nothing else, and the belief is the one input no protocol asks for. A preregistration that stated an expectation of the effect in the first measure with probability about one half, and among these three with probability about nine in ten, would let a reader check the choice of against the table, and would say — more usefully than any list — how exploratory the study really was.
The false positives a wrong list produces
Every expected power above counts a significant result from a list that did not contain the effect as though it were a detection. It is not: it is the list’s 5% family-wise false-positive rate landing on the wrong analyses, and it will be written up as a finding. The correction that makes the estimate worse is a reminder that what survives a threshold is selected, and a false positive from a wrong list is the most selected of all, since it is pure noise that cleared the bar.
Removing those false positives lowers every plan’s expected power by 5% of the probability that the list missed, which penalises short lists most, since they miss most often. The best number to name moves up for three of the four analysts drawn — from one to two for the most confident, from five to six, and from fifteen to all twenty for the least certain, whose best plan becomes simply exploring — and the break-even certainties rise by four to six points — 43.9% rather than 37.9% for a single name, 58.0% rather than 51.9% for three. The confident analyst’s single name loses its edge to a pair, which is the one place the answer shifts; the rule’s structure does not. But it is worth saying plainly that the single-name plan’s advantage is partly made of false positives when the name is wrong, which is the cliff the essay on naming one warned about, seen once more from the side.
The exact prices of a longer list, and the belief they leave to the analyst
The exact family-wise threshold for of twenty analyses correlated at 0.6 rises from 1.960 at one to 2.328 at three and 2.846 at twenty, and the power against an effect of two standard errors inside the named set falls from 51.6% to 39.1% to 22.7%.
Naming more pays when the added analysis carries at least 20.8%, 13.2% and 9.7% of the belief already named, for the second, third and fourth; and the best list, for analysts whose belief decays geometrically at four different rates, holds between four fifths and nine tenths of their belief.
The middle plans are best only for an analyst whose belief is concentrated on a few analyses. For one who is sure of a first choice and indifferent about the rest, every middle plan is worse than both naming one and exploring all twenty.
Every rate is exact under the one-factor normal model of correlated analyses: a one-dimensional integral over the shared component for each probability, and bisection on it for each threshold. The single-name figures agree with the simulation in the essay on naming one to within its error — 51.6% here against 51.5% there, 22.7% against 22.5%.
Not claimed: that analysts’ beliefs decay geometrically, or at any particular rate. The four curves are illustrations of how the answer depends on the belief, and the rule for one more is the part that does not depend on it. The equal correlation of every pair of analyses is the simplest structure that produces the shared-data effect; a real family has blocks — three scales of one construct correlated more strongly with each other than with the rest — and in such a family the cost of naming a block together is lower than this table’s, which strengthens the case for naming near-equivalents together.
Still open: the list that was written after looking
Everything here assumes the list was written before the data. The documents that most need the arithmetic are the ones where that is in doubt — a protocol amended during the trial, a registered primary outcome that quietly becomes secondary, an analysis plan finalised after an interim look. A list chosen with partial knowledge of the data is not a prior; it is a selection, and twenty analyses of nothing prices what selection does to a threshold that pretends it did not happen.
The quantity that would settle how much a late list costs is the correlation between the data used to choose it and the data used to test it. A list chosen from a pilot that shares no participants with the trial is a legitimate prior; one chosen from the first half of the trial’s own data is partly a look at the answer. Between those two the family-wise rate the list claims is wrong by an amount that depends on the overlap, and nothing here has measured it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- False discoveries that arrive together — both name correlation, false positive, familywise error rate, multiple comparisons, statistical power
- An order that spends the error rate — both name false positive, familywise error rate, multiple comparisons, statistical power
- One control, many arms — both name correlation, familywise error rate, multiple comparisons, statistical power
- Estimating how many nulls are true — both name correlation, multiple comparisons, statistical power
- The models that were never in the running — both name familywise error rate, multiple comparisons, statistical power
- The price of control — both name false positive, multiple comparisons, statistical power
Named objects
A flat tag is an object no other essay names yet.
CorrelationFalse positiveFamilywise error rateThe garden of forking pathsMultiple comparisonsPreregistrationResearcher degrees of freedomStatistical power