The analyses that were available and not run

Naming a handful in advance

A preregistration that names three of twenty analyses sits between naming one and naming none, and whether it beats both depends on one thing: how much of the analyst's belief the three hold. An analyst who is 50% sure of a first choice, and whose belief halves down the list, gets 34.8% expected power from naming three, 28.3% from naming one and 22.7% from exploring all twenty. Spread the same 50% evenly over the rest and naming three is the worst of the three plans. The rule for adding one more to the list needs no belief at all: the next analysis must be at least 20.8% as likely as the first to be worth naming.

Worth reading first: Twenty analyses of nothing.

What naming it in advance costs compared two plans against an effect that lives in one of twenty analyses. Name one analysis and test it at an uncorrected 5%, which has 51.5% power at two standard errors if the guess is right and nothing if it is wrong; or run all twenty and correct for them exactly, which has 22.5% power wherever the effect is. The two are worth the same when the analyst is 38% sure of the guess.

Real preregistrations are rarely either plan. A protocol names a primary outcome and two secondary ones; a registered report commits to “the effect on any of three measures of anxiety”; an analysis plan names four covariate sets and promises to correct for them. Every one of those is a middle plan, naming some number mm of the candidate analyses and correcting for those mm. The question left open was how the middle behaves — whether naming three captures most of what naming one captures, as the logarithmic cost of a correction suggests, or whether the middle is a compromise that gets the worst of both.

It is neither in general, and which it is in particular turns on one quantity the analyst has and never writes down.

How much naming more analyses in advance is worth, at an effect of 2 standard errors, for four analysts' beliefsTwenty analyses correlated at 0.6, one of which holds the effect. Each curve is an analyst whose belief falls by a fixed ratio from one choice to the next: 0.2, 0.5, 0.7, 0.9. The best number to name is 1, 3, 5, 15 respectively, and each best set holds between 80% and 90% of the analyst's belief. Every curve ends at 22.7%, the power of correcting for all twenty.00.2000.4005101520analyses named in advance, of twentyexpected powerbelief falls ×0.2 a step (first choice 80%)belief falls ×0.5 a step (first choice 50%)belief falls ×0.7 a step (first choice 30%)belief falls ×0.9 a step (first choice 11%)exact, correlation 0.6 · dashed: all twenty, correctedrings mark the best number to name
Fig. 1 Expected power against the number of analyses named in advance, for four analysts whose belief about which analysis holds the effect falls by a fixed ratio from each choice to the next. The dashed line is correcting for all twenty, where every curve ends; the rings mark each analyst’s best number to name. The slider sets the size of the effect.

What each extra name costs

Every number here is exact. The twenty analyses are correlated at 0.6, as they are when they share a dataset, and under the one-factor model that produces that correlation the chance that none of mm statistics clears a threshold is a single integral over the shared component. So the family-wise threshold for any mm can be solved for exactly rather than read off a simulation, and so can the power — which matters, because the differences between neighbouring values of mm are a few points and a simulation’s own noise would be the same size.

The threshold rises with each name, and it rises fastest at the start:

analyses named exact 5% family-wise threshold power if the effect is among them certainty needed to beat exploring
1 1.960 51.6% 37.9%
2 2.199 43.6% 45.8%
3 2.328 39.1% 51.9%
5 2.480 33.8% 61.3%
10 2.671 27.7% 77.8%
20 2.846 22.7% 100.0%

The last column is the break-even of the essay on naming one, generalised: the probability that the named set contains the effect at which naming that many is worth the same as exploring all twenty. It counts a false positive from a set that missed the effect as a detection, as the earlier comparison did; leaving those out raises every entry by four to six points and changes no conclusion.

Naming more analyses: the power if the set is right, and the certainty in the set that makes it worth it, correlation 0.6. At an effect of two standard errors, naming one analysis gives 51.6% power if it is the right one, three give 39.1%, and all twenty give 22.7%. The certainty that the named set contains the effect needed to beat exploring rises from 37.9% at one to 51.9% at three and 77.8% at ten.
Fig. 2 The power of a named set against an effect it contains, and the certainty that it contains the effect needed to beat correcting for all twenty, against the number named.

The shape the logarithmic argument predicted is there: the second name costs eight points of power, the third four and a half, the tenth a fraction of one. Most of the cost of a correction arrives with the first addition. But the certainty column shows what that concave cost does not imply. Naming three rather than one lowers the power by twelve and a half points and raises the certainty the plan needs from 37.9% to 51.9%, so a three-analysis plan beats exploring only if the analyst is more than half sure the effect is somewhere in the three. Whether that is easier to be than 37.9% sure of one is the whole question, and the arithmetic of the threshold cannot answer it.

The answer lives in the belief

What decides it is how an analyst’s belief is spread over their own ranking of the twenty. Two analysts can both be 50% sure of their first choice and differ completely about the rest.

An analyst whose belief halves down the list — 50% on the first choice, 25% on the second, 12.5% on the third — has put 87.5% of their belief into three analyses. For that analyst naming three is the best plan: 34.8% expected power, against 28.3% for naming one and 22.7% for exploring everything. The looser document is better than the strict one, by more than six points.

An analyst who is 50% sure of the first choice and has no idea about the rest spreads the other half evenly, 2.6% on each of the nineteen. For that analyst the second and third names add almost no belief and cost the full price in threshold, and naming three gives 23.8% — worse than naming one, and barely better than exploring all twenty. Every middle plan is worse than both ends, and the worst is around ten names, at 21.7%: less than exploring, because ten names pay most of the threshold of twenty and cover barely more belief than one.

How much naming more analyses in advance is worth when an analyst has a first choice and no view of the rest, effect 2 standard errors. Twenty analyses correlated at 0.6, one of which holds the effect. Each curve is an analyst sure of a first choice to the degree shown — 70%, 50%, 38%, 20% — with the rest of the belief spread evenly over the other nineteen. The best plan is one, one, one, all twenty respectively; no middle plan is best for any of them. Every curve ends at 22.7%, the power of correcting for all twenty.
Fig. 3 Four analysts who are sure of a first choice to different degrees — 70%, 50%, 38% and 20% — and spread the rest of their belief evenly over the other nineteen analyses. Every curve is lowest in the middle or falls all the way; the rings sit at one end or the other.

The figure makes the pattern general. With the rest of the belief spread evenly, no middle plan is best for any analyst: the three who are at least 38% sure of their first choice should name it alone, and the one who is 20% sure should explore everything. The analyst at 38% is the one the essay on naming one found indifferent between its two plans, and here the indifference extends to every plan in between being worse than either. A list longer than one is worth writing only if the second and later entries are themselves real guesses — which is a statement about the analyst’s knowledge, not about statistics, and is the reason the question could not be settled by the arithmetic of the threshold alone.

It also explains why twenty analyses of nothing and a long preregistered list can look so alike in their results. An analyst with a vague idea who lists ten outcomes to seem careful has run, in effect, the uniform analyst’s ten-name plan: most of the correction’s cost has been paid, little of the belief has been concentrated, and the significant result that emerges, if one does, has been selected nearly as hard as the smallest p-value from an unplanned search. The winner’s curse applies to it in full, and the preregistration has protected the error rate without protecting the estimate.

The hero figure draws four such analysts, with belief falling by factors of 0.2, 0.5, 0.7 and 0.9 a step. Their best numbers to name are 1, 3, 5 and 15, and their best sets hold 80%, 88%, 83% and 90% of their belief respectively. The number of analyses is not what is being chosen. What is being chosen is a share of belief, and at this effect size and correlation the right share is somewhere around four fifths to nine tenths — whatever number of analyses it takes to reach it.

That is a cleaner answer than the question expected. “Should a preregistration name one analysis or a handful?” has no general answer, and “name the smallest set of analyses that holds most of the analyst’s belief” does. An analyst sure of one measure names one. An analyst who knows it is one of three anxiety scales and cannot say which names three. An analyst who knows only that it is “some outcome” has nothing to name, and correcting for everything is the honest plan.

The rule for one more

The belief is the analyst’s and cannot be computed. What can be computed is the price it is being compared with, and that turns out to need no belief at all.

Going from mm named analyses to m+1m + 1 raises the expected power exactly when the added analysis’s probability of holding the effect, as a share of the probability the first mm already hold, exceeds

powerm−powerm+1powerm+1−5%,\frac{\text{power}_m - \text{power}_{m+1}}{\text{power}_{m+1} - 5\%},

the relative power lost to the stricter threshold. The ratio depends only on the family and the effect size — the threshold for mm and for m+1m + 1, and the power at each — so it is a table an analyst can hold their own belief against.

When naming one more analysis pays: the next one's share of belief it must carry, correlation 0.6. Adding a second named analysis pays when it is at least 20.8% as likely to hold the effect as the first, at an effect of two standard errors; a third, when it is 13.2% as likely as the first two together; a fourth, 9.7%. Against an effect of three standard errors the shares are 6.8%, 4.9% and 3.8%.
Fig. 4 The share of belief the next named analysis must carry, relative to the analyses already named, for naming it to raise the expected power, at effects of two and three standard errors.

At twenty analyses correlated at 0.6 and an effect of two standard errors:

  • a second analysis is worth naming if it is at least 20.8% as likely to hold the effect as the first;
  • a third, if it is at least 13.2% as likely as the first two together;
  • a fourth, 9.7%; a fifth, 7.7%; a sixth, 6.4%.

Against a larger effect the bar is much lower — 6.8%, 4.9% and 3.8% for the second, third and fourth at three standard errors — because a larger effect clears a stricter threshold easily and the price of a name is almost nothing.

How much naming more analyses in advance is worth, at an effect of 3 standard errors, for four analysts' beliefs. Twenty analyses correlated at 0.6, one of which holds the effect. Each curve is an analyst whose belief falls by a fixed ratio from one choice to the next: 0.2, 0.5, 0.7, 0.9. The best number to name is 2, 5, 8, 20 respectively, and each best set holds between 94% and 100% of the analyst's belief. Every curve ends at 58.5%, the power of correcting for all twenty.
Fig. 5 The four geometric analysts again, against an effect of three standard errors. Every curve’s best point moves right: the lists worth writing are longer, and they hold more of the belief.

The four geometric analysts show what the lower bar does to a whole plan. Against an effect of three standard errors their best lists are 2, 5, 8 and 20 names long rather than 1, 3, 5 and 15, and those lists hold between 94% and all of the analysts’ belief rather than four fifths to nine tenths. When the effect is large, missing it is the expensive mistake and a stricter threshold is cheap, so the right list is the one that almost certainly contains the effect. When it is small, the threshold is expensive and the right list is shorter than the analyst’s uncertainty would suggest. That is the practical form of a trade-off every study planner already makes about sample size: the effect the study expects decides not only how many participants it needs but how many analyses it can afford to name — and the second decision is usually made without the first in view, although the chance a trial succeeds depends on both. That is the same direction the break-even moved in the essay on naming one, where a large effect made exploring attractive; here it makes a longer list attractive, which is the same fact seen from the middle.

The correlation moves the bar the way the effective count would predict. Independent analyses make each name expensive — the second must carry 25.4% of the first’s belief — and analyses correlated at 0.85 make it cheap, 12.7%, because twenty versions of one question are closer to one question and a correction for them costs little.

What the rule says about real documents

The rule turns a debate about strictness into a question an analyst can answer honestly.

The single primary outcome is right more often than it is fashionable to say. An analyst who would put a second outcome at less than a fifth of the first’s likelihood gains nothing by naming it, and a regulatory insistence on one primary endpoint is, on these numbers, the correct plan for any sponsor who knows which endpoint the drug should move. The strictness is not buying rigour at a cost in power; for a confident analyst it is also the most powerful plan.

A short list of near-equivalent measures should be named together. Three scales of one construct, each about as likely as the others to show the effect, are the case the geometric analyst describes, and the rule says to name all three: each carries about half the belief of the set before it, well above the 20.8% and 13.2% bars. Choosing one of them at random for the sake of a single primary outcome throws away belief for no gain in rigour, since the correction for three correlated measures is small.

A long list named for the sake of completeness is the worst document there is. Ten outcomes listed because they might matter, with belief concentrated on one or two, pay nine tenths of the threshold of exploring everything and cover little more belief than the first name. The uniform analyst’s ten-name plan, at 21.7%, is below the power of simply exploring all twenty with the exact correction — so the long list is dominated by honestly admitting the analysis is exploratory. It has also, in practice, the reputation of a preregistration, which the exploratory plan does not.

And the belief should be written down. The rule needs the analyst’s belief about the list and nothing else, and the belief is the one input no protocol asks for. A preregistration that stated an expectation of the effect in the first measure with probability about one half, and among these three with probability about nine in ten, would let a reader check the choice of mm against the table, and would say — more usefully than any list — how exploratory the study really was.

The false positives a wrong list produces

Every expected power above counts a significant result from a list that did not contain the effect as though it were a detection. It is not: it is the list’s 5% family-wise false-positive rate landing on the wrong analyses, and it will be written up as a finding. The correction that makes the estimate worse is a reminder that what survives a threshold is selected, and a false positive from a wrong list is the most selected of all, since it is pure noise that cleared the bar.

Removing those false positives lowers every plan’s expected power by 5% of the probability that the list missed, which penalises short lists most, since they miss most often. The best number to name moves up for three of the four analysts drawn — from one to two for the most confident, from five to six, and from fifteen to all twenty for the least certain, whose best plan becomes simply exploring — and the break-even certainties rise by four to six points — 43.9% rather than 37.9% for a single name, 58.0% rather than 51.9% for three. The confident analyst’s single name loses its edge to a pair, which is the one place the answer shifts; the rule’s structure does not. But it is worth saying plainly that the single-name plan’s advantage is partly made of false positives when the name is wrong, which is the cliff the essay on naming one warned about, seen once more from the side.

The exact prices of a longer list, and the belief they leave to the analyst

The exact family-wise threshold for mm of twenty analyses correlated at 0.6 rises from 1.960 at one to 2.328 at three and 2.846 at twenty, and the power against an effect of two standard errors inside the named set falls from 51.6% to 39.1% to 22.7%.

Naming more pays when the added analysis carries at least 20.8%, 13.2% and 9.7% of the belief already named, for the second, third and fourth; and the best list, for analysts whose belief decays geometrically at four different rates, holds between four fifths and nine tenths of their belief.

The middle plans are best only for an analyst whose belief is concentrated on a few analyses. For one who is sure of a first choice and indifferent about the rest, every middle plan is worse than both naming one and exploring all twenty.

Every rate is exact under the one-factor normal model of correlated analyses: a one-dimensional integral over the shared component for each probability, and bisection on it for each threshold. The single-name figures agree with the simulation in the essay on naming one to within its error — 51.6% here against 51.5% there, 22.7% against 22.5%.

Not claimed: that analysts’ beliefs decay geometrically, or at any particular rate. The four curves are illustrations of how the answer depends on the belief, and the rule for one more is the part that does not depend on it. The equal correlation of every pair of analyses is the simplest structure that produces the shared-data effect; a real family has blocks — three scales of one construct correlated more strongly with each other than with the rest — and in such a family the cost of naming a block together is lower than this table’s, which strengthens the case for naming near-equivalents together.

Still open: the list that was written after looking

Everything here assumes the list was written before the data. The documents that most need the arithmetic are the ones where that is in doubt — a protocol amended during the trial, a registered primary outcome that quietly becomes secondary, an analysis plan finalised after an interim look. A list chosen with partial knowledge of the data is not a prior; it is a selection, and twenty analyses of nothing prices what selection does to a threshold that pretends it did not happen.

The quantity that would settle how much a late list costs is the correlation between the data used to choose it and the data used to test it. A list chosen from a pilot that shares no participants with the trial is a legitimate prior; one chosen from the first half of the trial’s own data is partly a look at the answer. Between those two the family-wise rate the list claims is wrong by an amount that depends on the overlap, and nothing here has measured it.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CorrelationFalse positiveFamilywise error rateThe garden of forking pathsMultiple comparisonsPreregistrationResearcher degrees of freedomStatistical power