An order that spends the error rate
Worth reading first: What the correction corrects.
Holm tests twenty hypotheses as a set. It sorts their p-values, compares the smallest with 5% divided by twenty and each next one with a slightly larger share, and stops at the first that fails. No hypothesis is more important than another, and with ten real effects of three standard errors among the twenty, each real effect is found 52.5% of the time — the price the price of control put on a guarantee about the whole set.
Trials seldom treat their hypotheses as a set. They declare a primary hypothesis and then secondary ones, in an order, and the declaration says which finding the trial is about. A procedure can use that order directly. The fixed sequence tests the first hypothesis at the full 5%; if the first is rejected, it tests the second at the full 5%; and it stops at the first hypothesis it cannot reject. Every test that is run gets the whole error rate.
The familywise rate still holds at 5%, and the reason is short. A false rejection anywhere in the list needs every hypothesis before it to have been rejected — including the first true null on the list, which was tested at 5%. So the chance of any false rejection is at most the chance of rejecting that first true null. The guarantee is the same one Holm makes, and it is the stronger of the two promises a correction can make: the chance of any false finding, not their share. What differs is which hypotheses get the power.
The whole 5% on each test in turn
With the ten real effects listed first, the fixed sequence finds the first 85.3% of the time. That is simply the power of one test at 5% against an effect of three standard errors, because that is what the first test is. Holm, sharing 5% across twenty, finds the same hypothesis 52.5% of the time, so the first hypothesis gains 32.8 points by being declared first.
The second hypothesis is reached only if the first was rejected, so its chance is the product of two powers, and the fifth the product of five. At the fifth position the fixed sequence finds its hypothesis 45.0% of the time, already below Holm; at the tenth, 20.4%. Averaged over the ten real effects it finds 46.10%, against Holm’s 52.53%.
The order moved power from the back of the list to the front and lost some on the way. That loss is not a flaw of the procedure; it is the point of it. A trial that declares an order has said the first hypothesis matters more than the tenth, and the fixed sequence takes the statement at its word.
It also spends its error rate in a particular place. On this list the familywise rate is 1.07%: the ten true nulls sit behind ten real effects, and a chain of ten rejections has to get through before any of them is tested at all. Most of the 5% is never used.
One true null at the head of the list
Now move one true null to the front of the same list. It is tested at 5% and rejected 5% of the time, and otherwise the list stops there. Every real effect behind it inherits that. The first real effect, now second on the list, is found 4.3% of the time; the fifth real effect, 2.2%; the tenth, 1.1%; averaged over all ten, 2.33%. Holm, which does not look at the order, still finds 52.53%.
That is the cliff a fixed sequence builds into a protocol. Its familywise rate on this list is 5.08% — the whole budget, spent on the one hypothesis that could not rightly be rejected — against 1.07% with the real effects first. The same procedure uses almost none of its error rate on one list and all of it on another, and the difference is which hypothesis somebody wrote down first.
A second procedure keeps the order and removes the cliff. Wiens’s fallback divides the 5% equally along the list, one twentieth to each hypothesis, and lets a hypothesis that is rejected pass its share on to the next; a hypothesis that is not rejected keeps its share from passing. With the real effects first it finds the first 49.1% of the time and the tenth 57.9%, rising along the list as rejected shares accumulate, and 55.87% of the ten overall — more than Holm. With the true null at the head it finds 55.87% again, because a failure at the head costs only the head’s own twentieth.
Fallback’s rise along the list comes from what a rejection carries. Its first hypothesis is tested at a twentieth of 5%, 0.25%, which is Bonferroni’s level, and against an effect of three standard errors that finds it 49.1% of the time. A rejection passes its 0.25% on, so the second hypothesis is tested at 0.5% when the first was rejected and at 0.25% when it was not, and so on down the list: every run of rejections raises the level of the next test, and a failure resets it to a twentieth. The real effects at the front of the list build up a level; the true nulls behind them start from a twentieth again, which is why they are almost never rejected.
Three hypotheses, which is the shape trials have
Twenty hypotheses in a fixed order is a thought experiment. A primary endpoint and two secondary ones is a protocol, and it is where the fixed sequence is actually used.
With three real effects of three standard errors, the fixed sequence finds the primary 85.3% of the time, Holm 82.0% and fallback 72.8%. Over all three, Holm finds 82.0% of the real effects, fallback 77.9% and the fixed sequence 73.5%. Among three hypotheses the fixed sequence’s advantage on the primary is 3.3 points over Holm, and it pays for them with 8.6 points on the set.
When the order is right in the stronger sense — strongest effect first, at three, two and a half and two standard errors — the primary still gets 85.3% against Holm’s 77.9%, and the mean over three is 58.8% against 61.7%. When it is backwards — weakest first — the fixed sequence finds the primary 51.7% of the time, still above Holm’s 45.1%, and the set 40.0% against 62.0%: the primary keeps its advantage however the effects are arranged, and the set pays more the worse the arrangement.
With a true null as the primary, the fixed sequence finds the two real secondaries 3.8% of the time. Holm finds them 76.6% of the time and fallback 76.1%. With the true null last, the fixed sequence finds the real ones 79.1% of the time, above Holm’s 76.5%, because a null at the end of the list costs nothing that comes before it.
The pattern fits in a sentence. A fixed sequence is a bet that the first hypothesis is real: when it is, the bet wins a few points on that hypothesis, and when it is not, the bet loses more than seventy points on everything declared after it.
Two routes to the price of an order
For independent tests the fixed sequence needs no simulation: the chance of reaching and rejecting the k-th hypothesis is the product of the first k powers. One test at 5% against three standard errors rejects 85.08% of the time, so the second of three equal effects is found 85.08% squared, 72.39% of the time, and the third 61.59%. Counted over twenty thousand families, the same positions come out at 85.30%, 72.86% and 62.20%.
Across fifteen positions in five scenarios the largest gap between the count and the product is 0.60 points. The five scenarios were counted on the same draws, so their counts drift together rather than independently — which is why every position of the equal-effects scenario is a little high at once, and why the product, which has no draws in it, is the number to quote.
Correlated endpoints
Trial endpoints are seldom independent, and neither are the comparisons a shared control arm makes, correlated at by that sharing alone. A primary outcome and its secondaries are measured on the same people, and a patient who responds on one tends to respond on the others. So the three equal effects were also counted with every pair of endpoints correlated at 0.5 and at 0.9.
The first hypothesis is still found 85.2% of the time at 0.5 and 85.3% at 0.9 — its own power at 5%, whatever the others do. The second and third are no longer products. At a correlation of 0.5 the fixed sequence finds them 76.2% and 69.7% of the time, against 72.9% and 62.2% independent; at 0.9, 81.2% and 78.7%. Holm moves the other way, to 80.4% for each at 0.5 and 78.2% at 0.9, because its first step reads the smallest of the three p-values, and the smallest of three correlated p-values is less small than the smallest of three independent ones. At a correlation of 0.9 the fixed sequence beats Holm at every one of the three positions.
The reason is conditioning. A trial whose primary endpoint crossed is, more often than not, a trial in which the shared part of the response was favourable, and the secondary endpoints inherit that. The product formula assumed each step down the list was a fresh chance; with correlated endpoints each step is partly the same chance again, and a chain of three highly correlated tests behaves more like one test than like three. With the strongest effect first — three, two and a half and two standard errors — and a correlation of 0.9, the fixed sequence finds the three 85.3%, 69.8% and 50.5% of the time, against Holm’s 73.4%, 60.3% and 49.0%.
Correlation rescues nothing behind a true null. With the null first, the fixed sequence’s familywise rate is 5.12% at 0.5 and 5.11% at 0.9, and it finds the real secondaries behind it 3.6% and 2.6% of the time; Holm still finds them 76.2% and 75.3% of the time. The cliff does not depend on how the endpoints move together, because it is the first test’s failure, and the first test’s chance of failing does not.
Where each procedure spends 5%
All three procedures keep the familywise rate at 5% on every list counted. They differ in where on the list the rate is spent, and so in how much of it any list uses.
The fixed sequence spends everything at the first true null. With a null first among three it rejects a true null 4.83% of the time, and among twenty 5.08%; with the real effects first among twenty, 1.07%. Fallback spends a twentieth at each position and passes on what is not needed: 1.65% with a null first among three, 3.05% with the real effects first among twenty. Holm spends by rank rather than position and uses 3.56% among twenty and 3.78% among three with a null first, whatever the order.
That is the same budget-framing a group-sequential trial uses across its looks, turned from time to hypotheses. O’Brien–Fleming spends almost nothing at the first look and nearly everything at the last; a fixed sequence spends everything at the first hypothesis and nothing afterwards. Holm is the Bonferroni of the set, sharing the budget by rank, and what the correction corrects showed it never does worse than dividing equally. Fallback sits between, and its evenness is why it is the only one of the three whose total power did not move when a null was put at the head.
The budget a procedure leaves unspent is a budget something else could have used. On the list with the real effects first, the fixed sequence spends 1.07% of its 5% and leaves 3.93 points unused, because its true nulls are almost never reached. Fallback spends 3.05% on the same list, and the difference is the level its later positions carry — which is why its tenth real effect is found 57.9% of the time where the fixed sequence’s is found 20.4%.
The three procedures are members of one larger family. Each passes the level of a rejected hypothesis on to others along declared routes: the fixed sequence passes everything to the next hypothesis, fallback passes a rejected share one step forward, and Holm passes it to every hypothesis still untested. The graphical procedures of Bretz and colleagues let a protocol draw those routes itself — which secondary inherits the primary’s level, and in what share. None of those is counted here, and each is the same bet made at a finer grain: a declaration, before the data, of where the error rate goes when a hypothesis is rejected.
An order chosen after the data is not an order
Every guarantee above assumes the list was written down before any hypothesis was tested. An order chosen after the p-values are seen — primary becomes whichever endpoint came out best — turns the fixed sequence into a procedure that tests its smallest p-value at the full 5%, which is uncorrected testing of the minimum of twenty. The familywise rate of that is the 64% twenty uncorrected tests produce, arrived at through a protocol that appears to control it.
The size of that failure is a closed form. Choosing the primary as the smallest of three true-null p-values and testing it at 5% rejects a true null with probability 1 − 0.95³, 14.26%; with twenty candidates, 64.15%. No step of the fixed sequence changed. The error rate changed because the first hypothesis became a maximum.
The same distinction ran through a look the trend asked for: a schedule fixed in advance spends exactly its budget, and a schedule chosen from the data spends more, by an amount nothing in the report reveals. An order is a schedule over hypotheses. It is worth exactly as much as the evidence that it was declared before it could be read off the results — a registration date, a protocol amendment history — and nothing in the reported p-values can supply that evidence.
What a declared order should say
Which procedure uses it. A fixed sequence and a fallback on the same list are different designs, with the same guarantee and power profiles that differ by more than seventy points on some lists.
Why the first hypothesis is first. The fixed sequence is a bet that it is real. A protocol that puts a speculative hypothesis first to give it the full 5% has put every hypothesis after it at risk of being untestable, and that risk is a number: under 4.3% power for each real effect behind a true null among twenty.
How correlated the endpoints are expected to be. Everything the fixed sequence offers the secondary hypotheses depends on it: at a correlation of 0.9 the third of three equal effects is found 78.7% of the time, and with independent endpoints 62.2%. A protocol choosing between a fixed sequence and Holm is choosing on that number whether or not it states it.
The power on each position, computed before the trial. For independent tests it is a product of single-test powers and needs no software beyond a normal table. A trial whose third secondary hypothesis has 31% power under the planned order has a secondary hypothesis in name.
What is exact here and what is counted
Exact. The fixed sequence’s power at each position for independent tests, as a product of normal tail probabilities; and the familywise guarantee of all three procedures, which is an argument about which hypothesis must be rejected first rather than a calculation.
Counted. Every power and familywise rate in the figures, over twenty thousand families of independent tests per scenario, and the fallback and Holm figures throughout, for which no product formula applies.
Particular to these lists. Effects of two to three standard errors; three or twenty hypotheses; one true null placed first or last, or ten at the back; and endpoints independent or sharing one correlation of 0.5 or 0.9. Endpoints correlated unequally — a primary closely tied to one secondary and loosely to another — are not measured, and for them neither the product formula nor the equal-correlation counts is the right reference.
Still open: intervals for the findings
Every procedure in this series has decided which hypotheses to report. None has said what to report about them. A finding goes out with an estimate and an interval, and the interval is computed as though the hypothesis had not been selected — for Holm, for a fixed sequence and for Benjamini–Hochberg alike. Intervals for the findings counts how often those intervals miss, which side they miss on, and what an interval widened for the selection costs.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- False discoveries that arrive together — both name error rate, false positive, familywise error rate, holm, monte carlo, multiple comparisons, statistical power
- Eight forecasters and one benchmark — both name bonferroni, error rate, familywise error rate, monte carlo, multiple comparisons, null hypothesis
- A coverage table with its own error — both name bonferroni, closed form, monte carlo, multiple comparisons, statistical power
- The models that were never in the running — both name error rate, familywise error rate, monte carlo, multiple comparisons, statistical power
- A boundary for giving up — both name closed form, error rate, monte carlo, statistical power
- A simulation that stops when it looks settled — both name closed form, error rate, monte carlo, statistical power
Named objects
A flat tag is an object no other essay names yet.
BonferroniClosed formError rateFalse positiveFamilywise error rateHolmMonte CarloMultiple comparisonsNull hypothesisStatistical powerStudy design