Testing at 0.05 every time the data is looked at
The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.
Stopping rulesslider: nominal level, 4 positionswide8 views
What else it draws
The same object, drawn to answer the other questions the essays put to it.
Forty trials in which nothing is happening, monitored at 5 interim points. The boundary is 1.96, 1.96, 1.96, 1.96, 1.96. 3 of the forty cross it somewhere and would be reported as significant.
test at 0.05 every look: 1.96, 1.96, 1.96, 1.96, 1.96. Pocock — a constant, higher boundary: 2.41, 2.41, 2.41, 2.41, 2.41. O'Brien–Fleming — strict early, nearly nominal at the end: 4.55, 3.22, 2.63, 2.27, 2.03. Bonferroni across looks: 2.58, 2.58, 2.58, 2.58, 2.58. Every one except the first spends the same total error rate; they differ in when they spend it.
test at 0.05 every look: rejects a true null 14.4%, finds a real effect 92%, uses 204 observations on average. Pocock: rejects a true null 4.9%, finds a real effect 82%, uses 256 observations on average. O'Brien–Fleming: rejects a true null 4.9%, finds a real effect 88%, uses 299 observations on average. Bonferroni across looks: rejects a true null 3.1%, finds a real effect 77%, uses 274 observations on average.
The dashed line is the smallest effect a trial can report and still stop at each look: 0.510 at 80 observations, 0.255 at 160 observations, 0.170 at 240 observations, 0.128 at 320 observations, 0.102 at 400 observations. The true effect is 0.16, so at 3 of the five looks a trial cannot stop without reporting more than it. 29 of these forty trials stop before the last look, each marked where it stopped.
Each column is one look of an O'Brien–Fleming trial; above the boundary a trial stops there. Highlighted are the outcomes that count as at least as extreme as the observed one when outcomes are ordered stagewise: at 80, z ≥ 4.56 (probability 2.54 × 10⁻⁶ with no effect); at 160, z ≥ 3.30 (probability 4.82 × 10⁻⁴ with no effect); at 240, none; at 320, none; at 400, none. The two-sided p-value is 9.69 × 10⁻⁴.
The upper line is the benefit boundary (4.56, 3.23, 2.63, 2.28, 2.04); the lower line is where the rule stops a trial for futility (0.40 at 80, 0.66 at 160, 0.95 at 240, 1.31 at 320). Of forty trials with a real effect, 29 cross for benefit and 11 are stopped for futility.
Exact. At an interim |z| of 1.0: end only 3.64%, +0.6 3.41%, +0.75 3.20%, +0.9 3.41%, every 0.125 3.00%, every 0.05 2.88%. At 2.5: end only 38.30%, +0.6 45.80%, +0.75 45.36%, +0.9 41.70%, every 0.125 50.08%, every 0.05 53.24%. The heavy line is the largest of the six at each z.
Where it is used
11 essays draw this figure, each at the numbers its own argument is about, so the same picture answers 11 different questions.
- Choosing n after looking Designs that change while they run
- When the looking happens Stopping rules
- A p-value that is not flat is not a p-value Tests, and the second number
- Spending the error rate Stopping rules
- The test with no table Series that move together
- Dropping the losers Designs that change while they run
- The winner's curse Tests, and the second number
- The effect a stopped trial reports Stopping rules
- The outcomes a trial could have stopped with Stopping rules
- A boundary for giving up Stopping rules
- A look the trend asked for Stopping rules