Series

Forking — the series

4 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. The false-positive rate against the number of analyses, on pure noise. The data has no effect in it. Each analysis is correct and each p-value is honest. With 20 correlated outcomes available, something reaches p below 0.05 58% of the time.

    Twenty analyses of nothing

    Twenty honest, correct analyses of data with no effect in it find something significant 57% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

    part 3 · testing
  2. What 20 analyses of one dataset are worth. The threshold giving a 5% family-wise error rate, read back as a number of independent analyses. At no correlation it is 20.05; at 0.6 it is 11.37; at 0.95 it is 2.58. Bonferroni divides by 20 throughout.

    How many analyses there really were

    Bonferroni divides by twenty because twenty analyses were run. Twenty analyses of one dataset are worth 11.37 independent ones at a correlation of 0.6 and 2.58 at 0.95, and the threshold that controls exactly the same error rate is measurable rather than assumed.

    part 4 · paths
  3. Naming the analysis in advance, against correcting for all 20 of them. The prespecified analysis detects an effect that is in it 52% of the time at two standard errors and an effect elsewhere 5% of the time. The corrected slate detects it 23% of the time wherever it is. The two are worth the same when the chance of having named the right analysis is 38% — and that figure rises to 91% at four standard errors.

    What naming it in advance costs

    Preregistration is argued for as free. Against an effect of two standard errors hiding in one of twenty analyses, naming the right one detects it 51.5% of the time and naming the wrong one detects it 4.7% of the time; correcting all twenty detects it 22.5% wherever it is. The two are worth the same when the chance of having named correctly is 38%.

    part 5 · paths
  4. What the family-wise correction does to the effect it lets through. At two standard errors the estimate that clears an uncorrected 5% threshold averages 1.35 times the truth, and the one that clears the family-wise threshold averages 1.69 times it. The correction fixes the error rate by demanding a larger estimate, and a larger estimate is a more selected one.

    The correction that makes the estimate worse

    Correcting for twenty analyses repairs the p-value by demanding a larger statistic, and a larger statistic is a more selected one. At two standard errors the surviving estimate averages 1.35 times the truth before the correction and 1.69 times it after — so the honest error rate is bought with a more inflated effect.

    part 6 · paths

All series