Field

An interval read beside something else

A 95% interval is almost never read alone. Read against a replication's estimate it captures an equal-sized one 83.42% of the time, because both estimates are uncertain, and ten replications of one study miss together: none of them lands outside 31.40% of the time against 16.32% if they were independent. Read against a second interval, two that just touch mark a difference with p = 0.0056 rather than 0.05, standard-error bars that touch mark p = 0.157, and in a figure of ten identical groups some pair of 95% intervals fails to overlap one time in seven. Read against a second verdict, two studies of one effect at 50% power disagree about significance half the time, and when they do the difference between them is significant in 9.75% of cases.
Where a replication's estimate lands against a 95% interval, replication the same size. The chance that a 95% interval contains a replication's estimate is 95.00% when the original landed on the truth, 82.99% one standard error away and 48.40% two away. Averaged over where originals land it is 83.42%, and 5.00% of originals capture a replication less than half the time.

Five times in six

A 95% interval is read as a 95% chance that a replication's estimate will land inside it. With the spread known and a replication of the same size, the chance is 83.42% โ€” five times in six โ€” because both estimates are uncertain. An original that landed two standard errors from the truth captures a replication 48.40% of the time; a replication a tenth the size lands inside 44.54% of the time; and among significant originals from studies with 17% power, 66.94%.

Two 95% intervals 2.772 standard errors of the difference apart, standard errors in the ratio 1. The intervals are separated, and the test of the difference gives p = 0.0056. Two 95% intervals with equal standard errors just touch at p = 0.0056.

Two intervals that overlap

Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.

Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time โ€” and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

All essays