Theme

The thread: Reversals that are nobody's mistake

Simpson's paradox, regression to the mean and the winner's curse all arise from correct arithmetic applied honestly. Each is shown as a region of a parameter space rather than one famous table, so the questions of how often and how large have answers.
fraction of group A given the treatmentgroup B0.00.51.032% reverseshaded: the overall comparison reversesthe rates are identical everywhere here Reversals that are not errors

Simpson's reversal is a region, not a table

The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.

0.8000.8500.9000.950255075100sample sizecoverage at a true proportion of 0.15n = 19n = 20WilsonWaldexact coverage at every n from 10 to 120more data is not automatically better here Intervals, counted

More data is not monotonically better

Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.

00.2500.5000.7501how common the condition ischance a positive result is true1 in 10,0001 in 1,0001 in 1001 in 101 in 11.8%66.7%one test, every prevalencethe base rate outweighs the test Reversals that are not errors

What a positive test is worth

A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.

02505007501e+3-0.50000.5001effect the study reportsstudiesthe truth, 0.3what gets published, 0.61shaded: reached p below 0.05power 20%, inflation 2.04× Tests, and the second number

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

first measurementsecond+0.55-0.59predicted 0.64 from the correlation alonenobody was treated Reversals that are not errors

Regression to the mean

Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.

00.2000.4000.6000.8005101520analyses available to the researcherchance of finding something significantthe nominal 5%if the analyses were independentcorrelated, as real ones areno effect present anywhereevery individual analysis is correct Tests, and the second number

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 58% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

the sample mean93.6%nominal 95%the sample maximum0.0%nominal 95%3,000 samples of 40the same procedure, two statisticsresampling cannot see past the data Intervals, counted

Where the bootstrap lies

Resampling is the most generally useful trick in the subject and it has a failure mode that is easy to state: it cannot see past the data. For a statistic that lives at the edge of the sample, coverage collapses from 95% to almost nothing.

All themes