Two-sample test — where it appears
Named by 5 essays across 3 fields — each of them below, with the objects they name alongside it.
A degrees of freedom that is not a count
The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.
Two intervals that overlap
Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.
The skewness of a difference
Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.
Significant in one, not in the other
Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.
An outcome cut in two
Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.
Named alongside it
The objects these essays reach for when they reach for this one.
Sample sizeStatistical powerBehrens–FisherError ratep-valueWelch testAllocationCentral limit theoremConfidence intervalCorrelationDegrees of freedomDichotomisation