Concept

Two-sample test — where it appears

A test of whether two groups share a mean, built from the difference of their estimates and its standard error. Pooled, Welch and proportion versions differ in what they assume about the groups' spreads and in which side their errors fall on.

Named by 5 essays across 3 fields — each of them below, with the objects they name alongside it.

The pooled two-sample test's size, with a true null everywhere. Forty units split between two groups, with the second group's variance a stated multiple of the first's, and the two population means equal. A 5% test should reject 5% of the time. The pooled test runs from 0.55% to 18.91% across this region; Welch's runs from 4.63% to 5.51%.

A degrees of freedom that is not a count

The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.

tails · Student
Two 95% intervals 2.772 standard errors of the difference apart, standard errors in the ratio 1. The intervals are separated, and the test of the difference gives p = 0.0056. Two 95% intervals with equal standard errors just touch at p = 0.0056.

Two intervals that overlap

Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.

alongside · Repetition
Welch's test on skewed groups: the low and high rejection rates in every cell, with equal means throughout. Each cell should read 2.5 / 2.5. Two identical exponentials at 20 and 20 read 2.04 / 2.22; the worst cell, a wide exponential against a normal at 8 and 32, reads 9.79 / 0.47.

The skewness of a difference

Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.

tails · Student
Two studies of the same effect, z statistics with mean 1.96: where one is significant and the other is not. Five hundred pairs. With the true effect identical in both, exactly one of the two is significant in 50.0% of pairs; among those, the difference between the two is significant in 9.7%. The dashed lines mark 1.96 on each axis; the diagonal lines mark a significant difference.

Significant in one, not in the other

Two studies of exactly the same effect, each with 50% power, disagree about significance half the time — and when they do, the test of the difference between them is significant in 9.75% of cases. A p of 0.01 beside a p of 0.20 is a difference with p = 0.36. Among four subgroups sharing one effect, at least one significant and one not happens 87.5% of the time, and the test that would tell a real difference apart needs four times the sample the effect itself needed.

alongside · Repetition
How much of a normal outcome's information survives cutting it into two, by where the cut is. For a small shift, a cut at the mean keeps 63.7% of the information, so the trial needs 1.57 times the sample. A cut at the top tenth keeps 34.2% and needs 2.92 times; a cut two standard deviations out keeps 13.1%.

An outcome cut in two

Replacing a measured outcome with whether it crossed a threshold keeps 63.7% of the information when the cut is at the mean, 34.2% at the top tenth and 13.1% two standard deviations out. A trial that needs 63 patients per arm on the measured outcome needs 102 cut at the mean and 185 cut at one and a half standard deviations. The responder rates that result read as a share of patients who respond — 50.0% against 69.1% — when every patient moved by the same amount; and a cut chosen after looking turns a 5% test into a 17.7% one.

planned · Power

Named alongside it

The objects these essays reach for when they reach for this one.

Sample sizeStatistical powerBehrens–FisherError ratep-valueWelch testAllocationCentral limit theoremConfidence intervalCorrelationDegrees of freedomDichotomisation

All concepts