Field

Shape, and what it does to a two-sample test

A t interval survives a skewed source in total and not in its tails. On an exponential source at 120 observations a 95% interval covers 94.81% while missing below the mean on 4.08% of samples and above it on 1.11%. The pooled two-sample test's size runs from 0.55% to 18.91% across forty units with a true null in every cell, and Welch's stays between 4.63% and 5.51% — on normal populations. On skewed ones Welch's tails are decided by one number, the skewness of the difference of the two means: two identical exponentials balance at twenty and twenty and reject low on 7.16% of samples at eight and thirty-two. And the side a decision reads needs its own repair: widened until its total coverage is exactly 95%, a t interval's upper limit on thirty exponential observations is still exceeded on 4.69% of samples, where Hall's transformation takes it to 3.31%.
Which side a 95% t interval misses on, exponential source. Both tails should be 2.5%. At 8 observations the interval falls short of the mean on 9.75% of samples and overshoots on 0.31%. At 500 they are 3.31% and 2.05%, and the total is 5.36% — which a coverage table reports as very nearly right.

Where the two tails disagree

A 95% t interval on an exponential source at 120 observations covers 94.81%, which reads as very nearly right. It misses below the mean on 4.08% of samples and above on 1.11% — one tail 63% too heavy and the other 56% too light, and the total is the statistic that hides it.

The pooled two-sample test's size, with a true null everywhere. Forty units split between two groups, with the second group's variance a stated multiple of the first's, and the two population means equal. A 5% test should reject 5% of the time. The pooled test runs from 0.55% to 18.91% across this region; Welch's runs from 4.63% to 5.51%.

A degrees of freedom that is not a count

The pooled two-sample test's size runs from 0.55% to 18.91% across forty units split five ways against five variance ratios, with a true null in every cell. Welch's runs from 4.63% to 5.51% — bought with a degrees of freedom that is a function of the data, not an integer, and not a count of anything.

Welch's test on skewed groups: the low and high rejection rates in every cell, with equal means throughout. Each cell should read 2.5 / 2.5. Two identical exponentials at 20 and 20 read 2.04 / 2.22; the worst cell, a wide exponential against a normal at 8 and 32, reads 9.79 / 0.47.

The skewness of a difference

Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.

Four 95% intervals for the mean of 30 exponential observations, split by the side they miss on. Both bars should read 2.5%. The t interval misses below the mean on 6.38% of samples. Widened until its total is exactly 5%, it misses below on 4.69% and above on 0.30%. Hall's transformation misses below on 3.31% and above on 1.96%.

The side a bound is read from

On thirty exponential observations the upper limit of a 95% t interval is exceeded by the true mean 6.38% of the time, against the 2.5% a safety margin set from it assumes. Widen the interval until its total coverage is exactly 95% and the upper limit is still exceeded 4.69% of the time. A symmetric repair fixes the number that is reported and not the one that is used; Hall's transformation, which bends the interval, takes the same rate to 3.31%.

All essays