Shape, and what it does to a two-sample test

The skewness of a difference

Welch's test holds its size to within half a point when both groups are normal. Give both groups the same skewed population and it still balances at twenty and twenty — and at eight and thirty-two it rejects low on 7.16% of samples and high on 0.66%. One number decides which: the skewness of the difference of the two means, which ranks twenty-five cells by their imbalance with a correlation of 0.997.

Worth reading first: The correction for not knowing the spread.

A degrees of freedom that is not a count measured Welch’s test across forty units split five ways and five variance ratios, and found its size between 4.63% and 5.51% everywhere. It ended by naming what the measurement had assumed away: every cell was two normal populations. Where the two tails disagree had shown what skew does to a one-sample interval — the misses pile onto one side and the total hides it — and the two failures had never been measured together.

This essay measures them together. The answer is tidier than either failure alone, because the whole region turns out to be governed by one computable number.

Welch's test on skewed groups: the low and high rejection rates in every cell, with equal means throughout. Each cell should read 2.5 / 2.5. Two identical exponentials at 20 and 20 read 2.04 / 2.22; the worst cell, a wide exponential against a normal at 8 and 32, reads 9.79 / 0.47.
Fig. 1 Welch’s two-sided 5% test on five pairings of populations and five splits of forty units, with the two population means equal in every cell. Each cell prints how often the test comes out significant because the first mean looked smaller, and how often because it looked larger. Both should read 2.5.

The region, split by side

Every cell holds a true null: the two populations have the same mean, so every rejection is an error, and the errors are split by the side they fall on. Twenty thousand pairs of samples per cell.

pairing 8 and 32 14 and 26 20 and 20 26 and 14 32 and 8
two identical exponentials 7.16 / 0.66 4.00 / 1.35 2.04 / 2.22 1.08 / 4.05 0.49 / 7.39
exponentials, the first twice as wide 9.65 / 0.41 7.12 / 0.60 4.98 / 0.90 3.35 / 1.50 1.64 / 2.97
a wide exponential against a normal 9.79 / 0.47 7.33 / 0.57 5.75 / 0.81 4.42 / 1.19 3.48 / 1.81
an exponential against its mirror image 8.24 / 0.58 6.28 / 0.84 5.68 / 0.84 6.16 / 0.69 8.19 / 0.51
a narrow exponential against a wide normal 4.49 / 1.23 3.33 / 1.79 2.90 / 2.15 2.70 / 2.38 2.55 / 2.74

The first row is the case people have in mind when they say a two-sample test is protected against skew by the symmetry of the comparison. It is protected at twenty and twenty, where the two tails read 2.04 and 2.22. Move to eight and thirty-two, with exactly the same two populations, and the low tail is 7.16% and the high tail 0.66% — a two-sided 5% test that is really a 7.8% test, with eleven of every twelve of its false positives saying the first group is lower.

The worst cells are not the ones with the most skewed populations. A wide exponential of eight against a normal of thirty-two rejects low on 9.79% of samples. An exponential against its own mirror image — two populations with skewness of opposite sign — is bad at every split, including the even one, where it reads 5.68 / 0.84. And a narrow exponential against a wide normal is nearly well-behaved, although one of its groups is as skewed as anything in the table.

Those five rows look like five separate facts about skew. They are one fact.

The number that decides the cell

The two-sample statistic is a studentised difference of two means, xˉ1xˉ2\bar x_1 - \bar x_2. Its sampling distribution is skewed if and only if that difference is skewed, and the third cumulant of a difference of independent means is the first mean’s third cumulant minus the second’s. Standardised:

γD  =  γ1σ13/n12    γ2σ23/n22(σ12/n1+σ22/n2)3/2\gamma_D \;=\; \frac{\gamma_1\,\sigma_1^3/n_1^2 \;-\; \gamma_2\,\sigma_2^3/n_2^2}{\bigl(\sigma_1^2/n_1 + \sigma_2^2/n_2\bigr)^{3/2}}

with γ\gamma each population’s skewness, σ\sigma its standard deviation and nn its sample size. Nothing in it comes from data: it is a property of the two populations and the design.

The gap between Welch's two tails against the skewness of the difference, twenty-five cellsFive pairings of populations at five splits each. Whatever makes the difference skewed — unequal sizes, unequal spreads, opposite skews — the cells fall on one rising curve, with a rank correlation of 0.997.-0.05000.0500.100-0.500-0.25000.2500.500skewness of the difference of the meanslow rejections minus hightwo identical exponentialsexponentials, the first twice as widea wide exponential against a normalan exponential against its mirror imagea narrow exponential against a wide normal20,000 pairs per cell, equal meansone number, whatever made it
Fig. 2 Every cell of the region, placed by the skewness of the difference of its two means and by the gap between its low and high rejection rates. Five pairings, marked separately, fall on one rising curve. The slider switches the vertical axis to the total rejection rate.

Across the twenty-five cells the rank correlation between γD\gamma_D and the gap between the tails is 0.997. Whatever made the difference skewed — unequal sizes of identical populations, unequal spreads, opposite skews, one skewed group against a symmetric one — the cells line up by γD\gamma_D alone.

Each of the rows that looked odd is explained by the formula in a line.

Two identical exponentials give γ1=γ2\gamma_1 = \gamma_2 and σ1=σ2\sigma_1 = \sigma_2, so the numerator is γσ3(1/n121/n22)\gamma\sigma^3(1/n_1^2 - 1/n_2^2): exactly zero at equal sizes and 0.47 at eight and thirty-two. The symmetry people credit to the comparison is a symmetry of the design, and unequal allocation destroys it.

An exponential against its mirror image gives γ2=γ1\gamma_2 = -\gamma_1, so the two terms add rather than cancel: 0.32 at equal sizes and 0.54 at the extremes. That row is bad everywhere because nothing in the design can make a sum of two positive numbers zero.

A narrow exponential against a wide normal has one skewed term, but the skewed group’s spread is half the normal’s, so its σ3\sigma^3 is an eighth of the normal’s and the denominator is dominated by the normal’s variance. γD\gamma_D is 0.25 at worst and 0.005 at thirty-two and eight. The skewed group is swamped by the variance of the symmetric one.

The size as well as the balance

The table’s totals move with γD|\gamma_D| too, though less cleanly.

Welch's total size against the size of the skewness of the difference, twenty-five cells. The total rejection rate rises with the magnitude of the skewness too, from about 4.3% where it is zero to over 10% where it is 0.65, with a rank correlation of 0.956.
Fig. 3 The same twenty-five cells, placed by the size of the skewness of the difference and by Welch’s total rejection rate. The dashed line is the 5% the test claims.

Where γD\gamma_D is zero the test is slightly conservative — 4.26% for two identical exponentials at twenty and twenty — and where it is 0.65 the total passes 10%. The rank correlation between γD|\gamma_D| and the total is 0.956. So a skewed difference does two things at once: it unbalances the tails, which is nearly perfectly predicted, and it inflates their sum, which is predicted well.

The inflation is worth stating as an error rate a reader would recognise. A two-sided 5% Welch test comparing a wide skewed group of eight with a normal group of thirty-two rejects a true null on 10.27% of samples, double its label — on a design that the normal-population grid in the essay on Welch’s degrees of freedom put comfortably inside half a point. The failure is not in the degrees of freedom. It is in the reference distribution being symmetric when the statistic is not.

Equal sizes are a defence against one case only

The first row makes it tempting to conclude that a balanced design is the protection, and for that row it is. The rest of the table says how narrow the protection is.

At twenty and twenty, with the design as symmetric as it can be made, the five pairings read:

pairing at 20 and 20 low high total
two identical exponentials 2.04% 2.22% 4.26%
exponentials, the first twice as wide 4.98% 0.90% 5.88%
a wide exponential against a normal 5.75% 0.81% 6.57%
an exponential against its mirror image 5.68% 0.84% 6.53%
a narrow exponential against a wide normal 2.90% 2.15% 5.05%

Equal sizes cancel the skewness of the difference only when the two populations have the same skewness and the same spread. The moment the spreads differ, the wider group’s σ3\sigma^3 dominates the numerator whatever the sizes are, and four of the five rows are unbalanced at the even split.

The pairing that matters most in practice is the third. Where unequal variances come from is usually a treatment that does more to some units than to others, and a treatment that adds a variable amount to a positive quantity adds skew in the same stroke — so the more variable group is typically the more skewed one, and the control group is closer to symmetric. That is a wide exponential against a narrow near-normal, and at the most balanced design available it rejects low on 5.75% of samples, more than twice its allotment, with seven of every eight false positives on the same side.

The side a one-sided comparison reads

A two-sided test’s imbalance is an irritation; a one-sided test’s is its error rate, and one-sided two-sample comparisons are what non-inferiority trials, safety comparisons and cost-effectiveness claims are built from.

A one-sided 2.5% Welch test of “the first group’s mean is not lower” rejects when the statistic falls below the lower critical value — exactly the low column. So the one-sided test’s size is that column read directly:

  • two identical exponentials at eight and thirty-two: 7.16%, nearly three times its 2.5% label;
  • a wide exponential against a normal at eight and thirty-two: 9.79%, nearly four times;
  • the same pair at twenty and twenty: 5.75%, more than twice.

The test in the other direction is correspondingly timid — 0.47% where 2.5% was promised in the worst cell — and has lost most of its power without anyone deciding to spend it. Nothing in a two-sided size study reveals either number, because the two errors are added before they are reported, which is the same habit that hid the one-sample imbalance behind a coverage of 94.81%.

The direction is fixed by the sign of γD\gamma_D. When it is positive the comparison is too quick to call the first group lower, and a non-inferiority claim about the first group — “it is not worse” framed as “its mean is not lower” — is too hard to make. When it is negative the reverse. So the same data can make a one-sided claim too easy in one framing and too hard in the other, and which is which is decided by the design rather than by the effect.

A design read through the formula

Consider a comparison of costs per patient between a new pathway and the usual one, forty patients in all. Costs are right-skewed, and the new pathway’s costs vary more because some patients use much more of it than others. Recruitment to the new pathway is slower, so the design ends up with eight on the new pathway and thirty-two on the usual one.

That is the second row’s leftmost cell with the roles as described: a skewed group twice as wide, with a quarter of the units. γD\gamma_D is 0.64. A two-sided 5% Welch test rejects a true null on about ten per cent of such trials, and almost every one of those rejections says the new pathway is cheaper. A trial designed to be fair in both directions is, before any data arrives, a trial that favours one conclusion.

The formula says what would have helped. Units in proportion to the spreads raised to the power three halves would put about thirty in the more variable arm, which recruitment would not allow. Short of that, the cells show that moving from eight to fourteen patients on the new pathway takes the low tail from 9.65% to 7.12% — still wrong — and that only a design near thirty makes the two tails agree. A reader of the published trial can do none of this arithmetic without the allocation, the two sample spreads and some idea of the outcome’s skew, and all three are usually reported.

Why the limit theorem does not rescue it

Each sample mean becomes normal as its sample grows — that is the central limit theorem — and a difference of two normal means is normal. So the skewness of the difference must vanish eventually, and it does: scale both groups up together and the numerator of γD\gamma_D falls like 1/n21/n^2 while the denominator falls like 1/n3/21/n^{3/2}, so γD\gamma_D falls like 1/n1/\sqrt{n}.

The same 1/n1/\sqrt{n} is what governs the one-sample imbalance, and it has the same consequence. The rate is slow, and it is the rate of the smaller group. At eight and thirty-two the eight-unit group’s term dominates the numerator, so doubling the larger group barely moves γD\gamma_D while doubling the smaller one cuts its term in the numerator to a quarter. The limit theorem does repair the test; it repairs it at the pace of the group with the fewest observations, and an unbalanced design is exactly one whose smallest group is small.

This is also why the correction for an estimated spread is no help. Welch’s degrees of freedom, like Student’s, widens the reference distribution symmetrically. A symmetric widening changes how often the test comes out significant in total and does nothing about which side the rejections fall on — which is the subject of the essay on the side a bound is read from.

The split that cancels it

The formula also says where to put the units. For two populations of the same shape — the same skewness — with spreads in the ratio r=σ1/σ2r = \sigma_1/\sigma_2, the numerator vanishes when

n1n2  =  r3/2\frac{n_1}{n_2} \;=\; r^{3/2}

That is not the allocation that minimises the standard error of the difference. Neyman’s allocation puts units in proportion to the spreads, n1/n2=rn_1/n_2 = r, and for two populations of different spread the two rules disagree.

Where Welch's two tails balance for two exponentials, one 2 times as wide as the other. The gap between the low and high rejection rates, as forty units are split between the groups. The skewness of the difference is zero at 29.6 units in the wider group, where the split is 2.83 to one; the split that minimises the standard error puts 26.7 there.
Fig. 4 Two exponentials, the first twice as wide as the second, with forty units split between them. The curve is the measured gap between Welch’s low and high rejection rates at each split. The solid rule marks where the skewness of the difference is zero, the dashed rule Neyman’s allocation.

With the first population twice as wide, Neyman’s rule puts 26.7 of forty units in the wider group and the cancelling rule puts 29.6 there. The measured gap between the tails crosses zero where the formula says: +0.83 points at twenty-eight units, −0.28 at thirty. At twenty-six units, the nearest split to Neyman’s, the gap is still 1.84 points — a low tail of 3.35% against a high one of 1.50%.

With spreads in the ratio 1.5 the rules put 24.0 and 25.9 units in the wider group, and the measured gap changes sign between twenty-four and twenty-six: +0.75 points at one, −0.03 at the other. At equal spreads they coincide at twenty, and so does the crossing.

The two rules answer different questions, and the difference between them is small in units — three of forty at a spread ratio of two. What makes it worth knowing is that one of them protects the test’s error rate and the other its precision, and a design chosen by the second on skewed outcomes has a two-sided test that is quietly one-sided.

What the table does not show

Everything above is drawn with the populations known, and a real comparison knows neither γ\gamma nor σ\sigma. Two things follow, and one of them is reassuring.

The reassuring one is that the direction is predictable without the numbers. On a quantity that is right-skewed in both groups — a duration, a cost, a concentration — γD\gamma_D has the sign of the term with the larger σ3/n2\sigma^3/n^2, which is the smaller or the more variable group. A design that puts fewer units in the more variable arm, which is common when that arm is the expensive one, produces rejections that say that arm is lower more often than they should.

The less reassuring one is that the sample skewness, which would be needed to compute γD\gamma_D from data, is biased low and very noisy at the group sizes where it matters, so a data-driven check tends to report less skew than there is. The design-time version of the rule, with the skewness taken from what is known about the outcome, is more trustworthy than the analysis-time one.

The skewness of the difference of the two means, in every cell of the same region. The skewness of x̄₁ − x̄₂, computed from the populations and the sizes alone. It is zero where two identical populations meet at equal sizes and 0.65 in the cell whose tails are furthest apart.
Fig. 5 The same region, with each cell printing the skewness of the difference of the two means instead of the two rejection rates. The cells the hero figure shaded darkest are the ones with the largest values here.

What the region establishes, and what it does not

The skewness of the difference ranks the cells by the gap between Welch’s tails, with a rank correlation of 0.997 across all twenty-five. Stated over the whole region rather than per pairing, because the claim is that the mechanism behind a cell does not matter once γD\gamma_D is known, and only a comparison across mechanisms can say that.

Where two identical skewed populations meet at equal sizes, γD\gamma_D is exactly zero and the tails balance, to within half a point. This is the case that makes the formula’s cancellation visible rather than assumed.

The measured gap changes sign at the split where γD\gamma_D does, to within one step of two units, at spread ratios of one, one and a half and two.

What does not survive is Welch’s test used on skewed groups of unequal size as though its level were two-sided. At eight and thirty-two, with a wide exponential against a normal, the low side alone rejects 9.79% of the time against the 2.5% it is allotted.

Every rate in the tables is a count over twenty thousand pairs of samples, so each carries its own sampling error: about 0.11 points for a rate near 2.5% and about 0.2 points for one near 10%. Differences between cells smaller than a few tenths of a point — the 2.04 against 2.22 at the balanced split, for instance — are inside that error and are not read as findings. The sign change at the cancelling split is read from differences of the same size, which is why it is claimed to within one step of the grid rather than at a point, and why the crossing is checked at three spread ratios rather than one. The same seed is used in every cell, so neighbouring cells share their luck; that makes the curve smoother than independent counts would and does not bias any single cell.

Not claimed: that γD\gamma_D predicts the total size as tightly as it predicts the imbalance — it does not, and the rank correlation of 0.956 is reported beside the 0.997 for that reason. Nor anything about populations with heavy tails and no skew, which the one-sample essay found do not unbalance an interval, and which have not been measured here in the two-sample case.

Still open: a sample’s own estimate of the skewness of its difference

The formula is exact for known populations. The practical version replaces γ1\gamma_1, γ2\gamma_2, σ1\sigma_1 and σ2\sigma_2 with sample estimates and reports γ^D\hat\gamma_D beside the test. Whether that estimate is good enough to flag the dangerous cells — whether, at eight and thirty-two, γ^D\hat\gamma_D is usually large when γD\gamma_D is 0.47 — is a measurement this region can make and has not. The sample skewness’s downward bias suggests it will under-flag, and the question is by how much.

The other open question is the repair. A skew-corrected reference distribution for the one-sample mean exists and is measured on the side a bound is read from. Its two-sample analogue corrects with γ^D\hat\gamma_D in place of the one-sample skewness, and whether it balances Welch’s tails across this region — including the mirrored pairing, which no allocation can fix — has not been counted.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AllocationBehrens–FisherCentral limit theoremError rateNeyman allocationSkewnessTwo-sample testWelch test