The skewness of a difference
Worth reading first: The correction for not knowing the spread.
A degrees of freedom that is not a count measured Welch’s test across forty units split five ways and five variance ratios, and found its size between 4.63% and 5.51% everywhere. It ended by naming what the measurement had assumed away: every cell was two normal populations. Where the two tails disagree had shown what skew does to a one-sample interval — the misses pile onto one side and the total hides it — and the two failures had never been measured together.
This essay measures them together. The answer is tidier than either failure alone, because the whole region turns out to be governed by one computable number.
The region, split by side
Every cell holds a true null: the two populations have the same mean, so every rejection is an error, and the errors are split by the side they fall on. Twenty thousand pairs of samples per cell.
| pairing | 8 and 32 | 14 and 26 | 20 and 20 | 26 and 14 | 32 and 8 |
|---|---|---|---|---|---|
| two identical exponentials | 7.16 / 0.66 | 4.00 / 1.35 | 2.04 / 2.22 | 1.08 / 4.05 | 0.49 / 7.39 |
| exponentials, the first twice as wide | 9.65 / 0.41 | 7.12 / 0.60 | 4.98 / 0.90 | 3.35 / 1.50 | 1.64 / 2.97 |
| a wide exponential against a normal | 9.79 / 0.47 | 7.33 / 0.57 | 5.75 / 0.81 | 4.42 / 1.19 | 3.48 / 1.81 |
| an exponential against its mirror image | 8.24 / 0.58 | 6.28 / 0.84 | 5.68 / 0.84 | 6.16 / 0.69 | 8.19 / 0.51 |
| a narrow exponential against a wide normal | 4.49 / 1.23 | 3.33 / 1.79 | 2.90 / 2.15 | 2.70 / 2.38 | 2.55 / 2.74 |
The first row is the case people have in mind when they say a two-sample test is protected against skew by the symmetry of the comparison. It is protected at twenty and twenty, where the two tails read 2.04 and 2.22. Move to eight and thirty-two, with exactly the same two populations, and the low tail is 7.16% and the high tail 0.66% — a two-sided 5% test that is really a 7.8% test, with eleven of every twelve of its false positives saying the first group is lower.
The worst cells are not the ones with the most skewed populations. A wide exponential of eight against a normal of thirty-two rejects low on 9.79% of samples. An exponential against its own mirror image — two populations with skewness of opposite sign — is bad at every split, including the even one, where it reads 5.68 / 0.84. And a narrow exponential against a wide normal is nearly well-behaved, although one of its groups is as skewed as anything in the table.
Those five rows look like five separate facts about skew. They are one fact.
The number that decides the cell
The two-sample statistic is a studentised difference of two means, . Its sampling distribution is skewed if and only if that difference is skewed, and the third cumulant of a difference of independent means is the first mean’s third cumulant minus the second’s. Standardised:
with each population’s skewness, its standard deviation and its sample size. Nothing in it comes from data: it is a property of the two populations and the design.
Across the twenty-five cells the rank correlation between and the gap between the tails is 0.997. Whatever made the difference skewed — unequal sizes of identical populations, unequal spreads, opposite skews, one skewed group against a symmetric one — the cells line up by alone.
Each of the rows that looked odd is explained by the formula in a line.
Two identical exponentials give and , so the numerator is : exactly zero at equal sizes and 0.47 at eight and thirty-two. The symmetry people credit to the comparison is a symmetry of the design, and unequal allocation destroys it.
An exponential against its mirror image gives , so the two terms add rather than cancel: 0.32 at equal sizes and 0.54 at the extremes. That row is bad everywhere because nothing in the design can make a sum of two positive numbers zero.
A narrow exponential against a wide normal has one skewed term, but the skewed group’s spread is half the normal’s, so its is an eighth of the normal’s and the denominator is dominated by the normal’s variance. is 0.25 at worst and 0.005 at thirty-two and eight. The skewed group is swamped by the variance of the symmetric one.
The size as well as the balance
The table’s totals move with too, though less cleanly.
Where is zero the test is slightly conservative — 4.26% for two identical exponentials at twenty and twenty — and where it is 0.65 the total passes 10%. The rank correlation between and the total is 0.956. So a skewed difference does two things at once: it unbalances the tails, which is nearly perfectly predicted, and it inflates their sum, which is predicted well.
The inflation is worth stating as an error rate a reader would recognise. A two-sided 5% Welch test comparing a wide skewed group of eight with a normal group of thirty-two rejects a true null on 10.27% of samples, double its label — on a design that the normal-population grid in the essay on Welch’s degrees of freedom put comfortably inside half a point. The failure is not in the degrees of freedom. It is in the reference distribution being symmetric when the statistic is not.
Equal sizes are a defence against one case only
The first row makes it tempting to conclude that a balanced design is the protection, and for that row it is. The rest of the table says how narrow the protection is.
At twenty and twenty, with the design as symmetric as it can be made, the five pairings read:
| pairing at 20 and 20 | low | high | total |
|---|---|---|---|
| two identical exponentials | 2.04% | 2.22% | 4.26% |
| exponentials, the first twice as wide | 4.98% | 0.90% | 5.88% |
| a wide exponential against a normal | 5.75% | 0.81% | 6.57% |
| an exponential against its mirror image | 5.68% | 0.84% | 6.53% |
| a narrow exponential against a wide normal | 2.90% | 2.15% | 5.05% |
Equal sizes cancel the skewness of the difference only when the two populations have the same skewness and the same spread. The moment the spreads differ, the wider group’s dominates the numerator whatever the sizes are, and four of the five rows are unbalanced at the even split.
The pairing that matters most in practice is the third. Where unequal variances come from is usually a treatment that does more to some units than to others, and a treatment that adds a variable amount to a positive quantity adds skew in the same stroke — so the more variable group is typically the more skewed one, and the control group is closer to symmetric. That is a wide exponential against a narrow near-normal, and at the most balanced design available it rejects low on 5.75% of samples, more than twice its allotment, with seven of every eight false positives on the same side.
The side a one-sided comparison reads
A two-sided test’s imbalance is an irritation; a one-sided test’s is its error rate, and one-sided two-sample comparisons are what non-inferiority trials, safety comparisons and cost-effectiveness claims are built from.
A one-sided 2.5% Welch test of “the first group’s mean is not lower” rejects when the statistic falls below the lower critical value — exactly the low column. So the one-sided test’s size is that column read directly:
- two identical exponentials at eight and thirty-two: 7.16%, nearly three times its 2.5% label;
- a wide exponential against a normal at eight and thirty-two: 9.79%, nearly four times;
- the same pair at twenty and twenty: 5.75%, more than twice.
The test in the other direction is correspondingly timid — 0.47% where 2.5% was promised in the worst cell — and has lost most of its power without anyone deciding to spend it. Nothing in a two-sided size study reveals either number, because the two errors are added before they are reported, which is the same habit that hid the one-sample imbalance behind a coverage of 94.81%.
The direction is fixed by the sign of . When it is positive the comparison is too quick to call the first group lower, and a non-inferiority claim about the first group — “it is not worse” framed as “its mean is not lower” — is too hard to make. When it is negative the reverse. So the same data can make a one-sided claim too easy in one framing and too hard in the other, and which is which is decided by the design rather than by the effect.
A design read through the formula
Consider a comparison of costs per patient between a new pathway and the usual one, forty patients in all. Costs are right-skewed, and the new pathway’s costs vary more because some patients use much more of it than others. Recruitment to the new pathway is slower, so the design ends up with eight on the new pathway and thirty-two on the usual one.
That is the second row’s leftmost cell with the roles as described: a skewed group twice as wide, with a quarter of the units. is 0.64. A two-sided 5% Welch test rejects a true null on about ten per cent of such trials, and almost every one of those rejections says the new pathway is cheaper. A trial designed to be fair in both directions is, before any data arrives, a trial that favours one conclusion.
The formula says what would have helped. Units in proportion to the spreads raised to the power three halves would put about thirty in the more variable arm, which recruitment would not allow. Short of that, the cells show that moving from eight to fourteen patients on the new pathway takes the low tail from 9.65% to 7.12% — still wrong — and that only a design near thirty makes the two tails agree. A reader of the published trial can do none of this arithmetic without the allocation, the two sample spreads and some idea of the outcome’s skew, and all three are usually reported.
Why the limit theorem does not rescue it
Each sample mean becomes normal as its sample grows — that is the central limit theorem — and a difference of two normal means is normal. So the skewness of the difference must vanish eventually, and it does: scale both groups up together and the numerator of falls like while the denominator falls like , so falls like .
The same is what governs the one-sample imbalance, and it has the same consequence. The rate is slow, and it is the rate of the smaller group. At eight and thirty-two the eight-unit group’s term dominates the numerator, so doubling the larger group barely moves while doubling the smaller one cuts its term in the numerator to a quarter. The limit theorem does repair the test; it repairs it at the pace of the group with the fewest observations, and an unbalanced design is exactly one whose smallest group is small.
This is also why the correction for an estimated spread is no help. Welch’s degrees of freedom, like Student’s, widens the reference distribution symmetrically. A symmetric widening changes how often the test comes out significant in total and does nothing about which side the rejections fall on — which is the subject of the essay on the side a bound is read from.
The split that cancels it
The formula also says where to put the units. For two populations of the same shape — the same skewness — with spreads in the ratio , the numerator vanishes when
That is not the allocation that minimises the standard error of the difference. Neyman’s allocation puts units in proportion to the spreads, , and for two populations of different spread the two rules disagree.
With the first population twice as wide, Neyman’s rule puts 26.7 of forty units in the wider group and the cancelling rule puts 29.6 there. The measured gap between the tails crosses zero where the formula says: +0.83 points at twenty-eight units, −0.28 at thirty. At twenty-six units, the nearest split to Neyman’s, the gap is still 1.84 points — a low tail of 3.35% against a high one of 1.50%.
With spreads in the ratio 1.5 the rules put 24.0 and 25.9 units in the wider group, and the measured gap changes sign between twenty-four and twenty-six: +0.75 points at one, −0.03 at the other. At equal spreads they coincide at twenty, and so does the crossing.
The two rules answer different questions, and the difference between them is small in units — three of forty at a spread ratio of two. What makes it worth knowing is that one of them protects the test’s error rate and the other its precision, and a design chosen by the second on skewed outcomes has a two-sided test that is quietly one-sided.
What the table does not show
Everything above is drawn with the populations known, and a real comparison knows neither nor . Two things follow, and one of them is reassuring.
The reassuring one is that the direction is predictable without the numbers. On a quantity that is right-skewed in both groups — a duration, a cost, a concentration — has the sign of the term with the larger , which is the smaller or the more variable group. A design that puts fewer units in the more variable arm, which is common when that arm is the expensive one, produces rejections that say that arm is lower more often than they should.
The less reassuring one is that the sample skewness, which would be needed to compute from data, is biased low and very noisy at the group sizes where it matters, so a data-driven check tends to report less skew than there is. The design-time version of the rule, with the skewness taken from what is known about the outcome, is more trustworthy than the analysis-time one.
What the region establishes, and what it does not
The skewness of the difference ranks the cells by the gap between Welch’s tails, with a rank correlation of 0.997 across all twenty-five. Stated over the whole region rather than per pairing, because the claim is that the mechanism behind a cell does not matter once is known, and only a comparison across mechanisms can say that.
Where two identical skewed populations meet at equal sizes, is exactly zero and the tails balance, to within half a point. This is the case that makes the formula’s cancellation visible rather than assumed.
The measured gap changes sign at the split where does, to within one step of two units, at spread ratios of one, one and a half and two.
What does not survive is Welch’s test used on skewed groups of unequal size as though its level were two-sided. At eight and thirty-two, with a wide exponential against a normal, the low side alone rejects 9.79% of the time against the 2.5% it is allotted.
Every rate in the tables is a count over twenty thousand pairs of samples, so each carries its own sampling error: about 0.11 points for a rate near 2.5% and about 0.2 points for one near 10%. Differences between cells smaller than a few tenths of a point — the 2.04 against 2.22 at the balanced split, for instance — are inside that error and are not read as findings. The sign change at the cancelling split is read from differences of the same size, which is why it is claimed to within one step of the grid rather than at a point, and why the crossing is checked at three spread ratios rather than one. The same seed is used in every cell, so neighbouring cells share their luck; that makes the curve smoother than independent counts would and does not bias any single cell.
Not claimed: that predicts the total size as tightly as it predicts the imbalance — it does not, and the rank correlation of 0.956 is reported beside the 0.997 for that reason. Nor anything about populations with heavy tails and no skew, which the one-sample essay found do not unbalance an interval, and which have not been measured here in the two-sample case.
Still open: a sample’s own estimate of the skewness of its difference
The formula is exact for known populations. The practical version replaces , , and with sample estimates and reports beside the test. Whether that estimate is good enough to flag the dangerous cells — whether, at eight and thirty-two, is usually large when is 0.47 — is a measurement this region can make and has not. The sample skewness’s downward bias suggests it will under-flag, and the question is by how much.
The other open question is the repair. A skew-corrected reference distribution for the one-sample mean exists and is measured on the side a bound is read from. Its two-sample analogue corrects with in place of the one-sample skewness, and whether it balances Welch’s tails across this region — including the mirrored pairing, which no allocation can fix — has not been counted.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A bound written for a coin — both name central limit theorem, skewness
- A count that has to be estimated — both name allocation, central limit theorem
- A flat point with more than one direction — both name central limit theorem, skewness
- A width promised for a difference — both name allocation, neyman allocation
- Balancing towards unequal targets — both name allocation, neyman allocation
- Guessing one arm in three — both name allocation, error rate
Named objects
A flat tag is an object no other essay names yet.
AllocationBehrens–FisherCentral limit theoremError rateNeyman allocationSkewnessTwo-sample testWelch test