Series

Pooling — the series

6 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. Eight groups, τ = 1 against a within-group spread of 3. Each row is a group. The hollow circle is the group's own mean, the filled one is the estimate after pooling, and the small mark is the truth the data was generated from. The group of 3 moves 75% of the way to the population mean of 0.10; the group of 40 moves 18%.

    Eight groups, one population

    Eight hospitals are neither one hospital nor eight unrelated problems. The two obvious answers cost 2.23 and 1.15 in squared error; the estimate between them costs 0.88, and the weight it uses is not a matter of taste.

    part 1 · hierarchical
  2. One group 6 population widths from the rest. Squared error for each group under partial pooling, with each group's own mean beside it. Seven of the eight are estimated better by pooling. The eighth, which was never from the population, is estimated 6.1 times worse — 13.9 against 2.3.

    When borrowing goes wrong

    Partial pooling wins on the total and can lose badly on one group. Placed six population widths out, the group that was never from the population is estimated six times worse than by its own mean — and nothing in the output says so.

    part 3 · hierarchical
  3. How often "there is no spread between the groups" is reported about data that has one. Every dataset here was generated with a real population spread of 1. The moment estimator is the difference between the observed spread and what noise alone would produce, clamped at zero, and the difference comes out negative often: at eight groups it reports exactly zero on 32.6% of datasets, which is an instruction to pool completely and give all eight groups the same estimate. The rate falls to 4.2% at 48 groups.

    When the spread estimates to zero

    The usual estimate of a population spread is a difference of two positive quantities, clamped at zero. On a third of eight-group datasets with a real spread in them the difference comes out negative, the estimate is exactly zero, and every group is pooled completely on data that said no such thing.

    part 4 · fullbayes
  4. What each estimator costs, against how many groups there are. Each point is 4,000 datasets, with the population spread estimated from the data rather than supplied. Partial pooling first beats BOTH of the estimators it sits between at 5 groups; below that, complete pooling — which estimates nothing at all — is the better answer. Its own cost falls from 1.751 at 2 groups to 0.795 at 40.

    The fewest groups that can borrow

    At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.

    part 5 · hierarchical
  5. What each group gains from being pooled, τ = 1. Eight groups whose sizes span a factor of 13.3. The smallest gains 2.076 of squared error, which is 36.8% of the total reduction; the largest gains 0.030, which is 0.5%. 2 of the eight account for half of everything pooling buys.

    Where the borrowing goes

    Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.

    part 6 · hierarchical
  6. Twelve groups from two clusters, τ = 1. Every group's truth is in one of two clusters, and the population's total spread is exactly τ = 1, so the analysis recovers τ̂ = 1.515 and every shrinkage weight is what it would be for a single normal population. The estimates are pulled towards the grand mean, which is the middle of the gap — a place 0 of the 12 truths are and 6 of the estimates end up.

    One population, or two

    Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.

    part 7 · hierarchical

All series