Hierarchy past one number

Borrowing towards a line

A group shrunk towards the average of all groups is being compared with groups it has nothing in common with. Fit a group-level predictor and it is shrunk towards what the predictor says a group like it should be — which halves the spread left to borrow against and takes a quarter off the squared error.

Worth reading first: The slope that borrows.

Sixteen hospitals, each reporting an outcome, each measured with its own precision. Partial pooling moves every one of them towards the average of all sixteen.

That is the right operation only if the sixteen are exchangeable — if, before seeing any outcomes, there was no reason to expect one hospital to differ from another. Usually there is a reason. One hospital is a regional trauma centre and another is a district general; one has twice the case mix severity of another; one treats a population whose average age is ten years higher.

If any of that is known before the outcomes are looked at, shrinking every hospital towards the same number is throwing it away.

16 groups shrunk towards a fitted line, at γ = 1.2Hollow circles are the groups' own values, filled ones the estimates after pooling, and the diagonal is the line fitted through them with each group weighted by how well it is measured — slope 1.21, intercept 0.20. The horizontal rule is where the same 16 groups would have been shrunk to with no covariate. The spread left to borrow against is 0.64 with the covariate against 1.28 without, so every group is pulled further in than it would otherwise have been.-202-101the group-level covariate zthe group's valueno covariate: everything shrinks hereθⱼ ~ N(α + γzⱼ, τ²), fitted by marginal likelihoodτ̂ falls from 1.28 to 0.64
Fig. 1 Sixteen groups against a covariate known before their outcomes were measured. The diagonal is the line fitted through them, each group weighted by how well it is measured; the horizontal rule is the single number they would all have been shrunk towards without it. Every group now moves towards what the line predicts for a group like it, and the spread left to borrow against has fallen from 1.28 to 0.64.

One change, in one place

The model is the same two levels with a different centre:

θjNormal(α+γzj,  τ2)\theta_j \sim \text{Normal}(\alpha + \gamma z_j,\; \tau^2)

instead of θⱼ ~ Normal(μ, τ²). The covariate zⱼ is known for every group before its outcome is measured; α and γ are fitted; τ is what is left over.

Everything else is unchanged. The shrinkage weight is still B = se²/(se² + τ²), each group’s estimate is still B·(centre) + (1 − B)·(own value), and the centre is now group-specific instead of shared. No new arithmetic appears.

What changes is what τ means. Without the covariate it is the whole spread between the groups; with it, it is the spread that remains once the covariate has been accounted for. On this data the first is 1.28 and the second is 0.64, so the covariate has explained three-quarters of the variance between the groups — and every group now borrows against half the spread it did before.

Because B falls as τ falls, that means every group borrows more. The covariate has not made the groups more independent of each other; it has made the population’s prediction for each of them sharper, which is worth listening to more.

The fit needs weights, and the weights need τ

There is a circularity in fitting that line and it is worth naming rather than stepping over.

The line should be fitted by weighted least squares, with each group weighted by 1/(seⱼ² + τ²) — a badly measured group should have less say in where the line goes. But τ is the residual spread about the line, which cannot be computed until the line is fitted.

The way out taken here is to scan: for every candidate τ on a grid, compute the weights it implies, solve for α and γ in closed form under those weights, evaluate the marginal likelihood, and keep the best. That is one pass and it has a useful by-product — the whole profile, rather than the single value at its top.

That profile is worth looking at for the same reason the posterior for τ was worth looking at in the field beside this one: it is usually flat enough that the single value at its maximum is a poor summary of it. The scan is cheap, so having it is free, and it makes the question “how well is τ determined here” answerable rather than assumed.

What it is worth

Six hundred datasets, sixteen groups each, generated from a known truth with a real covariate effect:

estimator squared error per group fitted spread
each group’s own value 0.691
pooled towards the grand mean 0.510 1.34
pooled towards the fitted line 0.384 0.63

Plain pooling takes 26% off the error of the groups’ own values; the covariate takes another 25% off that. And the fitted spread halves, which is the mechanism: the groups are being compared with groups like themselves rather than with all groups.

The two columns are not the same claim in different units, and the gap between them is worth noticing. The spread falls by 53% and the error by 25%. That is because a group’s estimate depends on the spread only through B, and B is a ratio: halving τ moves a group whose own standard error is comparable to τ from borrowing a third to borrowing two thirds, but moves a group measured far more precisely than τ hardly at all. The covariate helps the badly measured groups and barely touches the well measured ones, which is the same rule the whole field runs on, applied to a change in the population rather than to a difference between groups.

Both columns matter and it is possible to have one without the other. A covariate that explains nothing would leave the spread where it was and the error where it was; a covariate fitted to noise would reduce the apparent spread while leaving the error alone or making it worse. The library’s assertion therefore requires both — the spread falls and every group is estimated better — because either on its own is satisfiable by a covariate that is not doing anything real.

16 groups shrunk towards a fitted line, at γ = 0. Hollow circles are the groups' own values, filled ones the estimates after pooling, and the diagonal is the line fitted through them with each group weighted by how well it is measured — slope 0.01, intercept 0.20. The horizontal rule is where the same 16 groups would have been shrunk to with no covariate. The spread left to borrow against is 0.64 with the covariate against 0.64 without, which at this setting is the same number: a covariate explaining nothing leaves everything to explain.
Fig. 2 A covariate that explains nothing, on the same sixteen groups. The fitted slope is 0.01, the line is flat, and the spread left to borrow against is 0.64 with the covariate and 0.64 without — the same number, because there was nothing for it to explain. Nothing has been gained and, importantly, nothing has been lost.
16 groups shrunk towards a fitted line, at γ = 2.4. Hollow circles are the groups' own values, filled ones the estimates after pooling, and the diagonal is the line fitted through them with each group weighted by how well it is measured — slope 2.41, intercept 0.20. The horizontal rule is where the same 16 groups would have been shrunk to with no covariate. The spread left to borrow against is 0.65 with the covariate against 2.33 without, so every group is pulled further in than it would otherwise have been.
Fig. 3 And a covariate that explains a great deal. The fitted slope is 2.41 against a true 2.4, the spread without the covariate has climbed to 2.33 because the covariate’s own variation is now most of what is between the groups, and the spread with it is unchanged at 0.65. Every group is pulled hard towards the line.

The comparison across those three is the argument. The residual spread stays at about 0.64 whatever the covariate does, because 0.64 is what it actually is; the apparent spread without the covariate climbs from 0.64 to 2.33 as the covariate’s effect grows, because a covariate not in the model shows up as between-group variation. The population spread is not a property of the groups. It is a property of the groups and the model, and adding a predictor changes it.

The same arithmetic, read as a warning

That last sentence has an uncomfortable corollary and this site has a whole field about it.

If a covariate not in the model turns up as between-group spread, then a hierarchical model without the relevant covariate reports groups as differing when what differs is the covariate. Sixteen hospitals with genuinely identical quality but very different case mixes will produce a large τ̂, will be shrunk gently towards a common mean, and will be reported as differing from each other.

Two groups, one treatment, and both readings of the same numbers. The treatment wins in group A (93.0% against 87.0%) and in group B (73.0% against 69.0%), and loses overall (76.1% against 84.3%). Nothing here is a trick; the allocation differs between the groups.
Fig. 4 The extreme form of the same thing. Two treatments, one better in every subgroup and worse overall, because the subgroups differ in something not accounted for. A hierarchical model shrinking towards a grand mean is doing the arithmetic in the top row; adding the group-level covariate is doing it in the bottom.

Simpson’s reversal is that failure at its most dramatic, and the connection is not a metaphor: both are what happens when groups that differ in something unmodelled are combined as though they do not. What the hierarchical version adds is a measurement of how much of the between-group variation the omitted variable was responsible for — 1.28 against 0.64 here, so three quarters of the variance.

The direction of the error is worth stating precisely, because it is not the intuitive one. Omitting a relevant group-level covariate does not make the groups look more alike. It makes them look more different, because their genuine differences in the covariate are attributed to them rather than to it.

That has a consequence for how these analyses read. A hospital league table built from a hierarchical model without case mix will show a wide spread of hospitals and shrink each one only gently, because the model has been told the hospitals differ a lot. The same table with case mix in it will show a narrow spread and shrink hard, because most of what looked like difference has been explained. The second table is less dramatic and it is the one that is about hospitals.

Reading the profile

The scan produces the whole marginal likelihood as a function of τ, and looking at it is the cheapest diagnostic in this field.

Two readings are useful and neither needs any further computation.

How much of the profile is above zero. If the curve is still rising at the smallest τ on the grid, the data has not distinguished the residual spread from none, and the maximum is at the boundary. Every group will land on the fitted line, every estimate will be the line’s prediction, and the analysis has silently become a plain weighted regression on the group means.

How flat the top is. A profile that is nearly level across a wide range of τ is saying that the residual spread could be twice or half what was reported, and since B depends on τ, so could every shrinkage weight. That is the case where the plug-in intervals of the field beside this one are worst, and it is exactly the case a group-level covariate creates by reducing the residual spread.

Neither reading is a substitute for integrating over τ. Both are available for free from a calculation that has already been done, and both turn “the model fitted” into “the model fitted, and here is how well its top level was determined”.

What decides how much a slope borrows, at 10 observations throughout. Every point on this curve is a group of 10 observations. What changes along the axis is only where those 10 x values are placed. A group whose x values span 0.2 has a slope standard error of 3.72 and moves 97% of the way to the population slope; one spanning 2.9 has a standard error of 0.25 and moves 15%. The half-way point is at a spread of 1.25, where the slope's own standard error equals τ.
Fig. 5 Why the two columns move by different amounts. The weight is a ratio, so a change in the population spread slides a group along this curve rather than scaling its estimate — and the curve is steep in the middle and flat at both ends. A group already borrowing 90%, or already borrowing 10%, barely notices that τ has halved.

Why a spread that halves buys only a quarter

The two columns move by very different amounts — the spread falls 53% and the squared error falls 25% — and the reason given above, that B is a ratio and its curve is flat at both ends, has an exact form worth writing down.

Under this model a group’s expected squared error after pooling is

se²τ² / (se² + τ²)

which is smaller than either of the two quantities in it and is dominated by whichever is smaller. That single expression contains the whole gap between the columns. Reducing τ from 1.34 to 0.63 — which is what the covariate did, and which is 78% of the between-group variance, since 0.63²/1.34² = 0.22 — multiplies that error by τ₂²/(se² + τ₂²) over τ₁²/(se² + τ₁²), and the value of that ratio depends entirely on where the group’s own precision sits relative to the two spreads.

At the two extremes it is unambiguous. A group measured far worse than either spread has an error of about τ² whichever spread applies, so it gets the full 78% cut: the population’s prediction is all it has, and a better prediction is worth everything. A group measured far better than either has an error of about se² either way, and gets nothing at all: it was never listening to the population, so sharpening what the population says changes nothing about it.

Between them the crossing point is clean. Setting se² = τ₁τ₂ — the geometric mean of the two spreads, 0.92 here — makes the error ratio exactly τ₂/τ₁, so such a group’s squared error falls by precisely the 53% the spread did. That is the group for which the headline number and the actual gain coincide, and it is one particular group rather than the typical one.

So the 25% is a statement about the sixteen groups’ precisions, not about the covariate. It says most of them are measured well relative to a residual spread of 0.63, which is why the average gain sits nearer the zero end of the range than the 78% one. The same covariate on sixteen worse-measured groups would have bought most of the 78%, and on sixteen better-measured ones almost nothing — with the fitted spread falling by 53% in all three cases, because the spread is a property of the population and the gain is a property of the study.

Which groups the covariate is actually for

That accounting says where a group-level predictor is worth the trouble of finding, and it is not where the effort usually goes.

The gain is concentrated entirely in the groups whose own standard errors are large compared with the residual spread — the small hospitals, the short runs, the schools with forty pupils rather than four hundred. Those are exactly the groups a plain hierarchical model already shrinks hardest, which is to say the ones whose reported estimate is mostly the population’s opinion of them. A covariate replaces that opinion with a better one; for everyone else it replaces an opinion nobody was listening to.

The consequence for a league table is direct and slightly uncomfortable. Adding case mix changes the published position of the small units and barely moves the large ones, so the visible effect of the adjustment is concentrated in the part of the table that is least certain and most disputed. That is not a defect of the adjustment. It is what an adjustment is: the units with the least evidence of their own are the units for which what is assumed about them matters most, which is the same observation when borrowing goes wrong makes about a group that does not belong to the population at all.

What it does not fix

Two limits, both real.

The covariate has to be known before the outcome. A predictor chosen after looking at which groups came out high is not a covariate; it is a description of the outcome, and fitting it will reduce the apparent residual spread to nearly nothing while explaining nothing at all. This is the garden of forking paths at the group level, and the only defence is the same one: state the covariate before the outcomes are looked at.

The covariate itself is measured, and measurement error in it biases γ towards zero. Everything here assumes z is known exactly, which is reasonable for a design variable — how many beds, which region, which batch — and unreasonable for a summary of the group’s own patients, which is what a case mix index is. A covariate observed with error explains less than it should, leaves more residual spread than it should, and is therefore reported as less useful than it is. The direction is predictable and its size is not, without knowing how well z was measured.

A group can still be genuinely outside the model. Shrinking towards a line rather than a point does not help a group that is not from the population at all — it changes what it is being dragged towards, not whether it should be dragged. When borrowing goes wrong measures the cost of that, and the measurement is unaffected by whether the centre is a point or a fitted value.

There is a third thing that is not a limit but is easy to expect wrongly. Adding a covariate does not make the pooled estimates closer to the groups’ own values. It makes them further, because τ falls and B rises: with the covariate every group here moves further from its own observation than it did without it. The estimates are better and they trust the individual group less, which is the correct response to having a better prediction for it.

Where the sixteen came from

One practical note that the figures made unavoidable.

At twelve groups this essay’s figures collapsed. The residual spread after the covariate was estimated as exactly zero at every setting of the slider, every group landed precisely on the fitted line, and the picture showed complete pooling rather than the thing it is about. That is the atom at zero from the field beside this one, and fitting a covariate makes it more likely rather than less: the residual spread is smaller than the total spread by construction, and a smaller quantity estimated by the same subtraction lands below zero more often.

Sixteen groups and a slightly larger true spread put the figures back into the regime the essay is about. That is a legitimate choice for an illustration and it is not a fix for anything: an applied analysis with twelve groups and a good covariate should expect the residual spread to collapse to zero a substantial fraction of the time, with the consequences the other field documents — every group reported as identical to the line, and every interval far too narrow.

The general shape is worth carrying. Every improvement to a hierarchical model that reduces the residual spread also makes the residual spread harder to estimate, because it is the difference of two quantities that are now closer together. A better model is a model whose top level is less well determined, and the repair for that is the integral rather than the estimate.

What the three essays so far have taken

The weight has now been applied to a mean, a slope, a proportion and a group with a predictor attached, and each time the thing that decides how much a group borrows has been something different: the sample size, the arrangement of the x values, the value observed, and what a covariate already accounted for.

That is the argument this field was for. B = se²/(se² + τ²) is not four formulas. It is one formula whose two inputs mean different things in different problems, and the reason it looks like a statement about sample size in every introduction is that a mean is the one case where sample size is the only input.

One thing has been constant across all four and it is the thing the last essay takes: the hierarchy has had exactly two levels. Groups, and a population they come from. Adding a third changes nothing about the arithmetic and produces a quantity that the time-series field has already computed from an entirely different picture. What comes out of it is a number that decides how many independent observations a clustered study is worth — and, unlike everything in this essay, it can be computed before the study is run.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

ConfoundingHierarchical modelLeast squaresMarginal likelihoodPartial poolingRandom-effectsShrinkage