Hierarchy past one number

A prior on the spread between two sites

Integrate the between-site spread of a two-site study under a half-normal prior of scale 1 and the interval for the overall mean is 3.01 wide and covers 94.63% when the true spread is 1 — the supplied interval again, under another name. At a true spread of 3 it covers 68.38%. Every prior tried covers 95% averaged over its own draws and none covers it at every spread; each holds up to about its own scale, because two sites move the posterior median of the spread by less than a third of what the truth moves.

Worth reading first: Eight groups, one population.

A study run at two sites has one degree of freedom for the spread between sites. A level with two units showed what that means for the variance itself: it comes out exactly zero on a quarter of studies, and its middle half spans a factor of thirteen. What a two-unit study should report carried that through to the number a study prints, an interval for the overall mean. Three intervals answer two different questions. The fixed-effect interval covers the mean of the two sites in hand and only about half the time the mean of sites in general. The exact random-effects interval, a tt on one degree of freedom, covers the population mean at exactly its level at the cost of being more than ten times as wide. A spread supplied from outside the study covers at its level if the supplied value is right and not otherwise.

The fourth response was named and not measured: put a prior on the between-site spread and integrate it out. That is what a Bayesian analysis of a two-site study actually does, and it has a reputation for being the sensible middle. The prior supplies a distribution rather than a value, so the interval should sit somewhere between the supplied one and the exact one. The worry was that it would be narrow because the prior is confident, and would read as though the data had spoken. This essay measures where it sits, under five priors, against both the exact interval and the supplied one, at every true spread from zero to 3, against a within-site standard deviation of 1.2.

The setting and the arithmetic

Each study has two sites, ten observations a site, and a within-site standard deviation of 1.2. So each site’s mean carries a known sampling variance of s2=1.22/10=0.144s^2 = 1.2^2/10 = 0.144 around its site’s true mean, and the site means scatter around the population mean with the between-site standard deviation τ\tau. The quantity reported is an interval for the population mean, and the true spread τ\tau is varied from 0 to 3, four thousand studies at each.

With a flat prior on the overall mean the posterior has a closed shape. Given τ\tau, the overall mean is normal around the average of the two site means with variance (τ2+s2)/2(\tau^2 + s^2)/2. The posterior for τ\tau is the prior times (τ2+s2)−1/2exp⁡{−S/2(τ2+s2)}(\tau^2+s^2)^{-1/2}\exp\{-S/2(\tau^2+s^2)\}, where SS is the sum of squared deviations of the two site means from their average. The interval for the overall mean is therefore a mixture of normals all centred on the same point, and its half-width is the one value at which the mixture holds 95% of its mass. Everything here is computed on a grid of 320 values of τ\tau from 10−410^{-4} to 400, spaced evenly in its logarithm, so nothing is sampled except the studies themselves.

The likelihood factor is worth looking at before any prior is chosen. As τ\tau grows it falls only as 1/τ1/\tau. Two site means can say that the spread is not tiny, if they differ by a lot, but they cannot say that it is not huge. Whatever decides how wide the interval is, it is the prior’s upper tail, divided by τ\tau. A flat prior on τ\tau — the choice that sounds most like “letting the data speak” — gives a posterior that falls as 1/τ1/\tau, which does not integrate. With two sites, the flat prior is not a weak choice, it is no answer at all.

Coverage, at every true spread

Five proper priors are measured: half-normals of scale 0.5, 1 and 2, a half-Cauchy of scale 1, which is the common recommendation for a group-level standard deviation, and an inverse-gamma with both parameters 0.001 on the variance, which is the older default that recommendation was written against.

Coverage of the population mean from two units, by the true spread between units. 4,000 two-unit studies at each true spread, ten observations a unit, within-unit standard deviation 1.2. exact, on one degree of freedom: 94.97%, 95.50%, 94.95%, 94.83%, 95.15%, 95.20%, 94.67%. spread supplied as 1: 100.00%, 100.00%, 99.92%, 95.53%, 81.85%, 70.03%, 51.55%. half-normal, scale 1: 100.00%, 100.00%, 99.78%, 94.63%, 86.25%, 79.70%, 68.38%. half-Cauchy, scale 1: 100.00%, 100.00%, 100.00%, 98.98%, 95.50%, 92.38%, 88.00%. inverse-gamma(0.001, 0.001) on the variance: 100.00%, 100.00%, 100.00%, 99.15%, 97.05%, 95.28%, 92.65% — at spreads 0, 0.25, 0.5, 1, 1.5, 2, 3. The exact interval holds 95% at every spread; every integrated one over-covers when the spread is small and under-covers when it is large, and where it crosses 95% is set by the prior's tail.
Fig. 1 Coverage of the population mean by the exact interval on one degree of freedom, the interval with the spread supplied as 1, and the integrated intervals under three priors, at true between-site spreads from 0 to 3.

The exact interval holds 95% at every spread, as the arithmetic of a tt on one degree of freedom says it must: 94.97% with no spread at all, 94.83% at 1 and 94.67% at 3. Every integrated interval has a different shape. It over-covers when the true spread is small, falls through 95% somewhere, and under-covers beyond.

The half-normal of scale 1 covers 94.63% at a true spread of 1 and is 3.01 wide there. The interval with the spread supplied as 1 covers 95.53% and is 2.96 wide. Those are the same interval. Integrating under a prior whose scale is the truth gives back what supplying the truth gives, and it fails the same way when the truth moves: at a spread of 2 the half-normal covers 79.70% and the supplied interval 70.03%, and at 3 they cover 68.38% and 51.55%. The integration buys some protection — the half-normal’s posterior does move upward when the sites differ a lot — but much less than the exact interval’s, which does not move at all.

The other priors are the same picture with the crossing moved. The half-normal of scale 0.5 drops through 95% near a spread of 0.5 and covers 51.23% at 3. The half-normal of scale 2 holds until about 1.9 and covers 85.80% at 3. The half-Cauchy of scale 1 holds until about 1.6 and covers 88.00% at 3. The inverse-gamma holds until about 2.1 and covers 92.65% at 3. Each prior’s interval is honest up to roughly its own scale, and the prior’s tail decides how quickly the honesty is lost past it. Nothing in two sites’ data moves that crossing.

What each interval costs

The width figure is where the inverse-gamma prior changes the story.

Median interval width from two units, by the true spread between units. Median width over 800 studies at each spread. exact, on one degree of freedom: 4.62, 5.48, 7.11, 13.47, 17.88, 23.37, 35.05. inverse-gamma(0.001, 0.001) on the variance: 3.43, 3.62, 4.12, 8.21, 13.14, 19.69, 31.27. half-Cauchy, scale 1: 3.28, 3.35, 3.54, 4.66, 5.66, 6.94, 9.57. half-normal, scale 1: 2.44, 2.48, 2.56, 3.01, 3.36, 3.75, 4.42. The supplied interval is 2.96 wide and the fixed-effect one 1.13 at every spread.
Fig. 2 Median interval width at each true spread for the exact interval and three integrated ones, with the width of the supplied interval drawn across.

The exact interval’s median width grows with the truth, from 4.62 with no spread to 13.47 at 1 and 35.05 at 3. That is the tt on one degree of freedom multiplying half the observed gap between the sites, and the gap grows with the spread. The half-normal of scale 1 stays narrow everywhere — 2.44 with no spread, 3.01 at 1, 4.42 at 3. Its narrowness is the prior’s confidence, exactly as the worry predicted. The half-Cauchy is wider, 4.66 at 1 and 9.57 at 3, a third and a quarter of the exact interval’s width there.

The inverse-gamma is the surprise. Its reputation is for over-confidence. It piles prior mass near zero, and in studies with few groups it is supposed to pull the spread’s estimate down and the interval in. Here its posterior median for the spread is indeed the smallest of any prior when the truth is small, 0.20 with no spread at all. But its upper tail on τ\tau falls only as τ−1.002\tau^{-1.002}, so after the likelihood’s 1/τ1/\tau the posterior still has a heavy tail, and the interval follows the tail. At a spread of 1 it is 8.21 wide, more than half the exact interval’s width. At 3 it is 31.27, nearly all of it. With two sites the inverse-gamma’s problem is not that it is confident, but that it is nearly as wide as the exact interval while giving up the exact interval’s guarantee.

So the integrated intervals sort themselves on a single line, and the line is the prior’s tail. A light tail gives a narrow interval that is honest only where the prior’s scale is right. A heavy tail gives a wide interval that is honest further out. No prior escapes the trade, because nothing else in the problem can pay for width.

How far two sites move the posterior

The reason is visible directly in what the posterior says about the spread itself.

What two units move the spread's posterior to, against what the spread actually is. For each prior, the median over 800 studies of the posterior median of the between-unit spread, at each true spread. half-normal, scale 0.5: 0.27, 0.27, 0.30, 0.42, 0.53, 0.67, 0.89. half-normal, scale 1: 0.44, 0.44, 0.48, 0.70, 0.85, 1.03, 1.31. half-Cauchy, scale 1: 0.44, 0.46, 0.51, 0.76, 0.98, 1.28, 1.82. inverse-gamma(0.001, 0.001) on the variance: 0.20, 0.21, 0.25, 0.57, 1.03, 1.66, 2.73. The diagonal is a posterior that follows the truth. From 0 to 3 the half-normal of scale 1 moves by 0.87, a third of the truth's movement.
Fig. 3 The median, over studies, of each prior’s posterior median for the between-site spread, at each true spread, with the diagonal a posterior that followed the truth.

Under the half-normal of scale 1, the posterior median of the spread is 0.44 in a typical study when the true spread is zero and 1.31 when it is 3. The truth moved by three, and the posterior by 0.87, about 29% as far. The prior’s own median is 0.67, so the posterior sits close to the prior and is nudged by the data in roughly the right direction. The half-normal of scale 0.5 moves from 0.27 to 0.89, and the half-Cauchy from 0.44 to 1.82. Only the inverse-gamma follows the truth closely, from 0.20 to 2.73, and it does so because its posterior is so wide that its median is pulled about by the data’s single gap.

This is the same thing the full-Bayesian treatment of a hierarchy found with many groups, pushed to its limit. With many groups, the prior’s effect on the spread is visible where the groups are similar and vanishes where they clearly differ. With two groups, they never clearly differ in the sense that matters, because one gap is consistent with a wide range of spreads. The posterior median is a weighted compromise in which the prior’s weight never falls away.

The guarantee every prior keeps

None of this means the integrated intervals are wrong. They carry a guarantee, and it is a different one from the exact interval’s.

Each prior's interval, scored on truths drawn from that prior. 4,000 two-unit studies for each prior, each with its true spread drawn from the prior itself. Coverage of the population mean: half-normal, scale 0.5 95.23%, half-normal, scale 1 94.45%, half-normal, scale 2 95.23%, half-Cauchy, scale 1 94.42% — each at 95% within its counting error, which is the guarantee a posterior interval carries. At a fixed true spread of up to 3 the same intervals fall to 51.23%, 68.38%, 85.80%, 88.00%.
Fig. 4 Coverage of the population mean by each prior’s interval when the true spread of each study is drawn from that prior itself, beside the worst coverage the same interval reaches at a fixed true spread up to 3.

Draw each study’s true spread from the prior and score the interval on those studies, and every prior covers at its level within the counting error: 95.23% for the half-normal of scale 0.5, 94.45% for scale 1, 95.23% for scale 2 and 94.42% for the half-Cauchy. That is exactly what a posterior interval promises: coverage averaged over the prior. The same intervals at a fixed spread fall as low as 51.23%, 68.38%, 85.80% and 88.00%. The average guarantee is assembled from over-coverage at small spreads, where the prior puts most of its draws, and under-coverage at large ones, where it puts few.

The exact interval promises something stronger and pays for it. It covers at its level at every fixed spread, which is the promise a reader of a confidence interval assumes. A prior wrong about its own spread found the same split for a single mean, where a normal prior’s interval is calibrated over its own draws and miscalibrated at a fixed truth. Here the split is sharper, because the data are too thin to rescue a prior placed in the wrong spot.

Is there a prior whose interval is as short as the supplied one and as honest as the exact one? Among the five measured, no, and the reason is not a defect of any one prior. An interval that holds 95% at every true spread has to be wide whenever the two sites differ by a lot, because a large gap is exactly what a large spread produces and the interval cannot tell a large spread seen once from a small one seen unluckily. Every prior here is narrower than the exact interval precisely when the gap is large, and so each buys its shortness by promising less: 95% on average over a stated belief about the spread, rather than 95% whatever the spread is.

Where the prior stops being the answer

Two sites are the extreme case, and the arithmetic says how fast it eases. With KK sites the likelihood for the spread falls as τ−(K−1)\tau^{-(K-1)} rather than 1/τ1/\tau, so each added site takes some of the decision away from the prior’s tail. The same four thousand studies a setting, run at four sites and at ten, show how quickly.

At four sites the exact interval is a tt on three degrees of freedom, and its median width at a true spread of 1 drops from 13.47 to 3.01. The half-normal of scale 1 is 2.25 wide there, so the exact interval is a third wider rather than four and a half times as wide. At a true spread of 3 it covers 78.50%, up from 68.38% at two sites. At ten sites the exact interval is 1.48 wide at a spread of 1 and the half-normal 1.41, a saving of 5%, and the half-normal still covers only 87.22% at a spread of 3. The half-Cauchy covers 93.90% there, close to its level.

So the integrated interval’s two properties move at different speeds. Its saving in width, which is the reason to use it, disappears fast: by ten sites the exact interval costs almost nothing extra. Its loss of coverage when the prior’s scale is wrong disappears slowly, because a light-tailed prior keeps pulling the spread down until the data overwhelm it. With many groups, integrating over the spread is what makes a group’s interval cover, because the plug-in it replaces ignores the spread’s uncertainty altogether. With two, the comparison is not with a plug-in but with an exact interval, and there integration is a way of trading a guarantee for width. The trade is only worth considering where the width is large — which is where the guarantee matters most.

This is also the shape of borrowing in general on this field. How much a group borrows is decided by how much the data can say about the spread it borrows against. With two sites the data say almost nothing, so the borrowing is the prior’s. The between-site variance decides how many independent observations a grouped study is worth, as a study with crossed groupings found separately for each of its contrasts, and two sites is the case where that variance is known least.

What a two-site study with a prior should report

The prior on the spread, with its scale and its tail. The integrated interval is the supplied interval under another name whenever the prior’s scale is close to the truth, and close to the exact interval only when the tail is heavy enough. A reader cannot know which without the prior.

Which guarantee the interval carries. Coverage averaged over the prior, which every proper prior here delivers, is a statement about a population of studies whose spreads follow the prior. Coverage at a fixed spread is the reader’s usual assumption, and only the exact interval delivers it.

The same interval under a prior twice as wide. For the half-normal, doubling the scale from 1 to 2 widens the interval at a spread of 1 from 3.01 to 4.97 and raises its coverage at a spread of 3 from 68.38% to 85.80%. A study whose conclusion survives that doubling has a conclusion that does not rest on the prior’s scale. One whose conclusion does not survive it should say that the spread was supplied, not estimated.

The fixed-effect interval beside it. It is 1.13 wide at every spread and covers the mean of the two sites in hand. If the integrated interval is not much wider than that, the prior has decided that the two sites are close to the population, and that is a claim about sites, not a result about them.

What is claimed and what is not

Computed exactly on a grid. Every posterior, interval half-width and coverage indicator, on a 320-point logarithmic grid for the spread with the overall mean integrated analytically. A study’s coverage needs no root-finding: the population mean is inside a symmetric interval at level cc exactly when the posterior mass within that distance of the centre is at most cc.

Counted. Four thousand studies at each of seven true spreads, with the median widths and posterior medians over the first eight hundred, and four thousand studies for each prior with spreads drawn from that prior.

Not claimed. The within-site variance is known here, which is generous: with ten observations a site its estimate has eighteen degrees of freedom and matters little beside the spread’s one. Priors are on the spread’s standard deviation, apart from the inverse-gamma on its variance. Weakly informative priors built from an external database of similar studies are the obvious next candidate, and they are what the supplied interval becomes when it is honest about being uncertain.

Still open: a prior built from other studies

Every prior here was chosen by convention. A two-site study rarely stands alone. The same treatment has usually been run at two sites elsewhere, or at twenty, and the between-site spreads those studies saw form an empirical distribution. A prior built from that distribution is a supplied spread with its own uncertainty attached. It is the two levels at once arithmetic moved up one level, with studies as the top level and sites inside them. Whether such a prior is right on average by construction — and so honest in the averaged sense at no cost — and how far its interval’s coverage falls when the new study’s spread comes from a different population than the database’s, is the measurement that would turn the choice of prior from a convention into a calculation, and it has not been made.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Between group varianceCoverageCredible intervalDegrees of freedomHierarchical modelMonte CarloPosterior distributionPrior distributionRandom-effectsStudent's t