The spread, and its own uncertainty

A prior on the spread

Integrating over the population spread means putting a prior on it, which sounds like the objection rather than the repair. The prior's effect is measurable, it is invisible where the groups are clearly different, and the reflex choice for a scale parameter turns out not to have a posterior at all.

Worth reading first: What a prior is worth · The weight that decides.

Integrating over the population spread instead of estimating it fixes a real defect: it takes the coverage of a pooled interval from 79% to 95%. It also introduces something the estimate did not have. An integral needs a measure, so τ needs a prior, and a prior is exactly what the whole apparatus of partial pooling was supposed to avoid — the point of estimating τ from the data was that nobody had to supply it.

The objection is fair and it has an answer, but the answer is a measurement rather than an argument.

Three priors on the spread, at a true τ of 0.5The posterior for τ under a flat prior (mean 1.66), a half-Cauchy of scale 1 (1.32) and one of scale 0.25 (1.15). The three answers differ by 30% of the widest. The prior does visible work when eight groups cannot separate a small spread from none, and almost none when they can.00.2500.5000.750101234τ, the spread between the groupsposterior densityflat on τhalf-Cauchy, scale 1half-Cauchy, scale 0.25the same eight groups, three priors on τposterior means 1.66 · 1.32 · 1.15
Fig. 1 The same eight groups under three different priors on the population spread: flat on τ, a half-Cauchy with scale 1, and a half-Cauchy with scale 0.25. The three posteriors are visibly different objects, and their medians — 1.50, 1.23 and 1.08 — span 28% of the widest.

What a prior on a spread is being asked to say

A prior on μ, the population mean, is the familiar case and what a prior is worth measured it in observations. A prior on τ is a different kind of statement. It is not a belief about where the groups are; it is a belief about how different from each other groups of this kind tend to be, before any of them has been looked at.

That is a question a reader in a subject usually can answer. Eight hospitals’ mortality rates do not differ by a factor of ten. Eight schools’ coaching effects are not fifty points apart. The half-Cauchy is the usual form because it does the two things that statement needs: it is finite at zero, so it does not rule out the groups being identical, and it has a heavy tail, so a scale chosen too small does not stop the data saying otherwise.

The flat prior says less than that and is the default here for exactly that reason. It is proper in this model — the likelihood dies fast enough at large τ that a flat density still integrates — and it is the choice that makes the fewest claims while remaining a choice.

Where the prior does visible work

The three curves above are far apart. Move the slider and they close.

Three priors on the spread, at a true τ of 4. The posterior for τ under a flat prior (mean 5.62), a half-Cauchy of scale 1 (4.67) and one of scale 0.25 (4.64). The three answers differ by 17% of the widest. The prior does visible work when eight groups cannot separate a small spread from none, and almost none when they can.
Fig. 2 The identical calculation on data whose groups are genuinely far apart. The three curves have moved much closer together, and the two half-Cauchys are now nearly indistinguishable from each other — the tighter one’s scale of 0.25 has been overruled completely. The three medians are 5.17, 4.41 and 4.38.
Three priors on the spread, at a true τ of 0.1. The posterior for τ under a flat prior (mean 1.23), a half-Cauchy of scale 1 (0.92) and one of scale 0.25 (0.60). The three answers differ by 51% of the widest. The prior does visible work when eight groups cannot separate a small spread from none, and almost none when they can.
Fig. 3 And the other end: eight groups that are very nearly identical. Here the three medians are 1.11, 0.85 and 0.51, spanning 54% of the widest. Every one of those is an answer to a question the data did not answer, and they differ by a factor of two.

That is the shape of the answer to the objection, and the honest version of it is more interesting than the reassuring one.

The prior is not harmless. Where eight groups cannot distinguish a small spread from none, the three answers differ by a factor of two, and every group’s estimate and every group’s interval moves with them. Anyone who reports one of them without saying which prior produced it is reporting an assumption dressed as a measurement.

And it does not stop mattering. This is where the reassuring version of the story would say the data takes over and the prior washes out, and on eight groups it does not quite. Running the comparison across the whole slider gives a span of 54%, 42%, 28%, 21%, 17% and 15% as the groups separate. The trend is unmistakable and the limit is not zero.

The residual 15% is worth understanding, because it is not the prior refusing to be overruled about small τ — the half-Cauchy of scale 0.25 has given that up entirely by then. It is the upper tail. A flat prior puts far more mass at large τ than a half-Cauchy does, the posterior’s upper tail is the part eight groups constrain worst, and a summary that integrates over that tail inherits the difference. The median differs by 15%; the mean, which weights the tail more, by 17%.

So the correct statement is narrower than “the data takes over”. It is: the prior stops deciding whether the groups differ, and does not stop influencing by how much. With eight groups there is never enough information to render the choice irrelevant, and the honest report says which prior was used at every separation, not only at small ones.

This is a weaker version of what the prior washing out shows one level down, and the connection between the two levels is exact rather than decorative: the population distribution over group means is a prior for each group, and now the population’s own spread has a prior over it. The hierarchy is priors all the way up, and at each level the question is whether the level below carries enough data to overrule the one above.

Four priors meeting the same data, true proportion 0.4. A prior is worth exactly a + b observations: Jeffreys — Beta(½, ½) is worth 1, weakly informative — Beta(2, 2) is worth 4, confident, centred at 0.5 — Beta(20, 20) is worth 40, confident and wrong — Beta(30, 5) is worth 35. Each curve is the posterior mean as the data accumulates, and every one of them converges on 0.4.
Fig. 4 The same phenomenon one level down, for comparison. A prior on a proportion is worth a stated number of observations and is overwhelmed once there are more than that. What is different about a prior on τ is how few observations there are to overwhelm it with: a study with four hundred patients in eight hospitals has four hundred observations about the hospitals and eight about how much hospitals differ.

That last sentence is the whole reason this level behaves differently. The sample size for τ is the number of groups, not the number of observations, and the number of groups is usually small and usually not something the analyst chose.

The reflex choice, and what is wrong with it

There is a standard prior for a scale parameter — p(σ) ∝ 1/σ, flat on the logarithm — and it is the right answer often enough to be a reflex. It is scale-invariant, which is the property a genuine scale parameter should have: it says nothing about the units.

Applied to τ in this model it does not work, and it does not work in a way that produces no error message, no warning, and a perfectly plausible-looking answer.

The same posterior, taken on a 201-point grid and on a 1601-point one. Four curves and only two calculations. Under a flat prior the two grids give the same answer — posterior mean 1.225 against 1.226, a difference of 0.1%. Under a 1/τ prior they do not: 1.050 against 0.812, a fall of 22.7%, with the density at the smallest τ on the grid rising from 0.92 to 1.80. There is nothing to converge to: the likelihood is finite at τ = 0 and 1/τ is not integrable there.
Fig. 5 Four curves and two calculations, on eight groups that look nearly identical. Under a flat prior the two grids give the same posterior mean to four digits — 1.225 against 1.226. Under a 1/τ prior they do not: refining the grid eightfold takes the mean from 1.050 to 0.812, a fall of 23%.
The same posterior, taken on a 201-point grid and on a 6401-point one. Four curves and only two calculations. Under a flat prior the two grids give the same answer — posterior mean 1.225 against 1.226, a difference of 0.1%. Under a 1/τ prior they do not: 1.050 against 0.720, a fall of 31.4%, with the density at the smallest τ on the grid rising from 0.92 to 6.22. There is nothing to converge to: the likelihood is finite at τ = 0 and 1/τ is not integrable there.
Fig. 6 And the same comparison against a grid thirty-two times finer than the coarse one. The flat prior’s curves are still indistinguishable at 1.226. The 1/τ prior’s mean has moved again, in the same direction, to 0.720 — the second refinement moving it about as far as the first. That is the signature of a logarithmic divergence, and it is what “there is no limit” looks like when it is drawn rather than asserted.

The reason is short. The marginal likelihood at τ = 0 is finite and non-zero — the groups being identical is a possibility this model contains, not a degenerate limit of it. Multiplying a finite function by 1/τ near the origin gives something that behaves like 1/τ, and ∫1/τ dτ diverges. The posterior does not exist.

What makes it dangerous rather than merely wrong is that nothing in a single calculation reveals it. A grid integrates over a finite range starting at its smallest non-zero point; the arithmetic completes; a density is plotted, a mean is computed, an interval is quoted. Every one of those numbers is a property of the grid. Refine the grid and they all move, and they keep moving, logarithmically and forever.

The check that catches it is therefore not a check on any one answer. It is the comparison of two answers at two resolutions, and it needs the proper case beside it or it would be measuring quadrature error. The site’s own gate runs it on a harder dataset than the figures above — eight groups whose standard errors are large relative to their spread, the case where the likelihood at τ = 0 is largest — and there the flat prior’s mean moves 1.0% across an eightfold refinement while the 1/τ prior’s moves 51%, from 5.92 to 2.90. A factor of fifty-two between the two, on the same two grids.

Why the pathology hides

There is a second half to this and it is the part worth carrying to other problems.

The 1/τ prior is improper here on every dataset. But the visible misbehaviour depends entirely on how much likelihood there is near zero. On data whose groups are clearly different, the likelihood at small τ is so tiny that it suppresses the singularity almost completely, and the two grid resolutions agree to three digits. The posterior still does not exist; it simply takes an unreasonable amount of refinement to notice.

What eight groups say about τ, when the truth is 4. The posterior density for the population spread after eight groups whose standard errors run from 0.5 to 2.1. The shaded band is the central 95% interval, from 3.02 to 10.85; the posterior median is 5.17 and the mean 5.62. The vertical mark at 4.25 is the moment estimate that empirical Bayes substitutes and then treats as known.
Fig. 7 Data on which the pathology would be invisible. The posterior for τ has essentially no mass below 1, so a prior that misbehaves only near zero has nothing to misbehave on. A method validated here would be certified for use on the data where it breaks.

So the failure appears exactly where the prior matters, which is exactly where it would be relied on. That is the same structure as the plug-in interval in the previous essay — good where it is not needed, worst where it is — and it is worth stating as a general habit rather than as two coincidences. A method should be checked at the setting where it is load-bearing, and the setting where it is load-bearing is usually the one where the data is thinnest.

The disagreement, measured against the data’s own resolution

Fifteen per cent sounds either large or small depending on what it is set beside, and there is a natural thing to set it beside: how well eight groups measure τ at all.

The sample size for τ is the number of groups, and a variance estimated from J groups carries a relative standard error of about √(2/(J − 1)). At J = 8 that is 0.53 on τ², which is about 27% on τ itself. So the data’s own resolution on the population spread — the width of the range the same prior would report across repeated studies of the same population — is roughly twice the 15% the three priors disagree by at the top of the range, and about half the 54% they disagree by at the bottom.

That comparison is the honest scale for the objection this essay opens with, and it cuts both ways. At clearly-separated groups the prior contributes about half as much uncertainty as the sampling does, which is not negligible and is not the dominant term. At nearly-identical groups it contributes about twice as much, and there the choice of prior is the largest single thing deciding the answer.

The asymmetry that matters is that these two uncertainties are of different kinds. Sampling error is random and prior disagreement is systematic. Run the study again and the 27% resamples; run it again and the choice of half-Cauchy scale does not. Ten studies of eight hospitals each, all analysed under the same tight prior, will agree with each other closely and be wrong together — which is exactly the failure a meta-analysis is least equipped to notice, since it reads the agreement as evidence.

The same arithmetic says how much more data would settle it. The resolution improves as 1/√(J − 1), so halving the 27% takes four times the groups: twenty-nine hospitals rather than eight, to reach about 13%. The number of groups at which the prior stops mattering is a number of groups nobody has, and that is a property of the design rather than of the analysis.

Where the fifteen per cent stops falling

The span across the slider — 54%, 42%, 28%, 21%, 17%, 15% — is worth reading as a sequence rather than as a first value and a last one, because its shape says which of two things is happening.

The successive falls are 12, 14, 7, 4 and 2 points. Four-fifths of the total decline happens in the first half of the range, and the last three steps together move it by six points. A quantity heading for zero does not decelerate like that; a quantity heading for a floor does. The curve is not slow convergence to irrelevance, it is fast convergence to about 15% and then nothing.

The floor has a source, and it is named above: the upper tail. A flat prior and a half-Cauchy differ most in how much mass they place at large τ, the likelihood from eight groups constrains that region worst, and no amount of separation between the groups improves it — separating the groups moves the posterior into the badly-constrained region rather than out of it. That is why the median’s residual disagreement of 15% becomes 17% for the mean, which weights the tail more heavily: the two summaries are ranked by how much of the unconstrained part they look at.

The practical reading is a small one and it survives the whole essay. Report a posterior median rather than a mean for τ, and report an interval for it rather than either. The median is the summary least exposed to the part of the posterior eight groups cannot see, and an interval at least displays the width of what the prior is deciding instead of collapsing it into a point that looks measured.

The method that avoids the prior is using one

There is a rejoinder to all of this which sounds decisive and is backwards.

Empirical Bayes needs no prior on τ at all. It estimates τ and uses it. Whatever the difficulties above, they are the price of a Bayesian treatment, and the plug-in does not pay them.

Substituting τ̂ and proceeding is not the absence of a prior on τ. It is a prior on τ, and specifically it is a point mass at τ̂ — a distribution asserting that the population spread is exactly 1.07 and that no other value has any probability whatsoever. Set beside a flat prior, or a half-Cauchy, or anything else anyone might reasonably write down, that is not the least assuming choice available. It is by an enormous margin the most.

τ̂ across 2,000 datasets of 8 groups, true τ = 1.5. The population spread is not supplied to a hierarchical model — it is estimated from how far apart the group means are, after subtracting the noise that would separate them anyway. It averages 1.35 here against a true 1.5, and comes out exactly zero on 11% of datasets.
Fig. 8 Where the point mass would be put, over two thousand datasets. The moment estimate of τ has this whole distribution — piled at zero, spread across a factor of several, with a long tail. Empirical Bayes draws one value from it and then treats that value as certain. The prior it is implicitly using is a spike, and which spike depends on the seed.

That reframing is what makes the coverage number in the previous essay unsurprising rather than mysterious. An interval built on a prior that claims perfect knowledge of a quantity nobody knows is going to be too short, and 79% against 95% is the size of the effect. The choice is not between using a prior on τ and not using one. It is between a prior that says what it does not know and a prior that says it knows everything.

What to use instead

Three choices are defensible and this site’s figures use the first.

Flat on τ. Proper in this model for four groups or more, makes no claim about scale, and is the usual recommendation for the case where nothing is known. Its one cost is that it puts a great deal of prior mass at large τ, which pulls the posterior mean up and makes the interval for τ itself wide. Every number in this field’s figures comes from it.

Half-Cauchy with a stated scale. What to use when something is known. The scale is a statement about how different groups of this kind get, and it is checkable against the answer: if the posterior for τ is concentrated well inside the prior’s scale, the data settled it; if it is not, the prior is doing work and the analysis should say so.

A prior on log τ that is proper. The scale-invariance the 1/τ prior was reached for can be had without the impropriety, at the price of stating a range. This is honest and rarely used, because stating the range feels like more of an assumption than 1/τ does — which it is not. It is the same assumption, written down.

What is not defensible is 1/τ, and the argument against it is not aesthetic. It is that the number reported depends on how finely somebody integrated.

One practical consequence follows from the 15% that never went away. A prior sensitivity check is not optional at this level and it is cheap. Running the same analysis under two or three priors costs seconds — every figure in this essay is that comparison — and the output is one line: the estimates moved by this much, or they did not. Where they did not, the line is a small piece of evidence that the conclusion is about the data. Where they did, the line is the finding.

The line to watch is not the group estimates, which are stabilised by their own data and move less than τ does. It is the intervals, and particularly the intervals for the groups with the largest standard errors, which are the ones being told what to be by the population. Those are the quantities a prior on the population’s spread has the most purchase on, and they are the ones a reader is most likely to quote.

The measurement this leaves

Three proper priors, on data where eight groups cannot settle τ, gave posterior means spanning about a third of the largest. On data where they can, the same three agreed to two digits. An improper prior on the same data gave an answer that moved by half when the grid was refined by eight.

Those three sentences are the entire content of the prior question in this model, and none of them is a philosophical position. Each is a number that came out of running the calculation.

The reason the question can be settled that way rather than argued about is the one this whole site runs on. A prior is not a belief that has to be defended in the abstract; it is a component of a procedure, and a procedure can be run ten thousand times against a truth that is known because it was generated. The half-Cauchy of scale 0.25 is not wrong because it is presumptuous. It is wrong on data generated at τ = 4, where it reports 4.38 against a flat prior’s 5.17, and right on data generated at τ = 0.25, where its 0.74 is closer than the flat prior’s 1.26. Both of those are countable, and what is countable does not need to be believed — which is the same test the interval that integrates puts the interval itself through.

What has not been addressed is the other end of the range. The posterior for τ can be piled against zero rather than spread across a factor of four, and there is a particular thing that happens when the estimate of τ is not merely small but exactly zero — which the moment estimator reports on nearly a third of eight-group datasets that have a real spread in them. That is not a prior problem. It is what happens when a difference of two positive quantities is negative, and it is the next essay.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Empirical BayesFlat priorThe half-Cauchy priorHierarchical modelImproper posteriorMarginal likelihoodPosteriorPosterior meanPriorPrior sensitivity