A prior wrong about its own spread
Worth reading first: What a prior is worth · What the 95% refers to.
When the prior is confident and wrong measured a prior in the wrong place. Worth thirty-five observations and centred away from the truth, it produced intervals that covered nothing at all, intervals whose width looked honest, and a failure that only the prior-predictive distribution of the data could see. It ended by setting aside a different failure. A prior can be centred exactly where the truths are and still be wrong — about how confident it is, or about how far an unusual truth can sit from its centre. A prior centred on the truth contributes no displacement, so every piece of arithmetic in that essay reports no problem.
Shape is a property of a population, not of one truth. A prior that says a proportion is near 0.2 with the confidence of a hundred observations is a claim about where the truths of many studies lie, and whether the claim is over-confident depends on how far those truths actually spread. So every measurement here averages over a stated population of truths, and for a proportion that average is exact. The coverage of a credible interval over truths drawn from a population is a sum over the possible counts of the population’s own probability of each count times its own posterior probability that the interval built from that count contains the truth — Beta–binomial and Beta, with no simulation anywhere.
Centred on the truth and too sure of it
The population of truths is centred at 0.2 and is worth ten observations: a Beta distribution whose spread is about 0.12, so studies’ true proportions run from a few per cent to about a half. Each study runs twenty trials. The analyst’s prior is centred at 0.2, the right place, with a weight that ranges from two observations to four hundred.
At the population’s own weight the credible interval covers 95.0% of the population’s studies — exactly, because a Bayesian interval under the prior the truths are actually drawn from is calibrated on average over those truths. Below that weight it covers a little more or a little less: 94.7% at two observations and 96.0% at five. Above it, coverage falls steadily. A prior worth twenty observations covers 88.1%, thirty-five 75.9%, a hundred 48.2%, and four hundred 23.9%.
So a prior twice as confident as it should be already costs seven points, and a prior ten times as confident halves the coverage. None of it involves a conflict of location: the posterior mean is pulled towards 0.2 by exactly the right amount on average across the population. What is wrong is the posterior’s spread. The prior tells the posterior that a truth far from 0.2 is nearly impossible, so the interval for a study whose truth is 0.4 is built around a mean dragged towards 0.2 and is too narrow to reach back to 0.4. Every study whose truth sits away from the centre is under-covered, and in a population worth ten observations most of them do.
Where in the population the misses are
The average hides a sharper picture, because coverage at a single truth is all or nothing far from the centre. For a prior worth a hundred observations, the interval covers a truth of 0.2 100.0% of the time — every study whose truth sits at the centre is covered — and a truth of 0.3 only 39.2%. At 0.1 and at 0.4 it covers 0.0%: no count from twenty trials can pull a hundred observations’ worth of prior far enough. The prior worth ten, by contrast, covers 95.7% at 0.1, 97.8% at 0.2 and 87.4% at 0.4, with its shortfall spread thinly over the population’s upper tail.
So the over-confident prior does not cover half the studies a little short of their truths. It covers the studies near its centre perfectly and the rest not at all, and the 48.2% is the share of the population that happens to sit close enough to 0.2. That is the same zero the location failure produced in the earlier essay, arriving by a different route: there one prior was far from one truth, here one prior is far from most of a population’s truths at once.
The width cannot tell earned confidence from claimed
The earlier essay’s lead suggested that this failure, unlike the location one, might show in the width: an over-confident prior gives a short interval, and a short interval is visible. It is visible. What it cannot show is whether the shortness is earned.
The figure sets two cases beside each other at each prior weight. In one the prior is right: the truths really are drawn from a population worth that many observations. In the other the prior has the same weight and the truths come from the population worth ten. A reader of either study sees an interval, and can compute how much narrower it is than the interval a flat prior would give from the same twenty trials — the visible measure of how much work the prior is doing.
At a weight of a hundred the right prior’s interval is 0.436 of a flat prior’s width and covers 95%. The over-confident prior’s is 0.456 of a flat prior’s and covers 48.2%. At thirty-five the two shares are 0.641 and 0.661; at four hundred, 0.234 and 0.246. The width tells the reader how heavily the prior weighs against the data, and that number is almost the same whether the weight is justified or not, because it depends on the prior’s weight and the sample size and hardly at all on where the truths actually are. What a prior is worth established that a prior’s weight is a stated number of observations. A reader can see the number; nothing in the interval says whether the world agrees with it.
What the data can say
The data can say something, through the same check the earlier essay used: the prior-predictive distribution of the count, which the prior implies before any data arrives. A study whose count falls in that distribution’s 5% tail is a study the prior found surprising.
The check’s false-alarm rate under a correct prior sits between about 3% and 7% across the weights, because counts are discrete and the 5% tail is not exactly 5%. Against the over-confident priors it fires more often, and not much more. At a weight of twenty it flags 14.2% of studies, at thirty-five 17.3%, and from a hundred upwards 21.9% — a ceiling, because once the prior is confident enough its predictive distribution is nearly binomial at 0.2, and the counts in its tail are the same counts whatever the weight. Four in five studies under a prior ten times too confident pass the check, and the interval of each one covers its truth less than half the time.
That is the contrast with the location failure. A prior centred in the wrong place made the observed data look impossible — a probability of 10⁻⁸ in the earlier essay — so a predictive check flagged it on essentially every study. A prior wrong about its spread makes most data look perfectly ordinary, because the population’s typical study has a truth near the centre and produces the count the prior expects. Over-confidence is detectable only in the studies whose truths are far from the centre, and those are a minority in any one study’s eyes, even though they are where all the under-coverage is.
The remedy the numbers point to is the one a single study cannot apply and a collection of studies can: estimate the spread from the studies themselves, as the prior the data estimates did for a single variance. That is what a hierarchical model does, and a prior on the spread measured what it costs when the number of studies is small — with a history that agrees with itself as the warning that a spread estimated from a few agreeing studies can itself be confidently too small.
More trials per study
Larger studies help, slowly. With the prior worth a hundred observations and the population worth ten, eighty trials a study raise the coverage from 48.2% to 55.5%, and three hundred and twenty raise it to 70.0% — still far from 95% at sixteen times the sample. The check improves faster, because a larger study’s count is more informative about where its own truth lies: it flags 46.1% of studies at eighty trials and 55.9% at three hundred and twenty.
And the width stops hiding the problem only in the sense that the prior stops mattering. At three hundred and twenty trials the over-confident interval is 0.904 of a flat prior’s width and the earned one 0.877: both are close to the data’s own interval, and the coverage that is still lost belongs to the studies whose truths lie so far from 0.2 that even three hundred and twenty trials cannot fully undo a hundred observations’ worth of the wrong spread. The recovery rate is the one the earlier essay found for a prior in the wrong place — slow, because a displacement shrinks like one over the sample size and a standard error like one over its square root — applied here to every study whose truth is far from the prior’s centre.
Tails too thin for an unusual study
The second shape failure is about the tails rather than the confidence. Suppose the analyst’s prior has exactly the right bell — centred at 0.2, worth ten, matching the bulk of the population — but one study in ten comes from a different population, centred at 0.6. The prior, built from the usual studies, puts almost no weight there.
With the bell alone, the usual studies are covered 95.0% of the time, as the bell is right for them. The unusual ones are covered 65.6% of the time. The posterior for a study that saw twelve successes in twenty is still tied to a prior that calls 0.6 implausible, so its interval is pulled down from where the data put it and, a third of the time, below the truth. Averaged over the whole population the interval covers 92.1%, which reads as a mild shortfall and is in fact a large failure concentrated in a tenth of the studies.
The prior-predictive check does better here than against over-confidence, because an unusual study produces an unusual count: it flags 66.4% of the unusual studies and 3.2% of the usual ones. That is a real signal, and it is the flagged studies whose intervals most need replacing. But a third of the unusual studies pass the check, and those are the ones the bell misplaces least visibly.
A small flat component, and what it costs
The standard repair is a robust prior: mix a small share of a vague component into the confident one, so that a study whose data are far from the bell can be explained by the vague part instead of being dragged back. With a flat Beta(1, 1) as the vague part, the posterior is a mixture of two Betas whose weights the data decide, and its interval is computed exactly as the quantiles of that mixture.
Two per cent of flat prior lifts the unusual studies’ coverage from 65.6% to 84.0%. A tenth lifts it to 89.8%, and three tenths to 92.9%. The usual studies are barely touched — 95.1% at two per cent and 95.4% at a tenth — because for a count near the bell’s expectation the data assign the flat component almost no posterior weight. The cost is width: the average interval is 0.272 wide with the bell alone, 0.278 with two per cent of flat prior, and 0.284 with a tenth. Eighteen points of coverage on the studies that need it, for an interval 2.5% wider on average.
The robust prior also helps against over-confidence, which the bell alone cannot detect.
With the bell worth thirty-five and the same population, the usual studies are covered only 75.9% of the time and the unusual ones 11.2%. A tenth of flat prior brings those to 86.6% and 87.0%, and two per cent alone already brings the unusual studies to 77.6%. The width rises more here than for the right bell, from 0.209 to 0.254 at a tenth, because a prior that is too confident about everything has more studies for the flat part to rescue. The flat component does not make the bell less confident about studies near its centre; it gives studies far from the centre somewhere else to go, and those are the studies over-confidence and thin tails both fail.
What a credible interval’s report should add
The population the prior describes, not just its centre. A prior’s centre can be checked against the data; its spread is a claim about other studies, and the report should say what evidence it rests on.
The width relative to a flat prior’s, read for what it is. It measures how much the prior is doing. It does not measure whether the prior has earned it, and it reads the same for a prior that has and one that has not.
The prior-predictive check, with its power stated. Against a prior in the wrong place it catches nearly everything; against a prior ten times too confident it catches about one study in five. A passed check is weak evidence about spread.
Whether the prior has a vague component. Two per cent of it buys most of the protection against an unusual study at almost no cost, and an analysis that reports a single-bell prior has chosen to be fragile in a direction its data mostly cannot reveal.
Exact, over a stated population
Every coverage, width and flag rate here is an exact sum over the twenty-one possible counts, weighted by the population’s own probability of each count, with each coverage term a Beta probability under the population’s posterior. The population of truths is a Beta centred at 0.2 and worth ten observations, or that with one study in ten replaced by a draw from a Beta centred at 0.6 worth ten; the analyst’s priors are Betas centred at 0.2, optionally mixed with a flat Beta. The credible interval is the equal-tailed 95% interval of the posterior, found by bisection on the mixture’s distribution function where the prior is a mixture. The check flags a count in the prior-predictive tail at or below 5% in the direction observed. One sample size, twenty trials, is measured throughout; at larger samples the data overrule an over-confident prior sooner, by the arithmetic the earlier essay derived, and the coverages above are the small-sample case where the prior’s shape matters most.
Still open: how much flat prior to mix in
The mixture’s share was swept, not chosen. Every share above two per cent buys protection, and every share costs width and a little efficiency for the usual studies, so somewhere there is a share that is best for a given population — and the analyst does not know the population. The choice has the shape of a level set before the source is known: a minimax share across plausible populations, or a share that makes the worst coverage over a family of populations as high as it can be, or a share estimated from the collection of studies.
Whether a minimax share exists and is small, how much width it costs on the usual studies, and whether a share estimated from a dozen studies does as well as a fixed one, are measurements the same exact sums can make. They also bear on what a credible interval covers, which measured the credible interval as a frequentist procedure under a reasonable prior; a robust prior is the version of that procedure built to stay reasonable when the prior is not.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A count that bets against its interval — both name binomial proportion, coverage, credible interval, exact enumeration, prior
- An interval that covers and says nothing — both name binomial proportion, coverage, exact enumeration, interval width
- The check worth more than the check — both name binomial proportion, coverage, exact enumeration, interval width
- The interval that integrates — both name coverage, credible interval, prior
- The shortest interval, and the one that does not move — both name coverage, credible interval, interval width
- What a two-unit study should report — both name coverage, interval width, prior sensitivity
Named objects
A flat tag is an object no other essay names yet.
Beta distributionBinomial proportionCoverageCredible intervalExact enumerationHeavy tailInterval widthMixture priorPriorPrior predictivePrior sensitivityRobust prior