The prior, doing visible work

When the prior is confident and wrong

A prior worth thirty-five observations, centred in the wrong place, produces a 95% interval that covers nothing at all — and reports a width 5% narrower than an honest one. It takes seventeen thousand observations to repair, not thirty-five, and the worst study to run is the one whose sample size equals the prior's weight, exactly.

Worth reading first: What a credible interval covers · What a prior is worth.

A prior is worth a stated number of observations: for a proportion, exactly a+ba + b of them. That number is the field’s most useful result and it is usually read as a reassurance. A prior worth thirty-five observations is overwhelmed by a study of a few hundred, so a disagreement between the prior and the world is a temporary problem that data fixes.

Every clause of that is true about the posterior mean. None of it is true about the interval, and the difference is three orders of magnitude.

A prior worth 35 observations, moved across the range — truth 0.1, n = 20The same prior weight centred at each of 33 places. Its interval covers 100.0% where the centre is near the truth and 0.0% at its worst, while the mean width where it covers least is 0.221 against a flat prior's 0.263 on the same data.00.2500.5000.75010.2000.4000.6000.800where the prior is centredcoverage, and width as a share of the widestthe truth, 0.1coveragea flat prior's width here, 0.263width, widest is 0.261exact sum over 21 outcomes at each centreone of these two is on the page a reader sees
Fig. 1 One prior weight — thirty-five observations — centred at each of thirty-three places, against a truth of 0.1 and twenty trials. The coverage runs from 99.96% to 0.00%. The width, which is the quantity a reader sees, never leaves the neighbourhood of an honest interval’s.

Coverage is a function of a distance

The shape of the coverage curve is the essay’s first result and it is more brutal than “coverage falls”.

At a centre of 0.1, where the prior agrees with the truth, the interval covers 99.96% — better than nominal, because a correct prior is extra information and the interval is correspondingly short. At a centre of 0.3 it covers 12.16%. At a centre of 0.5 and everywhere above it, it covers 0.00%.

Not “poorly”. Zero. Across all twenty-one possible datasets, at a true proportion of a tenth, a 95% interval built on a prior worth thirty-five observations centred at a half contains the truth on none of them.

That is not a quirk of one setting. Once the prior’s weight is comparable to the sample size and the conflict is more than a few standard errors, the posterior is pinned near the prior and the interval around it is short — so the interval is somewhere the truth is not, every time, with no dataset in the sample space capable of dragging it back.

A confidence interval cannot do this. Its coverage is a function of the sample size and the true proportion, and at a tenth and twenty trials the Wald interval covers 87.6% however the study came out. A credible interval’s coverage is a function of the sample size, the true proportion and how far the prior is from it — and the third argument is unbounded.

The width says nothing

The second result is the one that makes the first dangerous, and it is visible in the same figure.

At a centre of 0.5 the interval covers nothing and its mean width is 0.2491. A flat prior on the same data gives a width of 0.2632 and covers 95.68%.

The interval that covers nothing is 5% narrower than the interval that covers what it claims. Not wider, which a reader might have read as caution, and not obviously anomalous in either direction. Across every prior weight on the slider, every centre whose interval has stopped covering still reports a width within a third of a flat prior’s on the same data — and usually inside ten per cent of it.

So the output of the failing analysis is a proportion, an interval of ordinary length, and a stated 95%. There is nothing in it to look at.

That is a shape that has now turned up in four unrelated places — a tick function returning an empty array, a class no stylesheet defines, a variance estimate returning its own boundary, and here a prior overriding the data — and the common feature is not that a number is wrong. It is that the output is well-formed, so every check that reads the output passes.

A weak prior in the wrong place is a different object

Before the mechanism, the obvious control: whether any of this is about confidence rather than about priors in general.

A prior worth 5 observations, moved across the range — truth 0.1, n = 20. The same prior weight centred at each of 33 places. Its interval covers 98.9% where the centre is near the truth and 39.2% at its worst, while the mean width where it covers least is 0.324 against a flat prior's 0.263 on the same data.
Fig. 2 The same sweep with a prior worth five observations rather than thirty-five. Its worst coverage is 39.2% — bad, and not zero — and its width where it covers least is 0.324 against a flat prior’s 0.263, which is the one case in this essay where the failing interval is visibly the wider one.

A weak prior in the wrong place degrades the interval and does not destroy it, which is what a weight of five against a sample of twenty should do. And the width behaves differently too: a weak prior centred near a half pulls the posterior towards the middle of the range, where a Beta posterior is at its widest, so the interval gets longer rather than shorter.

That reversal is worth noticing because it removes the only rule of thumb this essay might otherwise have offered. A failing interval is not systematically narrow or systematically wide; which it is depends on the prior’s weight and on where the truth sits relative to a half. There is no direction to watch.

The worst study to run is the one the prior is the size of

The obvious follow-up question is which studies are at risk, and the answer is not the one the “worth thirty-five observations” framing suggests.

Where a prior worth 35 observations does the most harm. The posterior mean is displaced from the truth by w(c − p)/(w + n) and the interval round it has a width of order the square root of p(1 − p) over n, so the displacement measured in standard errors grows like the square root of n over (w + n). That rises and then falls, and it peaks at n = 35 against a prior worth 35 — the sample size at which the data is worth exactly what the prior is worth. The peak here is 7.40 standard errors.
Fig. 3 The displacement of the interval from the truth, measured in the standard errors it will be compared against, plotted against the sample size. It rises, peaks and falls, and the peak is at n = 35 — the sample size at which the data is worth exactly what the prior is worth.

Two quantities are in play and they shrink at different rates. The prior displaces the posterior mean by w(cp)/(w+n)w(c-p)/(w+n), which falls like 1/n1/n. The interval’s own half-width is about p(1p)/n\sqrt{p(1-p)/n}, which falls like 1/n1/\sqrt n. What decides coverage is the ratio — the displacement measured in standard errors — and that goes as

wcpn(w+n)p(1p)\frac{w\,|c-p|\,\sqrt n}{(w+n)\sqrt{p(1-p)}}

which is zero at n=0n = 0, zero in the limit, and maximised in between. Differentiating puts the maximum at

n=wn = w

exactly, with no dependence on where the prior sits or on what the truth is. It holds at every weight the slider reaches: a prior worth five peaks at five observations, one worth eighty peaks at eighty.

The reading is worth stating plainly because it inverts the usual intuition. A tiny study is comparatively safe: the posterior is nearly the prior, and a prior’s interval is wide enough to contain a good deal. A large study is safe: the data has won. The study at risk is the one whose sample size is about what the prior is worth — which is, unhelpfully, the size of study a researcher who elicited that prior is most likely to be running.

At a prior worth thirty-five centred at 0.85 against a truth of 0.1, the peak displacement is 7.40 standard errors. An interval displaced by 1.96 standard errors covers about half the time. One displaced by seven covers never.

How much data actually repairs it

Since the displacement in standard errors falls like 1/n1/\sqrt n rather than 1/n1/n, the sample size that repairs the interval is not a small multiple of the prior’s weight.

Recovering from a prior centred at 0.85 when the truth is 0.1. A prior worth 35 observations, centred 0.75 away from the truth. Its 95% interval covers 0.0% at 10 observations and does not reach 90% until 17409 — 497 times the prior's own weight, because what must shrink is the displacement in standard errors and that falls only as one over the square root of the sample size.
Fig. 4 Coverage against sample size, each tick four times the last, for a prior worth thirty-five observations centred at 0.85 against a truth of 0.1. It is still at 8.1% at six hundred and forty observations and does not reach 90% until 17,409. The flat prior’s curve, for comparison, sits at its nominal level throughout.

Seventeen thousand four hundred and nine observations to repair a prior worth thirty-five. That is 497 times the weight, and the multiplier is not a constant: it grows with the conflict and with the weight, because both enter the displacement linearly while the repair is quadratic in it.

The closed form says why. An interval centred dd standard errors off the truth covers Φ(zd)Φ(zd)\Phi(z-d) - \Phi(-z-d), so 90% coverage tolerates d=0.6524d = 0.6524 standard errors and no more. Setting the displacement equal to that and solving gives

n(wcp0.6524p(1p))2n^{*} \approx \left(\frac{w\,|c-p|}{0.6524\sqrt{p(1-p)}}\right)^{2}

which is 17,991 here — within 3.3% of the counted 17,409, and above it, which is the direction dropping the ww beside the nn in the denominator pushes the estimate.

The two routes agree where the displacement is large and part company where it is small. At a prior worth five the counted crossing is 206 and the closed form says 367; at a centre of 0.4 the counted crossing is 2,294 and the closed form says 2,878. That is the asymptote behaving like an asymptote: it is derived by treating w+nw + n as nn, which is wrong by a factor when nn is small, and the counted number is the one to quote. A closed form that agrees everywhere would be suspicious; one that agrees where its own derivation applies is evidence.

Five priors, one picture

The field’s first coverage figure already contained this result and nobody read it that way.

What a 95% credible interval covers, n = 20. Computed by summing over all 21 possible counts rather than by simulating them. Jeffreys' prior covers close to 95% across the range; a confident prior centred in the wrong place covers almost nothing where the truth is far from it.
Fig. 5 Four named priors, their coverage summed over the sample space at each true proportion. Three of the curves sit near the nominal line across the range. The fourth — a prior worth thirty-five centred at 0.857 — reaches zero over a wide stretch of the left, which is the region where proportions are usually reported.

That figure was drawn to make the point that a credible interval’s coverage is a checkable number, and the wrong prior was in it as a demonstration that the number can be bad. What was never asked is what governs the curve’s shape, and the answer is the distance argument above: the curve is a function of cp|c - p| measured in standard errors, and the flat stretch at zero is where that distance exceeds about four.

Two things about the curve are worth reading off it now. The zero region is wide — it is not a point failure at one unlucky proportion but most of the lower half of the range. And the curve recovers above nominal where the truth passes near the prior’s centre, to 99.96%, because a correct confident prior is genuinely extra information. The same prior is better than everything else in the figure at one proportion and worse than anything a frequentist could construct at most of the others.

The sentence this qualifies

It has already been said here that the prior washes out.

Four priors meeting the same data, true proportion 0.25. A prior is worth exactly a + b observations: Jeffreys — Beta(½, ½) is worth 1, weakly informative — Beta(2, 2) is worth 4, confident, centred at 0.5 — Beta(20, 20) is worth 40, confident and wrong — Beta(30, 5) is worth 35. Each curve is the posterior mean as the data accumulates, and every one of them converges on 0.25.
Fig. 6 The earlier measurement, unchanged: four priors of different weights meeting the same data, and every one of their posterior means converging on the truth. Nothing in this figure is wrong and it is a statement about a point estimate.

The posterior mean does converge, at the rate that figure shows, and the rate is 1/n1/n. The essay that measured it was about the posterior mean and said so.

What the interval adds to that is a different schedule of convergence, because an interval is a statement about a displacement relative to a spread and the spread is shrinking too. By the time the displacement is small in absolute terms it is not small compared with a standard error that has shrunk faster. A reader who takes “the prior washes out” from the point estimate to the interval has carried a true statement across a boundary it does not cross.

The general form is worth having, because it is not about priors: a bias that falls like 1/n1/n is negligible for an estimate and not negligible for an interval, since the interval’s own scale falls like 1/n1/\sqrt n. The same arithmetic decides whether a bias correction is worth making in the forecast field, and whether a small misspecification matters for a standard error built on a wrong model.

Zero is a statement about the sample space, not about bad luck

A coverage of exactly 0.00% deserves a sentence of its own, because it is a stronger statement than any simulation could make and it is easy to read as a rounding.

The number is a sum over all twenty-one possible datasets. For each one, the interval is built and asked whether it contains 0.1. At a prior worth thirty-five centred at a half, the answer is no twenty-one times. There is no dataset this study could have produced whose interval would have contained the truth — not an unlikely one, not one in a thousand. The sample space does not contain a success.

That is only available because the arithmetic here is exact. A simulation of ten thousand studies would have reported 0.00% too, and would have meant “none of the ten thousand”, which is a different and weaker claim: it leaves open a rare dataset that the run missed. Summing over the sample space closes it, and closing it is the difference between this fails almost always and this cannot succeed — which is the reason summing a coverage over every count rather than simulating it.

The mechanism is the same displacement arithmetic. The most extreme dataset available — zero successes in twenty — gives a posterior Beta(17.5, 37.5), whose 95% interval runs from 0.2031 to 0.4458. The truth is 0.1. The most favourable outcome in the entire sample space misses by more than the interval’s own width, so every other outcome misses by more.

This is worth separating from the usual way an interval fails. A Wald interval at a small proportion fails on some datasets and the failure rate is the coverage; there is always a dataset that would have worked. Here the procedure’s whole range of outputs has been moved off the truth, which is a misspecification and not a miss.

What would have caught it

Nothing above is an argument against priors, and the field would be in poor shape if the conclusion were that a confident prior is too dangerous to use. The conclusion is narrower: the interval is the wrong place to look for the failure, and there is a right place.

What a prior centred at 0.85 expects to see in 20 trials. The prior predictive distribution of the count, in closed form from the Beta–binomial. It expects 17.0 and gives the observed 2 a probability of 1.55e-8; a count at least this extreme has probability 1.68e-8. The interval built from this prior carries none of that.
Fig. 7 What a prior centred at 0.85 expects to see in twenty trials, before any data arrives. It expects seventeen successes. It gives the observed two a probability of 1.5 × 10⁻⁸, and a count at least that extreme a probability of 1.7 × 10⁻⁸.

The prior predictive distribution — the distribution of the data implied by the prior, averaging over every value of the parameter it allows — is available in closed form for this model and needs nothing that was not already specified. It is the Beta–binomial, it takes one line, and it makes the conflict unmissable: the study that produced a zero-coverage interval was a study the prior considered impossible.

The same check passes where it should. Under a prior centred at the truth the observed two has probability 0.2255 and a tail of 0.6750 — an ordinary outcome, no flag. Under a prior centred at a half it is 1.6 × 10⁻³, which is the correct intermediate answer for an intermediate conflict.

So the failure is detectable, cheaply, before any interval is formed, using only the objects the analysis already contains. It is simply not detectable from the interval, which is the only thing most analyses report.

Where this leaves the earlier results

A credible interval has now been measured against the frequentist’s own criterion four times over, and the results have gone in both directions. Putting them together is the honest summary and it is not the summary either camp’s rhetoric offers.

Under a reasonable prior the credible interval beats the textbook interval on coverage — 95.7% against 87.6% at twenty trials and a tenth — and it keeps that advantage through a change of variable where the frequentist repairs do not. Under a confident prior in the wrong place it covers nothing at all, for as many observations as most studies will ever collect.

Both are consequences of the same property: a credible interval uses the prior. The interval’s quality is the prior’s quality, transmitted without attenuation, and there is no prior-free part of it to fall back on. What the frequentist interval offers instead is not accuracy but a guarantee that does not depend on an input — a floor rather than a ceiling.

Stated that way the choice stops being philosophical. It is a question about whether the prior is checkable, and the prior predictive above is the check. A prior that has been put against the data it implies and survived is an asset worth eight points of coverage. One that has not is an unexamined assumption with an unbounded downside and no symptom.

Three priors on the spread, at a true τ of 1. The posterior for τ under a flat prior (mean 2.18), a half-Cauchy of scale 1 (1.78) and one of scale 0.25 (1.69). The three answers differ by 23% of the widest. The prior does visible work when eight groups cannot separate a small spread from none, and almost none when they can.
Fig. 8 The same question one field over, where the prior is on a population spread rather than a proportion and cannot be dispensed with: three priors give posterior means of 2.18, 1.78 and 1.69, differing by 23% of the widest. There the prior is doing visible work because eight groups cannot settle the quantity, and no amount of care about the interval removes the dependence.

What is claimed here, and what is not

The claim is how a credible interval for a proportion behaves as a function of the distance between its prior and the truth: that coverage falls to zero rather than merely falling, that the width carries no signal of it, that the displacement in standard errors peaks at a sample size equal to the prior’s weight, that recovery takes hundreds of times the prior’s weight rather than a few multiples of it, and that the prior predictive detects the conflict before any interval exists.

Every coverage and width is an exact sum over the sample space. The two recovery numbers are a counted crossing and a closed form that share only the four inputs.

What stays out: hierarchical priors, where the prior’s parameters are themselves estimated and the conflict has somewhere to go — the field next door measures what a prior on a spread costs, and that is a different mechanism; robust priors with heavy tails, which bound the damage at the cost of the efficiency that made a confident prior attractive; and formal prior-data conflict diagnostics, of which the prior predictive tail used here is the simplest and not the best. The measurements are also all at one sample size for the coverage curves and one truth for the recovery curves; the peak result is the only one here that holds at every setting, and it holds because it is an identity rather than a measurement.

Still open: a prior that is confident about the wrong thing

Every conflict in this essay is a conflict of location — the prior centred somewhere the truth is not. A prior can also be wrong about shape: correctly centred and far too confident, or correctly centred with tails too thin for a truth that is genuinely unusual.

Those are different failures and the displacement arithmetic above says nothing about either. A prior centred exactly at the truth contributes zero displacement whatever its weight, so the whole of this essay’s arithmetic reports no problem — and an over-confident prior still produces an interval too short for its own uncertainty, which is a coverage failure of a different kind and one the width would, for once, show. Whether it shows enough to act on is a measurement nobody here has made.

The check, and the refusal

Two claims are gated, and they are the two halves that have to hold together.

That a confident prior far from the truth covers a small fraction of what it claims, and that the same prior in the right place covers what it claims — both required, because either alone is unremarkable and the pair is the finding. And that every centre whose interval has stopped covering still reports a width within a third of a flat prior’s on the same data, required across the whole sweep rather than at the single worst centre, because a width claim read off one setting is a fact about that setting dressed as a fact about the family — the commonest way a claim of this shape goes wrong.

The refusal is the one that keeps the essay from being a counsel of despair: nothing can detect a wrong prior, rejected. The prior predictive tail is required to fall below a millionth on the conflicting data and to stay above 5% on data the prior expects. A diagnostic that fired on everything would be no diagnostic, so the second half is the load-bearing one — and it is the half that a check written to demonstrate a failure would have been most likely to omit.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CoverageCredible intervalFlat priorInterval widthModel misspecificationPosteriorPriorPrior predictivePrior sensitivitySample size