What a credible interval covers
Worth reading first: What a prior is worth.
A 95% credible interval says the parameter lies inside it with probability 0.95, given the data and the prior. That is the statement readers wrongly attribute to a confidence interval, and here it is meant literally.
It is also not a coverage claim. Nothing about a credible interval promises that 95% of such intervals contain the truth across repeated samples. So the obvious question is what its coverage actually is — and it is a question with an exact answer.
The measurement is the same one
For a proportion, the coverage of any interval procedure is a finite sum: n + 1 possible counts, each with a known binomial probability, each producing exactly one interval. Add up the probabilities of the counts whose interval contains the truth.
Nothing in that recipe cares where the interval came from. A Wald interval, a Wilson interval and a credible interval from a Beta prior are all just rules mapping a count to a pair of endpoints, and the same sum applies to all three.
That is what makes this field comparable with the rest of the site rather than a separate argument. The two schools are put on one axis, and neither is simulated.
What the reasonable priors do
At a sample of twenty, with the coverage computed exactly at four true proportions:
| prior or method | p = 0.05 | p = 0.1 | p = 0.3 | p = 0.5 |
|---|---|---|---|---|
| Jeffreys, Beta(½, ½) | 98.4% | 95.7% | 94.7% | 95.9% |
| flat, Beta(1, 1) | 92.5% | 95.7% | 97.5% | 95.9% |
| Wald | 63.9% | 87.6% | 94.7% | 95.9% |
| Wilson | 92.5% | 95.7% | 97.5% | 95.9% |
The first row is the headline. A Bayesian interval on Jeffreys’ prior covers 95.7% where the textbook frequentist interval covers 87.6%, and at a proportion of 0.05 it covers 98.4% against the Wald interval’s 63.9%.
That is not the result either camp’s rhetoric predicts. The frequentist objection to Bayesian intervals is that they lack a coverage guarantee; the measurement is that on a standard non-informative prior they have better coverage than the frequentist interval in universal use.
Jeffreys’ prior is the usual default for exactly this reason. It is chosen to be invariant under reparameterisation — the answer does not depend on whether the proportion or its log-odds is treated as the parameter — and the good frequentist behaviour is a consequence rather than the design goal.
The near-coincidence worth noticing
The flat prior and Wilson’s interval agree at every proportion in that table, to every digit shown. They are not the same interval: at a count of zero the flat-prior credible interval is [0.0012, 0.1611] and Wilson’s is [0, 0.1611].
Sweeping two hundred proportions, they differ at ten of them, and where they differ the gap reaches 0.165 — a large disagreement confined to a few places. Both facts come from the same source. Coverage is a sum over discrete counts, so two intervals that include the same counts have identical coverage however different their endpoints are, and the disagreements appear only where a boundary crossing falls between them.
The near-agreement is not a coincidence in the deeper sense. Wilson’s interval inverts the score test, and inverting a test is closely related to what a flat prior does. But “closely related” is doing real work in that sentence, and the honest summary is that they are cousins that mostly agree and come apart at the boundary — which is exactly where the count is zero and where the practical consequences are largest.
Where Wald’s thirty-six points go
The 63.9% at a proportion of 0.05 has a single cause, and naming it turns a bad row into a legible one.
With twenty observations at p = 0.05, the count is zero with probability . A Wald interval at a count of zero is , which is the single point zero — it has no width at all, and it cannot contain 0.05 or any other positive proportion.
So 35.85% of the sample space is lost on one count, before any question of the interval being too narrow at the counts where it exists. That leaves 64.15%, against the 63.9% the exact sum reports, so essentially the whole of Wald’s shortfall at this proportion is the degenerate interval and almost none of it is the approximation being poor elsewhere.
That is worth separating because the two failures have different remedies. A too-narrow interval is repaired by a better variance expression or a continuity correction; a zero-width interval at a boundary count is repaired only by not using an expression that collapses there, which is exactly what adding half an observation to each cell — Jeffreys’ prior — does.
Which of the two good rows is better depends on what is being asked
Jeffreys and the flat prior are both defensible and the table separates them, in a way that does not produce a winner.
Taking the absolute distance from 95% at each of the four proportions: Jeffreys is out by 3.4, 0.7, 0.3 and 0.9 points, averaging 1.33. The flat prior is out by 2.5, 0.7, 2.5 and 0.9, averaging 1.65. Wald is out by 31.1, 7.4, 0.3 and 0.9, averaging 9.93.
So Jeffreys is closer on average and the flat prior has the smaller worst case — 2.5 points against 3.4. The two are not ranked; they are two answers to two different questions, and which one a reader wants depends on whether they are protecting an average over the proportions they might meet or a single proportion they are worried about.
The direction of each miss is the more useful half. Jeffreys is over-covering at 0.05, by 3.4 points, which is conservative. The flat prior is over-covering at 0.3 and under-covering at 0.05, by two and a half points each way. On an interval whose whole purpose is a guarantee, an over-coverage of three points and an under-coverage of two and a half are not the same size of mistake, and the average treats them as though they were.
What a bad prior costs
The other half of the table is the part that answers the objection honestly.
A confident prior centred at a half — Beta(20, 20), worth forty observations — covers 0.0% at a true proportion of 0.05, and 41.6% at 0.3. A confident prior centred in the wrong place, Beta(30, 5), covers 0.0% almost everywhere and 0.6% at a proportion of 0.5.
Zero. Not “poor”: across every one of the twenty-one possible samples, the interval fails to contain the truth. With forty observations’ worth of prior insisting the proportion is near a half, twenty observations cannot move the posterior far enough for its 95% region to reach 0.05.
So the objection to priors is correct, and it is correct in a specific and bounded way. A prior that is both confident and wrong produces intervals that never cover, and the failure is not subtle. What the measurement adds is the boundary: the same machinery on a prior nobody would object to produces intervals better than the frequentist standard.
The distinction is between having a prior and having a confident wrong one, and it is the distinction the general objection elides.
Why this is not a criticism of the method
An important piece of bookkeeping, because it would be easy to read the figure as a scoreboard.
Coverage is a frequentist criterion. A credible interval does not claim it, and measuring a method against a property it never advertised is the kind of move this site complains about elsewhere. A Bayesian is entitled to say that the confident wrong prior gave exactly the right answer given that prior, and that the prior was the error.
Three reasons the measurement is worth making anyway.
Readers assume it. A 95% interval is read as covering 95% of the time by nearly everyone who encounters one, whatever the label. A number that will be misread is worth knowing.
It is the only common axis. Comparing methods requires a criterion both can be evaluated on, and coverage is the one available. The alternative is that the two schools never meet, which is approximately what has happened.
And it converts the prior from a stance into a component with a cost. The confident-wrong prior’s 0% is what makes “state the prior and show the answer under another one” a demand with force behind it rather than a courtesy.
Width, because coverage is never the whole criterion
The same trade the rest of this site insists on. An interval covering 100% of the time is available at no cost by reporting [0, 1], so coverage alone ranks nothing.
At a sample of twenty and a true proportion of 0.15, the credible interval on Jeffreys’ prior reaches its coverage on less mean width than Clopper–Pearson, the conservative exact frequentist interval. Both hold the level; one spends less resolution doing it.
That is the argument for the Bayesian interval in practice, and it is a narrowly quantitative one. It is not that the statement it makes is more natural, though it is. It is that on the frequentist’s own two criteria — cover at the stated rate, and be as short as possible while doing so — a credible interval on a standard non-informative prior is at least competitive and often better.
Where the coverage claim stops being available
One boundary, stated because the whole essay has leaned on the exact sum.
The sum works because a proportion has a finite sample space and a conjugate prior. Neither holds generally. For a continuous parameter with an arbitrary prior the posterior has no closed form, the interval comes out of a sampler, and its frequentist coverage can only be estimated by simulation — with the Monte Carlo error that implies.
So the sharpness here is a property of the case, not of the field. What survives generalisation is the habit rather than the arithmetic: a credible interval’s coverage is a real quantity, it is worth measuring, and the answer depends on the prior in a way that can be reported.
And the specific finding survives too, because it is about the case that comes up most. A proportion in a small sample is the single most common inferential situation in applied work, and in that situation the interval taught first covers 87.6% while a Bayesian interval on a default prior covers 95.7%.
That last figure carries the closing point. Increasing the sample from twenty to eighty repairs the reasonable priors and does very little for the confident wrong one — because the prior is worth thirty-five observations and eighty is not enough to overwhelm a prior that puts almost no mass where the truth is.
How much data it takes is the arithmetic from the previous essay, and this is what it looks like when the answer is “more than the study collected”.
The shape of the failure, not just its size
The coverage table gives four points. The curve gives the shape, and the shape says something the points cannot.
For the reasonable priors the curve oscillates a point or two around 95%, in the same jagged way every interval for a proportion does, because the sample space is discrete and no choice of prior changes that. The wobble is a property of counting, not of philosophy, and it appears identically in the frequentist and Bayesian curves.
For the confident wrong prior the curve is not oscillating around anything. It is flat at zero across most of the range and rises to a peak wherever the prior happens to sit. That is a qualitatively different picture, and the difference is diagnostic: a procedure whose coverage wobbles near nominal has a discreteness problem, and one whose coverage is zero over a region has a prior problem.
The practical value of the distinction is that the second is fixable and the first is not. Nobody can remove the oscillation. Anybody can widen a prior.
What the confident prior is actually doing
It is worth being concrete about the mechanism, because “the prior was wrong” is the kind of explanation that explains nothing.
Beta(30, 5) is worth thirty-five observations and has a mean of about 0.857. Against twenty observations at a true proportion of 0.1, the posterior is Beta(32, 23) — mean 0.58, and its central 95% region runs nowhere near 0.1.
The data said 2 successes in 20. The prior said, in effect, 30 successes in 35. The posterior combined them into 32 in 55 and reported that with confidence. Every step is correct, and the answer is nowhere near the truth, because the prior contributed nearly two thirds of the evidence and was wrong.
Nothing about the interval’s construction is defective. The interval faithfully reports the posterior, the posterior faithfully combines the inputs, and one of the inputs was false. That is worth stating plainly because it identifies where the check has to go: not on the interval, and not on the combination, but on the input — which means showing the answer under a prior that is not the one preferred.
What happens as the sample grows
The natural hope is that this is a small-sample problem and disappears with data. It does, and the rate is worth knowing because it is slower than expected in the case that matters.
The reasonable priors are already at nominal by n = 20 and stay there. The confident wrong prior needs enough data to overwhelm thirty-five observations’ worth of false information, and “overwhelm” here does not mean “outweigh slightly” — it means the posterior’s 95% region has to travel far enough to reach the truth.
At n = 80 the confident wrong prior is still failing. That is four times the sample size and more than twice the prior’s weight, and it is not sufficient, because the interval has to move a long way and the prior is still contributing three tenths of the answer.
The general lesson is one this site keeps arriving at: the sample size at which a problem disappears is usually much larger than the sample size at which it stops being visible. The confident wrong prior produces a narrow, confident, wrong interval at every sample size shown, and at no point does the output look distressed.
Reporting, and what would make this checkable by a reader
The field’s demand, made specific to intervals.
Give the prior as parameters. Beta(½, ½), not “non-informative”. A reader can then compute its weight, and the weight against the sample size is the whole question.
Give the interval under a second prior. Preferably one deliberately unsympathetic to the conclusion. The cost is recomputing a Beta quantile.
And say which statement is being made. A credible interval and a confidence interval can have the same endpoints — they do exactly, for a mean with a flat prior — and mean different things. A reader cannot tell from the numbers which one is intended.
None of this requires resolving the disagreement between the schools, and none of it requires believing the prior. It requires the prior to be visible, which is the only thing that lets a reader do the arithmetic this essay has been doing.
The measurement this field was built to make
Standing back from the table, the finding is a single sentence with two halves that are usually stated by opposing sides.
A credible interval on a standard non-informative prior has better frequentist coverage than the frequentist interval in universal use — 95.7% against 87.6% at n = 20 and p = 0.1, and 98.4% against 63.9% at p = 0.05. And a credible interval on a confident wrong prior has essentially no coverage at all, failing on every one of the twenty-one possible samples.
Both halves come from the same exact sum, computed the same way, and neither is an argument. Whichever side of the dispute a reader arrived on, one of those numbers is unwelcome — which is the usual sign that the measurement was worth taking rather than the position worth restating.
A note on what was compared
One boundary on the comparison above, since it has been made repeatedly and could be over-read.
The frequentist intervals here are the four in common use for a proportion, and the Bayesian intervals are equal-tailed credible intervals from Beta priors. Neither list is exhaustive, and a determined frequentist can construct intervals with better properties than Wald — Wilson is one, and it appears in the table performing well.
The claim is not that one school wins. It is that on this problem, at these sample sizes, the default Bayesian answer outperforms the default frequentist answer by a wide margin, and both defaults are what actually gets used.
That matters more than a like-for-like tournament between best-available methods, because the interval a reader meets in a paper is the one the software produced by default — and the defaults are what this table compared.
A comparison of best-available methods is a different and less useful exercise, because the winner of it is not what anybody runs.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A prior on the spread — both name flat prior, posterior, prior, prior sensitivity
- What the plug-in forgets — both name coverage, flat prior, posterior, prior
- The p-value a replication gets — both name flat prior, posterior, prior
- A block size that changes — both name coverage, sample size
- A coverage table with its own error — both name coverage, sample size
- A hole no sample size fills — both name coverage, jeffreys' prior
Named objects
A flat tag is an object no other essay names yet.
CoverageCredible intervalFlat priorJeffreys' priorPosteriorPriorPrior sensitivitySample size