An interval read beside something else

Two intervals that overlap

Two 95% intervals that just touch are read as a difference at the edge of significance. With equal standard errors their difference has p = 0.0056, not 0.05; two intervals can overlap by 58.6% of an arm and still differ at exactly 5%; standard-error bars that just touch mark p = 0.157; and when the two estimates are correlated at 0.8, touching intervals conceal a difference of 6.2 standard errors. Read as a test, non-overlap needs 1.66 times the sample for the same power.

Worth reading first: Twenty intervals and one expected miss.

Five times in six found that a 95% interval captures a same-sized replication’s estimate only 83.42% of the time, because the difference of two independent estimates has a standard error 2\sqrt{2} times either one’s. The same fact governs a reading made far more often than a replication: two intervals drawn side by side — treated against control, one year against the next, one subgroup against another — and a verdict taken from whether they overlap.

The verdict is usually one of two sentences. “The intervals overlap, so there is no significant difference.” “The intervals do not overlap, so the difference is significant.” The second is true and badly calibrated. The first is false.

Two 95% intervals 2.772 standard errors of the difference apart, standard errors in the ratio 1The intervals are separated, and the test of the difference gives p = 0.0056. Two 95% intervals with equal standard errors just touch at p = 0.0056.first groupsecond groupdifference 2.772 standard errors · p = 0.0056the intervals do not overlapindependent estimates, normal samplingtouching means p = 0.0056
Fig. 1 Two independent estimates with equal standard errors and their 95% intervals, set 2.772 standard errors of the difference apart, where the intervals just touch. The p-value of the difference is printed below. The slider moves the estimates apart or together.

Where touching intervals put the difference

Two independent estimates with standard errors se1\mathrm{se}_1 and se2\mathrm{se}_2 have 95% intervals that just touch when their difference is 1.96(se1+se2)1.96(\mathrm{se}_1 + \mathrm{se}_2). The difference itself has standard error se12+se22\sqrt{\mathrm{se}_1^2 + \mathrm{se}_2^2}. So, with r=se2/se1r = \mathrm{se}_2/\mathrm{se}_1, touching intervals sit

ztouch  =  1.96(1+r)1+r2z_{\text{touch}} \;=\; \frac{1.96\,(1 + r)}{\sqrt{1 + r^2}}

standard errors of the difference apart. At equal standard errors that is 1.9621.96\sqrt{2} = 2.772, a two-sided p-value of 0.0056.

Touching 95% intervals mark a difference significant at about one half of one per cent, not at five. The intervals were each built to hold a parameter 95% of the time, and the event “both hold their own parameter” is not the event “the difference is zero is rejected at 5%”; the sum of two arms is longer than the standard error of the difference, because standard errors add in quadrature and arms add in a straight line.

The figure’s slider shows the rest of the range. At 2.4 standard errors of the difference the intervals overlap by 13% of an interval’s width and the difference has p = 0.0164: significant, with overlapping intervals. At 1.96 they overlap more, and the difference is significant at exactly 5%.

How much overlap a 5% difference has

The practical rule of thumb that follows has an exact version. At the separation where the difference is significant at 5%, equal-sized 95% intervals overlap by 3.92se2.772se=1.148se3.92\,\mathrm{se} - 2.772\,\mathrm{se} = 1.148\,\mathrm{se}, which is 58.6% of one arm — the half-width of either interval.

So two 95% intervals with equal standard errors can overlap by more than half an arm and still differ significantly. The rule “overlap of about half an arm or less means p below 5%” is the rounding of this number, and it is conservative in the direction that protects a reader: an overlap of exactly half an arm is a little past significance, not short of it.

With unequal standard errors the allowable overlap grows. When one standard error is three times the other the overlap at the 5% point is 83.8% of the shorter interval’s arm, because the longer interval contributes most of the length and little of the precision.

The ratio of the standard errors

The unequal case moves the touching p-value as well.

The p-value of a difference when two error bars just touch, by the ratio of their standard errors. Two 95% intervals with equal standard errors touch at p = 0.0056, and at a ratio of ten at p = 0.0319. Two ±1 standard error bars touch at p = 0.157 and 0.274. Neither kind of bar marks 5%.
Fig. 2 The p-value of the difference when two error bars just touch, against the ratio of the two standard errors on a logarithmic axis, for 95% intervals and for bars of one standard error. The dashed line is 5%.
ratio of standard errors separation when 95% intervals touch p-value
1 2.772 0.0056
2 2.630 0.0085
3 2.479 0.0132
5 2.306 0.0211
10 2.145 0.0319

As one estimate becomes much more precise than the other, the touching point approaches 1.96 standard errors of the difference and the p-value approaches 5%, because the precise estimate is nearly a known constant and the comparison is nearly a one-sample test of the imprecise one against it. At any finite ratio touching still marks something stricter than 5%.

Standard-error bars fail in the opposite direction. Bars of one standard error each touch at a separation of (1+r)/1+r2(1 + r)/\sqrt{1 + r^2} standard errors of the difference — 1.414 at equal standard errors — which is a p-value of 0.157. At a ratio of ten it is 0.274. Two standard-error bars that just fail to overlap mean almost nothing, and a figure that draws them invites a reader to see significance in a difference that is one standard error and a half of its own noise.

The interval that does mark 5%

If an overlap reading is going to be made, the intervals can be drawn to make it correctly.

The confidence level at which two touching intervals mark a 5% difference, by the ratio of their standard errors. At equal standard errors, 83.4% intervals touch exactly when the difference has p = 0.05. At a ratio of three it is 87.9% and at ten 92.7%; it reaches 95% only when one estimate has no uncertainty at all.
Fig. 3 The confidence level at which two touching intervals correspond exactly to a 5% test of the difference, against the ratio of the two standard errors.

The multiplier that makes touching equivalent to z=1.96|z| = 1.96 is 1.961+r2/(1+r)1.96\sqrt{1 + r^2}/(1 + r). At equal standard errors that is 1.386, which is an 83.4% interval — the same 83.42% that measured replication capture, because both are the statement that a difference of two equally precise estimates has 2\sqrt{2} times their standard error. At a ratio of three the level is 87.9%, and at ten 92.7%.

A figure meant for comparing groups by eye can therefore draw 83% intervals and say so, and then “the bars do not overlap” means p<0.05p < 0.05 for equal standard errors, give or take a few tenths of a percentage point at moderate ratios. That is a legitimate design choice. It is also an interval that holds its own parameter only 83% of the time, so a figure that draws them has traded one reading for another and has to say which one it supports.

Correlated estimates

Every calculation above assumes the two estimates are independent. Many comparisons are not: a measurement before and after a treatment on the same people, two conditions tested on the same participants, this quarter’s rate and last quarter’s on a panel that overlaps.

The size of a difference that two just-touching 95% intervals conceal, by the correlation between the estimates. For independent estimates with equal standard errors, touching intervals mark a difference of 2.77 standard errors. At a correlation of 0.5, 3.92; at 0.9, 8.77 — a p-value below 6e-10 already at 0.8.
Fig. 4 The difference, in its own standard errors, at which two 95% intervals with equal standard errors just touch, against the correlation between the two estimates. The dashed line marks the 1.96 of a 5% test.

With correlation ρ\rho the difference has standard error se2(1ρ)\mathrm{se}\sqrt{2(1 - \rho)}, so a positive correlation shrinks it while leaving each estimate’s own interval unchanged. Touching intervals then sit

correlation separation when 95% intervals touch
0 2.77 standard errors of the difference
0.5 3.92
0.8 6.20
0.9 8.77

At a correlation of 0.8, two 95% intervals that just touch conceal a difference with a p-value below one in a billion. A pair of before-and-after intervals that overlap heavily can sit on a paired difference that is overwhelming, and the overlap cannot see it, because each interval describes the variation between people and the paired difference has removed exactly that.

The direction is the dangerous one for designs where correlation is built in, and a shared unit induces it as surely as a shared control group correlates the tests that use it. The correct interval for a paired design is the interval for the difference, and two separate intervals are not a weaker version of it; they answer a question about a different source of variation. For a negative correlation the reverse holds and non-overlap overstates the evidence, but designs that induce negative correlation between the two estimates being compared are rare.

Non-overlap as a test

The true sentence — “the intervals do not overlap, so the difference is significant” — is worth pricing, because its truth is bought by running a much stricter test than the reader thinks.

How often two 95% intervals fail to overlap, against how often the difference tests significant, standard errors in the ratio 1. At no true difference the ordinary test rejects 5.00% of the time and non-overlap occurs 0.56% of the time. At a true difference of 2.8 standard errors — 80% power for the ordinary test — the intervals fail to overlap 51.1% of the time.
Fig. 5 Two ways of calling a difference real, against the true difference in standard errors of the difference: the ordinary test at 5%, and “the 95% intervals do not overlap”.

With equal standard errors, reading non-overlap as significance is a test at level 0.56%. Its power, beside the ordinary test’s:

true difference ordinary test intervals do not overlap
0 5.0% 0.6%
2.0 standard errors 51.6% 22.0%
2.8 standard errors 80.0% 51.1%
3.5 standard errors 93.8% 76.7%
4.0 standard errors 97.9% 89.0%

At the effect the ordinary test detects 80% of the time, the overlap reading detects it half the time. Reaching 80% power by overlap needs a true difference of 3.61 standard errors rather than 2.80, which for a fixed effect is 1.66 times the sample.

So the two verdicts in the opening are two different failures. “Overlapping, so no difference” is wrong about differences between 1.96 and 2.77 standard errors, which are significant and overlap. “Not overlapping, so significant” is correct and spends two thirds more data than it needs to, while quietly running at a level of half a per cent — which is a stricter criterion than almost any study would choose on purpose, applied because it was the one the picture made easy.

A figure of many groups

The two-group arithmetic is the best case for the overlap reading, because a figure with two groups contains one comparison. Most figures with error bars contain more: a bar for each of several treatments, sites, years or subgroups, and a reader scanning for the pair that does not overlap.

In a figure of identical groups, how often some pair of error bars fails to overlap. Every group has the same true mean and the same standard error. With ten groups, some pair of 95% intervals fails to overlap 14.61% of the time, some pair of 83.4% intervals 62.74%, and some pair of ±1 standard error bars 92.32%.
Fig. 6 In a figure of identical groups — the same true mean and the same standard error in each — the chance that at least one pair of bars fails to overlap, against the number of groups, for 95% intervals, 83.4% intervals and bars of one standard error.

Some pair fails to overlap exactly when the range of the estimates exceeds the length of two arms, so the chance is a statement about the range of kk normal draws and has a one-dimensional integral for it. With nothing different between the groups at all:

groups 95% intervals 83.4% intervals ±1 standard error
2 0.56% 5.00% 15.73%
5 4.43% 28.58% 61.84%
10 14.61% 62.74% 92.32%
20 37.54% 91.84% 99.77%

The strictness that makes two touching 95% intervals a test at half a per cent is used up by the number of pairs. Ten groups make forty-five pairs, and a reader scanning a ten-bar figure of identical groups finds a non-overlapping pair about one time in seven. The 83.4% intervals that fix the two-group reading make it much worse here: they are calibrated to 5% per pair, and a figure of ten of them shows a spurious separation more often than not. Standard-error bars, the commonest error bar in print, show one nearly every time.

That is the familiar multiplicity problem wearing a picture instead of a table — twenty analyses of nothing with the analyses done by eye — and it has the familiar consequence. There is no single kind of bar that makes both the two-group reading and the many-group reading correct, because the two readings ask for different error rates. A figure that invites pairwise comparisons among many groups needs intervals built for the family of comparisons, and one that is only meant to show each group’s own uncertainty should say that comparing bars is not what it is for.

A study against the pooled estimate

A forest plot invites the same reading in a form where the correlation is built in and nobody draws it. Each study’s interval is drawn above a diamond for the pooled estimate, and a study whose interval does not reach the diamond is read as an outlier; one whose interval overlaps it is read as consistent.

The pooled estimate contains the study. In a fixed-effect analysis where a study carries a share ww of the total weight, the pooled standard error is w\sqrt{w} times the study’s and the two estimates are correlated at w\sqrt{w} — both facts working in the direction that makes overlap conceal a difference. Touching then happens at

z  =  1.96(1+w)1wz \;=\; \frac{1.96\,(1 + \sqrt{w})}{\sqrt{1 - w}}

standard errors of the difference between the study and the pooled estimate:

the study’s share of the weight separation when the intervals touch p-value
5% 2.46 0.0139
10% 2.72 0.0065
25% 3.39 0.00069
50% 4.73 0.0000022

A study carrying a quarter of the weight can sit 3.39 standard errors away from the pooled estimate with its interval just touching the diamond. The heavier the study, the more it drags the pooled estimate towards itself and the more of the difference the overlap hides — which is backwards from what a reader scanning for inconsistent studies needs, since a heavy inconsistent study is the one that matters most.

The comparison a forest plot is being scanned for is a test of heterogeneity, and it exists in closed form: each study’s deviation from the estimate pooled without it, or the overall statistic that sums the standardised deviations. Groups that borrow from each other are exactly the case where that comparison is the point of the analysis, and where reading it off overlapping lines on a plot answers a different question with a ruler.

Why the arms add in a straight line

The whole essay rests on one geometric fact, and it is worth seeing why it is not a technicality.

Each interval’s arm is a length along the same axis: 1.96se11.96\,\mathrm{se}_1 to the right of the first estimate, 1.96se21.96\,\mathrm{se}_2 to the left of the second. For the arms to meet, the gap between the estimates has to be the sum of those lengths. Nothing about that addition involves probability; it is where two line segments end.

The uncertainty in the difference, on the other hand, comes from two independent sources of error, and independent errors do not add their sizes, they add their variances. Two errors of one standard error each produce a difference whose error is 2\sqrt 2, not 2. The overlap reading takes a probabilistic quantity and measures it with a ruler, and the ruler overstates it by (1+r)/1+r2(1 + r)/\sqrt{1 + r^2} — by 41% at equal standard errors, which is exactly the gap between 1.96 and 2.772.

Correlation changes the variance addition and not the ruler, which is why its effect in the table above is so large. A positive correlation subtracts from the variance of the difference while leaving both arms exactly as long as they were.

What to draw instead

The reading that the overlap is standing in for has a direct picture: the difference between the two estimates, with its own interval.

For two independent groups that interval is the difference plus or minus 1.96se12+se221.96\sqrt{\mathrm{se}_1^2 + \mathrm{se}_2^2}, and whether it contains zero is exactly the 5% test. For a paired design it is the interval of the within-unit differences, which is the only one of the three pictures that can see a correlation. For several groups it is a set of intervals for the chosen contrasts, with the multiplicity handled in how they are built rather than left to a reader’s eye.

None of this discards the per-group intervals, which answer their own question — where each group’s mean is — correctly. It adds the one interval that answers the question the reader was going to ask anyway, and it removes the need to do the ruler arithmetic, the ratio of standard errors and the correlation in one’s head. A p-value on its own does not say how large the difference is; an interval for the difference says both how large it is and whether it is distinguishable from zero, which is the pair of numbers the overlap reading was trying to extract from two pictures of something else.

What the formulas establish, and what they do not

Two touching 95% intervals with equal standard errors mark a p-value of 0.56%, and unequal standard errors move it towards, but not to, 5% — 3.19% at a ratio of ten. Intervals at the overlap level touch exactly at p = 5%, and for equal standard errors that level is 83.42%.

Touching 95% intervals always mark a p-value below 5% and touching standard-error bars always one above it, at every ratio of standard errors from one to ten.

Reading non-overlap as significance never rejects more often than the ordinary test, at any true difference, at both ratios drawn.

In a figure of identical groups the chance that some pair of bars separates rises with the number of groups for every kind of bar, and for two groups it equals the touching p-value — the check that ties the range integral behind the many-group table to the two-group closed form, so that the two tables cannot silently disagree. The forest-plot table is the same touching formula with the correlation and the ratio of standard errors both set by the study’s weight, and it inherits that check. None of the intervals in any table is simulated; like the twenty intervals whose misses were counted, these are sampling statements, but here each has a closed form and needs no count.

What does not survive is overlap read as no significant difference. Two estimates 2.4 standard errors of the difference apart, with equal standard errors, have 95% intervals that overlap and a difference with p = 0.0164.

Not claimed: anything for estimates that are not approximately normal, or for intervals that are not symmetric about their estimates — an interval for a proportion near zero or a ratio on a logarithmic scale has arms of unequal length, and the arithmetic of arms adding in a straight line then has to be done arm by arm. The correlation figures assume equal standard errors; with unequal ones the concealed difference is smaller at the same correlation, and has not been tabulated.

Still open: one interval that excludes zero and one that does not

Overlap compares two intervals with each other. The commoner comparison in a results section compares each with zero: the effect in one group has an interval that excludes zero, the effect in the other has one that includes it, and the sentence that follows is that the effect is present in the first group and absent in the second.

That reading is not about overlap at all, and it is wrong more often. It compares two significance verdicts rather than two estimates, and the difference between a significant and a non-significant result need not itself be anywhere near significant. How often two studies of an identical effect split that way, how rarely the split corresponds to a real difference, and what it costs to test the difference properly, are measured in significant in one, not in the other.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Confidence intervalCorrelationError barsp-valuePaired comparisonStandard errorStatistical powerTwo-sample test