What a two-unit study should report
Worth reading first: Eight groups, one population · The prior the data estimates.
The essay before this one established that a variance estimated from two units is uninformative: a scaled chi-square on one degree of freedom, with an interquartile range spanning a factor of thirteen and a ten-to-ninety range spanning a factor of a hundred and seventy-one, coming out exactly zero more than a quarter of the time.
That is a statement about a nuisance parameter. What a two-site study actually prints is an interval for its overall mean, and three things can be done about the variance on the way to one: treat the sites as the population and ignore it, estimate it from the single degree of freedom available, or take it from outside the study.
They give different intervals, and — the part that is easy to miss — two of them are intervals for different quantities.
Neither number is a mistake. They are answers to different questions printed in the same place.
The two questions
A two-site trial can be asked about two things and the words used to report it do not distinguish them.
These two sites. The estimand is the average of the two sites that were actually studied, and the only uncertainty is the within-site sampling error. There is no between-site variance in the question at all, because the sites are not a sample of anything.
Sites in general. The estimand is the mean of the population the two sites were drawn from, and the uncertainty has two parts: the within-site error, and the fact that two sites are a very small sample of sites.
The first is a statement about a study and the second is a statement about the world, and the second is almost always what a reader takes away — a trial run at two hospitals is read as evidence about the treatment, not about those two hospitals. The fixed-effect interval covers that reading 54.77% of the time.
The confusion is easy to make and hard to notice because the point estimate is the same. Both analyses report the average of the two site means, and they differ only in the width around it and in what the width is a statement about. A reader comparing two papers with the same estimate and very different intervals will reach for a difference in sample size or in analysis quality, and the difference is in the question.
The interval that is exact, and why it is
The random-effects interval at two units has an unexpectedly clean property: it is exact, not approximate.
With two units the between-unit standard error is half the gap between the two site means, on one degree of freedom, and the interval is an ordinary t interval with one degree of freedom. It covers the population mean 94.96% of the time, which is its level, and it does so at every combination of within- and between-site variance — the t argument does not care how the two are divided, because both enter the site means the same way.
So the previous essay’s finding about the variance being uninformative does not make the interval invalid. It makes it wide.
The multiplier is 12.706, against 1.96 for a normal. Its tenth and ninetieth percentiles of width are 2.39 and 31.67, a factor of thirteen — the previous essay’s chi-square arriving as a width rather than as a variance.
That is an honest interval and it is nearly useless. It covers what it says it covers, and a study reporting one has told a reader that the data supports almost nothing about sites in general — which is true, and is the finding.
The one interval that over-covers, and why that is not a fourth option
The table has a fifth row worth reading: the exact random-effects interval, scored against the two sites in hand, covers 98.25%.
That is over-coverage and it is what a reader would expect — an interval built to account for between-site variation contains the average of the two studied sites more often than 95% of the time, because part of its width is doing a job the narrower question does not need done.
It matters because it rules out the obvious compromise. A reader tempted to report the exact interval and interpret it as a statement about these two sites gets an interval that is correct in the weak sense of covering at least its level, and is eleven times wider than the interval that question actually needs. Paying for a between-site variance and then not making the claim it buys is the worst available arrangement: the width of the general claim and the content of the specific one.
Where the cost actually sits
The two-unit case reads as pathological and the sweep says it is the end of a curve.
Knowing the variance rather than estimating it is worth a factor of 4.4 at two units, 1.8 at three, 1.4 at four and 1.05 at twenty. By ten units the two intervals are within a fifth of each other and by twenty the question has stopped mattering.
So the whole of what this essay is about lives at two units and three. That is not a small corner: two-site trials, two-arm cluster studies, before-and-after designs with two periods and multi-country studies with two countries are all in it, and they are in it because of a count nothing in their output reports.
A second reading of the same curve is worth having, because it is the one a design meeting can act on. The exact interval’s width at K units is proportional to , and that product falls from 8.98 at two units to 2.48 at three, 1.59 at four and 0.47 at twenty. Almost all of the fall is in the first step: going from two units to three cuts the interval to a third, where going from ten to twenty cuts it by a quarter.
The reason is that the t multiplier and the square root are both improving at once, and the t multiplier improves fastest where the degrees of freedom are fewest — from 12.706 to 4.303 in one step. A study contemplating a second site and a third is not making the same kind of decision twice.
The variance supplied from outside
The third option is to take the between-unit variance from somewhere else — an earlier study, a meta-analysis, a regulator’s assumption — and build a normal interval with it. It covers 95.14% when the value is right, at a width of 2.97, which is between the other two.
What it has to be right about is the question.
The error is one-sided in consequence. Too small breaks the promise badly — 59.74% at a quarter of the truth — and too large costs only width, 99.98% at twice the truth and 100.00% at four times.
That shape has a practical reading: a supplied variance should be supplied generously, and the sensible thing to supply is an upper bound rather than a best estimate. A bound from a previous literature is usually available where a point value is contested, and the loss function says the bound is the right object.
It also means the third option is not a free lunch. It converts a wide but honest interval into a narrow one whose validity depends on a number the study did not measure, and there is no diagnostic inside a two-site study that could check it — one degree of freedom is one degree of freedom whatever the interval is built from.
The three options as one table
Collecting them gives the choice a two-site study actually faces, with the price of each.
A fixed effect for the unit costs nothing and answers the narrow question. Width 1.127, coverage 96.37% of the mean of these two sites, 54.77% of the population mean. It needs no input the study does not have.
The estimated between-unit variance answers the broad question and is exact. Width 13.02, coverage 94.96% of the population mean. It needs no input either, and it says the data supports almost nothing.
A supplied between-unit variance answers the broad question at width 2.97 and coverage 95.14%, and it needs a number from outside whose error is not symmetric in consequence: a quarter of the truth takes the coverage to 59.74%.
There is no fourth column with a short interval, the broad question, and no external input in it, and the reason is not that nobody has built one. Two units is one degree of freedom, and an interval for a population from a sample of two either uses a t on one degree of freedom or uses something the sample did not supply. The table is a complete account of the options because the arithmetic is.
What a defensible report looks like
Add units before adding observations. The curve says the return to a third site is enormous and the return to a twenty-first is nothing, so a study with a fixed budget and a choice between more observations per site and more sites has an answer at the small end that is not close. This is the same arithmetic a design effect makes visible: observations inside a unit repeat each other and units do not.
Say which estimand the interval is for. “The overall mean” is ambiguous between two quantities whose intervals differ by a factor of eleven, and both analyses are routinely described with the same phrase. One sentence — these two sites or sites of this kind — settles it.
Report the number of units, not just the number of observations. A study with two hundred observations at two sites and one with two hundred at twenty sites are the same size by every count an ordinary methods section makes, and their intervals for sites in general differ by a factor of thirteen. That is the same complaint a cluster-robust interval read against the wrong reference makes about rows and clusters, arriving one level up.
Report the gap between the two sites. It is the entire between-site evidence the study contains — with two units, half the gap is the between-unit standard error — so printing it lets a reader construct any of the three intervals above. It also makes the design’s limitation visible in a way a width does not, because a reader can see that one number is carrying the whole claim. That is the same discipline a second number beside a promise asks for everywhere here.
Prefer the wide honest interval to a narrow assumed one, unless the assumption is external and stated. The exact interval at two units is enormous and it is the only one of the three that is right without an input the study did not produce. Where a supplied variance is used it should be named, sourced, and chosen from the generous end.
Do not read a narrow interval from a two-unit study as reassuring. Of the three options, the narrowest by far is the one that answers the question a reader is least likely to be asking, and the second narrowest depends on an unverifiable input. A two-site study with a tight interval about sites in general has either assumed a variance or reported the wrong estimand, and which of the two is not visible in the interval.
Do not let the estimate’s stability stand in for the interval’s. The point estimate of the overall mean is the same under all three analyses and is perfectly well behaved; it is the width that carries the design’s limitation. A study reporting a stable estimate across analyses has demonstrated that its estimate is stable, which is the property that was never in doubt. It is the same confusion a coverage that says nothing is about: one good property is not a summary of a procedure.
And consider whether the question can be narrowed. A study that cannot say anything useful about sites in general may be able to say something useful about these two, and the fixed-effect interval is a perfectly good interval for that — 1.127 wide and covering 96.37% of its own estimand. Narrowing the claim to what the design supports is a better response than widening it to a factor of thirteen or assuming a number.
The count that is doing the work
One reading collects everything above, and it is the same reading two other fields arrive at from their own directions.
The quantity that governs an interval about a population is the number of units drawn from that population, and it is not the number of observations. Two hundred observations at two sites give an interval about sites in general on one degree of freedom; the same two hundred at twenty sites give one on nineteen. The multiplier falls from 12.706 to 2.093 and the square root from to , and between them they account for the factor of thirteen.
That is the same arithmetic a cluster-robust standard error runs into, where the consistency is in the number of clusters and a study with five has five of whatever the limit is in. It is the same arithmetic the design effect makes visible, where observations inside a unit repeat each other and units do not. And it is the reason a variance from two units is uninformative rather than merely imprecise: one degree of freedom is not a small amount of information about a variance, it is the smallest amount that is not zero.
A methods section that reports n and not K has withheld the number the reader needs, and no amount of care in the analysis puts it back.
What is claimed here and what is not
Twenty thousand studies at each setting. Every coverage figure is a count over that many simulated two-site studies, so a reading of 94.96% carries a standard error of about 0.15 points and the difference between 94.96% and 95% is not a finding. What is a finding is 54.77% against 96.37%, which is four hundred standard errors.
Ten observations per site, one variance ratio. The sweeps use a between-site standard deviation of 1 against a within-site one of 1.2 with ten observations each, so the site means carry about equal contributions from the two sources. A design where within-site error dominates makes the fixed-effect interval’s failure less severe and the exact interval’s width less extreme; one where between-site variance dominates makes both worse.
The exact interval’s exactness is a t argument, not a simulation finding. covers at its level whenever the unit means are normal, whatever the two variance components are, and the measurement at 94.96% is a check on the implementation rather than evidence for the property. What the simulation supplies that the theory does not is the width and its spread, which is what this essay is about.
The fixed-effect reading is scored against the realised site effects. Its estimand is the average of the two site effects that were actually drawn, so it is a moving target across replications — which is correct: that is what “these two sites” means, and a fixed-effect analysis is a statement about whichever units were studied.
Normal site effects throughout. The exact interval’s exactness rests on the two site means being normal, and with two of them there is no way to check. A heavier-tailed distribution of site effects would make the t interval’s coverage depart from its level, and the departure would be undetectable by any diagnostic the study could run — which is a limitation of the design rather than of the interval, and is the reason a two-unit study’s assumptions are load-bearing in a way a twenty-unit study’s are not.
And the supplied-variance sweep assumes the supplied value is the only thing wrong. A study taking a variance from elsewhere is also assuming that elsewhere’s units resemble its own, which is an exchangeability assumption with no test available at two units. The sweep prices getting a number wrong; it does not price the assumption that the number is about the same thing.
Still open: what an integrated variance buys
Three responses were named at the previous essay and this one measures two of them properly plus a third that is between the two. The fourth — putting a prior on the between-unit variance and integrating it out — is the one a Bayesian analysis of a two-site study actually does, and it is not measured here.
There is a specific reason to expect it to be interesting rather than a compromise. A prior on the variance is a supplied input like the third option, but it supplies a distribution rather than a value, so its interval should sit between the supplied one and the exact one — and where exactly it sits is entirely a property of the prior’s tail, since with one degree of freedom the data cannot pull the posterior away from it. The uncomfortable possibility is that the resulting interval is narrow because the prior is confident and reads as though the data had spoken.
What the coverage of such an interval is, against both estimands, under a prior that is right and under one that is not, and whether there is a prior whose interval is as short as the supplied one and as honest as the exact one, is the measurement this field has not made.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A ratio that changes between blocks — both name conservative interval, coverage, degrees of freedom, interval width
- One population, or two — both name hierarchical model, partial pooling, random-effects, variance components
- Robust is not free — both name coverage, interval width, monte carlo, t interval
- The condition that cannot be dropped — both name conservative interval, coverage, degrees of freedom, interval width
- The fewest groups that can borrow — both name degrees of freedom, hierarchical model, partial pooling, variance components
- The slope that borrows — both name hierarchical model, partial pooling, random-effects, study design
Named objects
A flat tag is an object no other essay names yet.
Conservative intervalCoverageDegrees of freedomEffective sample sizeEstimandHierarchical modelInterval widthMonte CarloPartial poolingPrior sensitivityRandom-effectsStudy designT intervalVariance components