What a summary of a scatter is a property of

R² is a property of the design

One line, one residual spread, five studies that differ only in how far apart they placed their x values. R² runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in every one of them. Nothing about the relationship changed.

Worth reading first: Four datasets, one summary.

Adding a useless predictor raises R-squared, and four datasets can share one. Both of those hold the summary still and move the data.

Here is the reverse. The relationship is held absolutely fixed — the same slope, the same intercept, the same residual standard deviation — and only the design moves.

One relationship at five designs, residual spread 1.00Every panel has the same slope of 1, the same intercept of 0 and the same residual standard deviation of 1.00. Only the range of x differs. R-squared runs from 0.021 to 0.849, and the estimated residual spread is 0.9932 in all five.same line, same noise, different range of xR² 0.02±0.25R² 0.08±0.5R² 0.26±1R² 0.58±2R² 0.85±4the line is y = x in every panelR² from 0.02 to 0.85
Fig. 1 Five studies of y = x with a residual spread of 1. The only difference between the panels is how far apart the x values were placed. R2R^2 runs from 0.021 to 0.849 and the estimated residual spread is 0.993 in all five.

The arithmetic is one line

For a straight-line relationship with slope β and residual standard deviation σ,

R2  =  β2var(x)β2var(x)+σ2.R^2 \;=\; \frac{\beta^2 \operatorname{var}(x)}{\beta^2 \operatorname{var}(x) + \sigma^2}.

Three quantities. Two of them — β and σ — are the relationship: how much y moves per unit of x, and how far the points scatter around the line. The third, var(x), is how the study was run. It says nothing about the relationship at all; it says how wide a range of x the investigator chose to look at.

And it is the only one of the three that is under anyone’s control. Doubling the range of x quadruples var(x) and drives R2R^2 towards 1, with the line and the noise untouched. Across the five panels the range runs from ±0.25 to ±4, var(x) from 0.022 to 5.61, and R2R^2 from 0.021 to 0.849 — a factor of forty.

Which numbers move, and which do not

what is reported how far it travels across the five designs
R2R^2 ×18.83
the slope’s standard error ×16.00
the fitted slope ×1.0009
the estimated residual spread ×1.0000
What changing only the design does to each reported number. Each bar is the largest value divided by the smallest across five designs of the same relationship. R-squared spans a factor of 18.8 and the slope's standard error 16.0; the fitted slope spans 1.0006 and the residual spread 1.0000, both of which are 1 to within the counting.
Fig. 2 The same five designs, with each reported quantity’s largest value divided by its smallest. Two of the four are facts about the study and two are facts about the relationship.

The bottom two rows are the relationship and the top two are the design. The slope is unchanged because it is the thing being estimated; the residual spread is unchanged because it is the other thing being estimated.

The residual spread’s ×1.0000 is exact rather than a measurement that happened to come out flat. The grid of x values at every half-width is one grid scaled, and the projection that produces residuals is invariant to scaling x, so the residuals from the ±0.25 design are the same numbers as the residuals from the ±4 design. The invariance is an identity, and it is worth knowing that because it says the flatness is not an artefact of a large simulation.

What this does to comparing two studies

Two papers report R2R^2 for the same relationship measured in different populations. One reports 0.85 and one reports 0.26. That difference is entirely consistent with the two having measured exactly the same thing with exactly the same precision per observation, and having sampled x over ranges differing by a factor of four.

So R2R^2 is not comparable between studies unless the range of the predictor is comparable, and the range of the predictor is almost never reported in a form that lets a reader check. A paper that prints “R2R^2 = 0.26” without the spread of its x values has printed a number its reader cannot interpret, and the missing information is one standard deviation.

That is a stronger claim than the usual caution about R2R^2. The usual caution is that a high R2R^2 does not mean a good model — which is also true. This one is that a higher R2R^2 does not mean a stronger relationship, even between two correct models of the same thing.

The number that is invariant, and what it is for

If two of the four reported numbers are the study and two are the relationship, the obvious question is why the study’s numbers are the ones printed.

The residual standard deviation σ̂ is the answer to how far is a point from the line, in the units the outcome is measured in. It is what a prediction interval is built from, so it is the number that says whether the model is useful for anything a reader will do with it: a model of blood pressure with a residual spread of 4 mm Hg is useful and one with 25 mm Hg is not, and the two can have the same R2R^2.

It is not a universal summary and it has one limitation that R2R^2 does not. It is in the outcome’s units, so it cannot be compared between studies measuring different things. That is the property R2R^2 was invented to supply, by dividing by the outcome’s variance — and the division is exactly what brings the design in, since the outcome’s variance in the sample is β2var(x)+σ2\beta^2\mathrm{var}(x) + \sigma^2, which contains the design.

So the two summaries trade a real thing against a real thing. R2R^2 is comparable across units and contaminated by the design; the residual spread is uncontaminated and tied to its units. Printing both costs one number and leaves a reader able to tell which is which, and a paper printing only R2R^2 has chosen the contaminated one.

One relationship at five designs, residual spread 3.00. Every panel has the same slope of 1, the same intercept of 0 and the same residual standard deviation of 3.00. Only the range of x differs. R-squared runs from 0.002 to 0.384, and the estimated residual spread is 2.9795 in all five.
Fig. 3 The same five designs with the residual spread tripled. R2R^2 now runs from 0.002 to 0.384 — every panel worse than before — and the estimated residual spread reads 2.979 in all five, which is the number that says what changed.

The same fact in three fields’ own words

Different disciplines meet this and give it different names, which is part of why it is not one well-known result.

Restriction of range, in psychometrics and personnel selection. A test’s validity — its correlation with later performance — is measured on the people who were hired, and they were hired partly on the test. The observed correlation is computed on a truncated range and is much lower than the population’s.

What an eligibility window does to a correlation of 0.6. Observations outside a window of the given half-width are discarded and the correlation refitted. At a window of ±0.25 the reported correlation is 0.107; at ±3 it is 0.595. The line is the closed-form correction and the points are counted over fitted samples.
Fig. 4 A population correlation of 0.6, measured on the observations whose predictor falls inside a window. At a window of ±0.25 standard deviations the reported correlation is 0.107. The line is the standard correction and the points are counted.

The numbers are severe. A population correlation of 0.6, measured only on the observations whose predictor lies within a quarter of a standard deviation of the mean, reports 0.107. Within half a standard deviation, 0.209. Within a full one, 0.375. The relationship is 0.6 throughout.

The correction is a closed form — it needs the ratio of the truncated predictor’s standard deviation to the untruncated one, and nothing else — and the counted values agree with it at every window to within 0.002. So this is a case where the repair is complete and available, and the reason it is not routine is that the untruncated standard deviation is usually unknown.

Attenuation by selection, in genetics and epidemiology: a cohort recruited on a criterion has a compressed distribution of that criterion, and every association with it is attenuated.

Signal-to-noise, in engineering, which is the same quantity written the other way up: R2/(1R2)R^2/(1 - R^2) is β2var(x)/σ2\beta^2\mathrm{var}(x)/\sigma^2, a ratio of signal power to noise power, and everyone in that field knows the signal power is a property of the transmitter rather than of the channel.

What a reader can recover, and what they cannot

Given a paper reporting R2R^2, a reader is not helpless, and it is worth saying exactly how far they can get.

With the outcome’s standard deviation in the sample, which is often printed in a table of descriptive statistics, the residual spread follows: σ^=sy1R2\hat{\sigma} = s_y\sqrt{1 - R^2}. That is the whole recovery, it needs one number that is usually available, and it converts the contaminated summary into the uncontaminated one.

With the predictor’s standard deviation as well, the slope follows too, since β̂ = R · sy/sx in a simple regression. At that point the reader has reconstructed the entire fit from three summary numbers.

Without either, nothing is recoverable. R2R^2 alone is a single number constraining a two-dimensional family — every pair (σ, var(x)) satisfying the expression above is consistent with it — and the family is unbounded in both directions.

So the practical demand is the same one the reversal’s own field makes about reporting counts rather than rates: print the spreads beside the summary. One extra row in a table of descriptive statistics makes every R2R^2 in the paper interpretable, and its absence makes them all uninterpretable. That asymmetry is what makes the omission worth naming rather than tolerating — the cost of fixing it is a row and the cost of not fixing it is the number.

The design choice this licenses, and the one it does not

If R2R^2 rises with the range of x, the obvious move is to widen the range. That is a real design principle and it has a limit.

The licensed version. For a relationship that really is linear over the wider range, spreading the x values raises R2R^2, raises the precision of the slope, and costs nothing. It is why a dose-response study uses widely separated doses and why a calibration uses standards across the whole working range. The extreme is placing half the observations at each end, which maximises var(x) for a fixed range and is the optimal design for estimating a slope.

The unlicensed version. Widening the range beyond where the relationship is linear does not measure the relationship better; it measures a different relationship. The points at the ends carry the most leverage, so a curvature that is invisible in the middle sets the slope, and R2R^2 stays high because R2R^2 is high whenever var(x) is large — including when the model is wrong.

That is the trap worth naming. A wide design produces a high R2R^2 whether or not the model fits, because the term that makes R2R^2 large is in the numerator regardless. A curve fitted with a straight line over a wide range can report an R2R^2 of 0.9 and residuals with an obvious arc in them, and only one of those two is usually looked at.

Four datasets, slope 0.50, R² 0.67. Every one of these fits reports the same slope to two decimals and the same R². Only the first is a linear relationship with noise: the second is a curve, the third is a line with one outlier, and the fourth has its slope set by a single point.
Fig. 5 The quartet’s second dataset is exactly that case at a narrow range. Widen its x values and its R2R^2 rises without its curvature going anywhere.

What it does to a study designed on a pilot’s R-squared

The practical damage is done at the design stage rather than at the reporting one, and it has a specific shape.

A pilot study reports an R2R^2 of 0.26. The main study is sized to detect the relationship, and the sizing is done from the pilot’s R2R^2 — which is the standard route, because R2R^2 is the number the pilot printed. If the main study will sample x over a wider range than the pilot did, the calculation is conservative and the study is over-sized; if narrower, it is under-powered, and by a factor that has nothing to do with anything the investigators considered.

The correct input is the residual spread and the main study’s own planned var(x), which is a quantity the design already fixes. Nothing about the calculation is hard. It is not done because the pilot reported the wrong number and the main study’s designers had no way to recover the right one.

The same error runs through a power calculation based on a published effect in the same way: the published quantity is a standardised one, standardised by a spread that belongs to the study that published it.

With several predictors it gets worse rather than better

Everything above is one predictor. With several, the design contributes through the whole covariance matrix of the predictors rather than through one variance, and there are more ways for it to enter.

The R2R^2 of a multiple regression is the share of the outcome’s variance explained by the fitted combination, so it depends on the spread of each predictor and on how the predictors are correlated with each other. Two studies with the same relationships, the same residual spread and the same marginal spreads can report different R2R^2 because their predictors are correlated differently — which is a fact about who was sampled rather than about any of the relationships.

The direction is not fixed. Predictors correlated with each other and with the outcome in the same direction inflate R2R^2 relative to independent ones; predictors correlated with each other and contributing oppositely can deflate it. So the multiple-regression R2R^2 inherits the single-predictor contamination and adds one whose sign cannot be argued in advance.

And the two summaries that were clean stay clean. Each coefficient is still the thing it estimates, and the residual spread is still the spread around the fitted surface in the outcome’s units. What the extra predictors do change is the coefficients’ standard errors, which rise with the correlation between predictors — so a study with a high R2R^2 can have every individual coefficient poorly determined, which is the ordinary observation about collinearity arriving here as one more thing R2R^2 cannot distinguish.

Adding a predictor that is pure noise raises R2R^2 by 1/(n − 1) in expectation, and that is the same statement one more time: the number rises for a reason having nothing to do with the relationship. Three mechanisms, one summary, and no way to tell from the summary which of the three moved it.

Two routes, and the reading that does not survive

The closed form and a count agree at every design.

The closed form is the expression at the top of this page, evaluated at each design’s var(x). Nothing is sampled.

The count generates four thousand datasets at each design, fits each by ordinary least squares, and averages the fitted R2R^2. It touches none of the algebra: it does not know what var(x) is and computes its R2R^2 from a residual sum of squares.

The two do not agree, and that is the interesting part. The average fitted R2R^2 is above the population value at every design — 0.0454 against 0.0214 at the narrowest, 0.8548 against 0.8486 at the widest — because a fit spends two degrees of freedom chasing whatever noise the sample happened to carry. R-squared is biased upwards by the fitting, by (1R2)/(n2)(1 - R^2)/(n - 2) to first order, which is largest exactly where R2R^2 is smallest.

Adding that bias to the closed form reproduces the counted average to within 0.006 at every design, which is a much stronger statement than agreement to a loose tolerance would have been: it would fail if either route were wrong, and also if the bias had the wrong sign or the wrong size.

The reading that does not survive is R2R^2 taken as a property of the relationship. That reading requires the R2R^2 at the narrowest design to equal the R2R^2 at the widest, and they are 0.021 and 0.849. The positive companion is exact rather than approximate: the estimated residual spread is identical to within 10⁻⁹ across the same designs, so noise that depended on the design would be visible immediately rather than absorbed into a tolerance.

One thing the two routes do not settle, and it is worth naming so the agreement is not over-read. Both compute R2R^2 for a relationship that is genuinely linear over the whole range sampled. Neither says anything about what R2R^2 does when the model is wrong — a curve fitted with a line has an R2R^2 that depends on the design and on how much of the curvature the range happens to cover, and the two contributions are not separable by anything on this page.

Where the effect is largest, and it is not where it is discussed

Restriction of range is usually raised in one setting — selection on the predictor — and the arithmetic does not care how the range came to be narrow.

Observational studies of common exposures have narrow ranges by nature. A study of dietary sodium in one country samples a predictor whose population spread is a fraction of the global one, and reports a correlation attenuated by exactly the factor above. Two countries studying the same biological relationship report different strengths, and the difference is their diets’ variability.

Repeated measurements narrow the range without anyone choosing to. Averaging a predictor over several occasions reduces its measurement error and therefore its observed variance if the error was a large share of it — and the direction is the opposite of attenuation, which is why the two effects are hard to hold together.

A well-controlled experiment narrows the range on purpose. Controlling every nuisance variable tightly is what an experiment is for, and it compresses the outcome’s total variance, so the R2R^2 of the effect of interest rises while every other relationship in the same data has its own R2R^2 suppressed. An experiment is a design with a deliberately weird var(x), and its R2R^2 is the least transportable number it produces.

The common thread is that none of these is an error, in the sense the reversal keeps returning to. Every correlation reported is the correlation in the population sampled. What fails is the step where a reader takes it as a property of the relationship, and that step is invited by the number’s name.

Still open: what a comparable summary would look like

R2R^2 cannot be compared between studies and the residual spread can only be compared between studies measuring y in the same units. Neither is a summary a meta-analysis could pool across fields.

The candidate is the residual spread divided by the outcome’s own standard deviation in a reference population rather than in the study’s sample — which is R2R^2 computed with somebody else’s var(x). That is a well-defined quantity, it is comparable, and it needs a reference population that someone has to nominate.

Whether such a reference is available in practice, what it costs to be wrong about it, and whether the resulting number behaves better than the residual spread in the units of the outcome are all open here. What is not open is the negative half: the number currently printed is a function of a design choice, and no amount of care in reading it recovers what it would have been under a different one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

CorrelationEstimated varianceExperimental designResidual plotRestriction of rangeSummary statisticsVariance explained