Series

Levels — the series

7 essays on one idea, from the one that introduces it to the one that assumes the rest.
  1. Six groups of 10, each fitting its own slope, then borrowing. Each faint line is one group's own least-squares slope through its own centre; each solid line is that slope after pooling towards the population slope of 0.79. Every group has the same 10 observations. The group whose x values span 0.4 has a slope standard error of 1.86 and moves 91% of the way in; the group spanning 2.0 has a standard error of 0.37 and moves 28%.

    The slope that borrows

    Pooling a mean makes it look as though how much a group borrows depends on how much data it has. Pool a slope instead and the illusion breaks — ten groups with ten observations each can borrow anything from 28% to 91%, decided entirely by where those ten observations were placed.

    part 1 · multilevel
  2. Ten groups of 10, pooled on the log-odds scale. Each row is a group. The hollow circle is its own proportion, the filled one is the estimate after pooling, and the vertical rule is the pooled population proportion of 31.3%. One group saw no events at all, and its raw proportion of zero becomes 21.2% — an estimate the group's own data cannot produce and the population's can. The arrows are not the same length, and none of the groups differs in size.

    Pooling a proportion

    A proportion cannot be shrunk on its own scale — an estimate would leave the interval, and how much information a count carries depends on where it sits. Move to log-odds and the approximation works, at the price of a group that saw nothing having no estimate at all until the correction supplies one.

    part 2 · multilevel
  3. 16 groups shrunk towards a fitted line, at γ = 1.2. Hollow circles are the groups' own values, filled ones the estimates after pooling, and the diagonal is the line fitted through them with each group weighted by how well it is measured — slope 1.21, intercept 0.20. The horizontal rule is where the same 16 groups would have been shrunk to with no covariate. The spread left to borrow against is 0.64 with the covariate against 1.28 without, so every group is pulled further in than it would otherwise have been.

    Borrowing towards a line

    A group shrunk towards the average of all groups is being compared with groups it has nothing in common with. Fit a group-level predictor and it is shrunk towards what the predictor says a group like it should be — which halves the spread left to borrow against and takes a quarter off the squared error.

    part 3 · multilevel
  4. What 20 clusters of 20 correlated observations do to a 95% interval. Each study has 400 observations arranged as 20 clusters of 20. The lower points are the counted coverage of the usual interval, which treats them as 400 independent observations; the curve through them is 2Φ(1.96/√deff) − 1 with deff = 1 + 19ρ, computed before any data was drawn. At ρ = 0.81 the interval covers 36% rather than 95%. The upper points treat the cluster as the unit and need no variance components at all.

    Two levels at once

    A third level of grouping adds no new arithmetic and produces one number — the design effect — that decides how many independent observations a clustered study is worth. It is the same quantity the time-series field computes for autocorrelated data, arrived at from a completely different picture.

    part 4 · multilevel
  5. What 240 observations are worth, by which question is asked. 8 rows and 10 columns with 3 observations in each cell — 240 in all, each belonging to one row and one column, neither nested in the other. the overall mean: variance 0.1703 against a naive 0.0060, a design effect of 28.4 and 8.5 effective observations; a difference between two rows: variance 2.0682 against a naive 0.0960, a design effect of 21.5 and 11.1 effective observations; a difference between two columns: variance 1.0845 against a naive 0.1200, a design effect of 9.0 and 26.6 effective observations.

    Two groupings that cross

    Pupils belong to a school and to a neighbourhood, and neither is nested in the other. There is then no design effect: the overall mean is worth 8.5 independent observations out of 240, a row difference 11.1 and a column difference 26.6, and which grouping matters depends on the question rather than on the study.

    part 5 · multilevel
  6. What a variance estimated from K units is worth. The between-unit mean square is a scaled chi-square on K − 1 degrees of freedom, so the estimator's whole distribution is decided by the number of units. At two units its interquartile range spans a factor of 13.03 and its ten-to-ninety range a factor of 171.3, and it comes out exactly zero on 26.7% of studies. The closed form and 3,000 simulated studies agree to 0.051 at every quantile.

    A level with two units

    A variance estimated from two units is a scaled chi-square on one degree of freedom. Its interquartile range spans a factor of thirteen, its ten-to-ninety range a factor of a hundred and seventy-one, and it comes out exactly zero on 26.7% of studies — so the design effect it decides runs from 1.00 to 7.01 against a truth of 4.69.

    part 6 · multilevel
  7. Each interval covers one question and not the other. Coverage of each interval for the overall mean, scored against both estimands, over 20,000 two-site studies of 10 observations apiece. The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and over-covers the two sites in hand at 98.25%. Both are correct; they are answers to different questions printed in the same place.

    What a two-unit study should report

    The fixed-effect interval covers the mean of the two sites in hand 96.37% of the time and the population mean 54.77%. The random-effects interval covers the population mean 94.96% — exactly its level, from one degree of freedom — and is 11.6 times wider.

    part 7 · multilevel

All series