Every essay — page 5
Stopping rules
A p-value is defined relative to a sampling plan, so two experiments with identical data and different stopping rules have different p-values. That sounds philosophical and it is arithmetical: testing five times at the nominal level rejects a true null 14% of the time.
Groups that borrow
Eight hospitals are neither one hospital nor eight unrelated problems, and the estimate between those two answers is not a compromise. It is a weighted average whose weight — se²/(se² + τ²) — is decided by how well each group is measured and by nothing else, and the population spread in it is estimated from the data rather than assumed.
Eight groups, one population
Eight hospitals are neither one hospital nor eight unrelated problems. The two obvious answers cost 2.23 and 1.15 in squared error; the estimate between them costs 0.88, and the weight it uses is not a matter of taste.
The weight that decides
B = se²/(se² + τ²) is not a compromise between two answers. It is exactly the posterior mean's weight, it agrees with a numerical integration to ten digits, and an argument that mentions no population at all arrives at almost the same estimator.
The prior the data estimates
A hierarchical model needs a population spread, and it does not ask for one. It reads τ off the distance between the group means — biased six per cent low, exactly zero on 53% of datasets where the groups are identical — and the prior stops being a belief.
When borrowing goes wrong
Partial pooling wins on the total and can lose badly on one group. Placed six population widths out, the group that was never from the population is estimated six times worse than by its own mean — and nothing in the output says so.
The fewest groups that can borrow
At three groups the estimator that shrinks towards its own data's mean returns the group means untouched, on every dataset, because its constant is J − 3. At two it expands instead of shrinking. And the number of groups at which partial pooling starts to be worth doing is five, or two, or never — it depends on how far apart the groups are.
Where the borrowing goes
Pooling cuts the total squared error across eight groups by 56%. Two of the eight take 61% of that reduction, the four best-measured groups share 11% between them, and the largest group gets 1.5% of what the smallest does. The headline is a fact about the groups nobody was asking about.
One population, or two
Group effects from two clusters rather than one bell, with the same total spread. The analysis recovers the same population spread, uses the same weight for every group, and reports nothing unusual — while 46% of its estimates land in a region holding 6.6% of the truths.
Decided before the data
Every other field here repairs an analysis after the fact. These figures are about the arrangement that makes the repair unnecessary: blocking removes exactly the variance the blocks carry, randomisation buys a known reference distribution rather than balance, and varying every factor at once estimates each of them from every run.
The variance removed before the data
Arranging forty units in pairs rather than assigning them at random cuts the variance of the estimated effect to a fifth — and the fifth is knowable in advance, because it is exactly the share of the variance the pairs do not carry.
Randomisation is not balance
A third of all ways to split sixteen units leave the two halves more than half a standard deviation apart on a covariate. What randomisation delivers is not balance but a known reference distribution — and it makes a test exact with no assumption about the data's shape at all.
One factor at a time
Changing one thing per experiment estimates each effect from two conditions; changing everything at once estimates each from every run. The ratio is (k+1)/2 and it is exact — and when two factors interact, the one-at-a-time design recommends a setting it never tried.
How many subjects
Sixty-four per arm for 80% power at half a standard deviation — a power figure that could only be simulated, with nothing to disagree with, until the non-central t was written. Two routes now, agreeing to within the simulation's own error.
The word a fraction costs
A half fraction estimates each main effect as an exact sum of that effect and everything it is confounded with — no error term, no sample-size argument. With every interaction at 0.8 the design reports a true effect of −1 as −0.20, and the design cannot test the assumption that makes the number mean anything.
The design that refuses the corners
Box–Behnken runs three factors in fifteen runs and puts none of them at a corner, which is what makes it usable where a corner cannot be run. It predicts the corner 1.84 times worse than the seventeen-run design that goes there, and 1.31 times worse at the middle of a face, and all three numbers are matrix computations with no simulation in them.
The run that did not happen
Lose one run from any orthogonal design and every coefficient's variance is multiplied by exactly 1 + 1/(N − p), and every pair of coefficients acquires a correlation of exactly 1/(N − p + 1) where there was none. The price is set by the design's spare capacity and by nothing else, and a saturated design cannot survive it at all.
Weighting one sample into another
A weight turns the sample that was assigned into the sample a coin would have assigned, and the exchange is exact: integrated over the population, the standardised difference on every covariate goes to machine zero whatever the assignment rule was. What it charges is observations — a treated arm worth 94% of itself where the assignment is nearly a toss-up and 12% of itself where it is nearly decidable — and past that point trimming does not repair the estimate, it replaces the question. Weighting by an estimate of the probability turns out twice as precise as weighting by the probability itself, and weights fitted to balance the covariates directly reach the precision no estimator can beat — while staying exactly balanced, and silent, on every moment they were not told about.
A score that balances
Weighting each unit by one over its own assignment probability drives the standardised difference between the arms from 0.8310 to 2.8×10⁻¹⁷ — exactly, not nearly. A score fitted without the second covariate leaves that covariate at 0.7057, further apart than doing nothing at all.
How many observations a weight leaves
Kish's effective sample size is exact — for an outcome whose mean does not move with the covariates the weights are built from, the studentised variance reads 1.0680 where the formula says one. For the population's own outcome the same reading is 6.769, rising to 52.497.
The region with no comparison
A trimmed interval covers the average effect over everybody 90.8% of the time at six hundred rows and 41.0% at nine thousand six hundred, while covering the average effect over the units it kept 94.3% and 96.0% throughout. An interval that gets worse as the sample grows is an interval about something else.
Either model, but not neither
The augmented estimator's bias is −0.0085, −0.0083 and −0.0016 wherever one nuisance model is right, against components off by 0.8064 and 0.8190. One step past the overlap sweep it is the least biased estimator on the table at 0.0857 and the worst on it at 1.9265.
The estimated weight is the better one
The propensity is known exactly here, so it can be weighted by — and estimating it from the same data and weighting by that gives a variance ratio of 0.4769 on paired draws. The reason is a projection: the draw's own imbalance explains 56.33% of the true-weight variance and 0.05% of the estimated-weight one.
A weight fitted to balance
Weights fitted so that each arm's weighted covariate means equal the sample's leave a difference of 1.4×10⁻¹⁴ between the arms and give the estimate a third of the variance of weights fitted by likelihood — 0.011883, within a relative 5.8% of the bound no estimator can beat. In the world where the assignment carries a square nobody named, the same exact balance leaves the square further apart than no weighting at all, and where the outcome carries it too the estimate is wrong by 0.6973 with an interval that covers 1.5%.
The moments a balance is told
Weights fitted to balance the covariates' means were wrong by 0.6973 in the world where both the assignment and the outcome carry a square. Told the squares and the product as well, the same construction is off by −0.0045 there and its interval covers 91.0%. The failure moves up a moment rather than away: with a cube in both, the second-moment balance is off by 0.3099 and leaves the cube twice as far apart as no weighting. And where overlap is thin, 37.0% of samples have no such weights at all.
When the observations repeat each other
Every standard error on this site divides by √n, which claims the observations carry independent information. In time order they usually do not: at a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, its 95% interval covers 47%, and two series that wander are called related three times out of four.
The observations that repeat each other
Almost every standard error divides by √n, which claims the observations carry independent information. At a lag-one correlation of 0.8 a fifty-point series is worth about six independent observations, and its 95% interval covers 47%.
Two walks and a finding
Regress one random walk on another, independently generated, and the slope is significant 76.7% of the time with a median R² of 0.17. Nothing connects the two series, nothing in the output says so, and more data makes it worse.