Depth

Series — page 2

A field says what an essay is about. A series follows one idea essay by essay — from the question that introduces it to the one that assumes all the others.
The coverage is exact and it is not the nominal rate. ⌈(m+1)(1−α)⌉/(m+1) against m, the number of calibration points, at α = 0.05. It is a closed form and needs no data. It never falls below 95.0% and never reaches 1−α+1/(m+1), the two bounds the rank argument gives. It equals 95.0% exactly at 10 of the 182 sizes drawn — the sizes where (m+1)α is a whole number, which are 20 apart — and sits above it everywhere else, worst at 38 points where it is 97.4359%, or 2.4359% of coverage nobody asked for. Below 19 points there is no such order statistic and the interval is the whole line, which is where the curve starts.

Exchangeability

  1. 1 Coverage from exchangeability alone
  2. 2 What the split costs
  3. 3 Marginal is not conditional
  4. 4 The score is the modelling
  5. 5 When the order matters
  6. +2 more
7 essays · conformal
The first stage an instrument needs is set by the violation nobody can see. The error each estimator converges on when the instrument has a direct effect of 0.05 on the outcome — a path the exclusion restriction asserts is zero and no sample can check. The instrument's error is δ/π exactly, so it is the reciprocal of the very quantity that made the method work: 1.0000 at a first stage of 0.05 and 0.0833 at 0.60. Least squares carries the confounding instead, at 0.3440 at a first stage of 0.30. The two cross at π = 0.1389, and the crossing is exactly δ times 2.7778 — the first stage an instrument needs is proportional to the violation it is assumed not to have, and below that line the method being corrected is the better estimator.

Exclusion

  1. 1 The assumption nothing tests
  2. 2 Weak, and back where it started
  3. 3 What the first stage does not know
  4. 4 Whose effect it is
  5. 5 Two instruments that disagree
  6. +2 more
7 essays · instrument
More blocks, and the light tail is called wrong more often. The share of records whose three-way family call is right, over 400 records at each block count, at 100 readings a block. A 95% interval for the shape is formed and the record is recorded as calling Fréchet, Weibull or Gumbel according to whether that interval sits above zero, below it, or straddles it. The two signed parents go from 51.5% and 71.0% at 20 blocks to certainty by 100. The light-tailed one goes the other way — 83.0%, 75.5%, 55.0%, 36.5%, 4.8% — because its estimate sits at about -0.1157 whatever the record length, and a longer record only shrinks the interval onto that number.

Extremes

  1. 1 Three shapes, one limit
  2. 2 The maximum converges slowly
  3. 3 The threshold is a dial
  4. 4 A level with no data in it
  5. 5 Two intervals for one return level
  6. +2 more
7 essays · extreme
Six groups of 10, each fitting its own slope, then borrowing. Each faint line is one group's own least-squares slope through its own centre; each solid line is that slope after pooling towards the population slope of 0.79. Every group has the same 10 observations. The group whose x values span 0.4 has a slope standard error of 1.86 and moves 91% of the way in; the group spanning 2.0 has a standard error of 0.37 and moves 28%.

Levels

  1. 1 The slope that borrows
  2. 2 Pooling a proportion
  3. 3 Borrowing towards a line
  4. 4 Two levels at once
  5. 5 Two groupings that cross
  6. +2 more
7 essays · multilevel
Three mechanisms leave the slope alone; one does not. The bias of the complete-case slope under each of four missingness rules, counted over 4000 studies of 200 rows at 35.0% missing, with the closed form printed beside each count. Missingness that depends on nothing, on the regressor, or on the second covariate leaves the slope exactly where it was — the closed forms are zero to machine precision and the counts are -0.0005, -0.0005 and -0.0011 against standard errors of about 0.0018. Missingness that depends on the outcome moves it by -0.1635, which is 27.3% of the slope being estimated. The same share of rows is lost in every case.

Missingness

  1. 1 Three mechanisms and one dataset
  2. 2 Dropping the incomplete rows
  3. 3 One imputation is not an observation
  4. 4 The variance between imputations
  5. 5 An imputation model the analysis does not contain
  6. +2 more
7 essays · missing
One wrong model, four designs, four slopes. The slope a straight line converges to when the truth is a quadratic, under four covariate distributions, by two routes: the population projection in closed form, and the mean of 2500 fitted slopes at 200 rows apiece. The even spread over [0, 2] gives 1.6000 and the same spread moved to [1, 3] gives 2.6000, while widening it to [0, 4] gives 2.6000 — the same number as the shifted one, because a symmetric design's target is the truth's tangent slope at the design's own mean and does not read the spread at all. An exponential spread with the SAME mean as the first gives 2.6000. So two studies of one world, each fitting the same wrong model, honestly report slopes 1.0000 apart, and neither is making an error.

Misspecification

  1. 1 What a wrong model estimates
  2. 2 The bread and the filling
  3. 3 Robust is not free
  4. 4 Three corrections and a leverage
  5. 5 Right for the wrong reason
  6. +2 more
7 essays · sandwich
Three series and one relation between them. Above, three series generated from Δy = Πy₋₁ + ε with Π of rank 1. Below, the combination y1 −y2. It stays inside a band of 9.5 while the series themselves travel 28.4. The count of combinations that behave this way is the rank of Π, and it is what every method in the field sets out to estimate.

Rank

  1. 1 Three series and a count
  2. 2 Which series goes on the left
  3. 3 Counting what is still wandering
  4. 4 The weight that is a vector
  5. 5 The rank is a decision
  6. +2 more
7 essays · systems
Two independent random walks, 100 steps. Nothing connects these two series: each is generated from its own independent draws. Regressing one on the other gives a slope with t = -10.9, R² = 0.55 and p = 0.0e+0 — a result that would be reported as a finding by any standard output.

Spurious

  1. 1 Two walks and a finding
  2. 2 What differencing costs
  3. 3 The regression that is not spurious
  4. 4 The test with no table
  5. 5 The cliff that is a slope
  6. +2 more
7 essays · timeseries
A weight that balances, and one that unbalances. The standardised difference between the arms on each covariate, integrated over the population rather than counted in a sample. Unweighted, the arms differ by 0.8310 on the first covariate and 0.6015 on the second, which is what makes the raw difference of arm means 2.7102 against a true average effect of 1.0000. Weighting each unit by one over its own assignment probability removes both differences exactly — -2.78e-17 and -5.69e-19, which is machine precision and not a small number — because the weighted density of the treated arm is the population's own whatever the propensity is. Weighting by a score fitted without the second covariate balances the first to 0.0035 and pushes the second out to 0.7057, further apart than doing nothing.

Weighting

  1. 1 A score that balances
  2. 2 How many observations a weight leaves
  3. 3 The region with no comparison
  4. 4 Either model, but not neither
  5. 5 The estimated weight is the better one
  6. +2 more
7 essays · weights
What a 95% credible interval covers, n = 20. Computed by summing over all 21 possible counts rather than by simulating them. Jeffreys' prior covers close to 95% across the range; a confident prior centred in the wrong place covers almost nothing where the truth is far from it.

Credible

  1. 2 What a credible interval covers
  2. 3 Where the two schools agree
  3. 4 The interval that integrates
  4. 5 The shortest interval, and the one that does not move
  5. 6 An interval for something else
  6. +1 more
6 essays · bayes
Two factors that interact — where each design looks. The true response at the four corners is -10, 6, 4, 0. One factor at a time visits three of them, sees that raising either factor alone helps, and recommends raising both — a corner it never ran, and one that is worse than either single change. It picks the best corner 0.0% of the time against the factorial design's 99.4%, on the same number of runs.

Factorial

  1. 1 One factor at a time
  2. 2 The design that cannot see a curve
  3. 3 Three levels, and the ring where the design says the same thing
  4. 4 The word a fraction costs
  5. 5 The design that refuses the corners
  6. +1 more
6 essays · design
Twenty walks up the same hill, σ = 2. Each walk fits a plane to the same four-corner factorial, takes its gradient as a direction, and steps along it until a run comes in below the one before. The true optimum is the cross. 80% of the walks stop before the best point on their own path — not because the direction was wrong, but because one noisy run is enough to stop them, and the direction error costs only 3.9% of the available gain.

Optimum

  1. 1 Walking up the gradient
  2. 2 The optimum is a ratio, and its interval is sometimes the whole line
  3. 3 Where the enumeration stops
  4. 4 The sign the curvature has
  5. 5 When the best setting is outside the region
  6. +1 more
6 essays · surface
Eight groups, τ = 1 against a within-group spread of 3. Each row is a group. The hollow circle is the group's own mean, the filled one is the estimate after pooling, and the small mark is the truth the data was generated from. The group of 3 moves 75% of the way to the population mean of 0.10; the group of 40 moves 18%.

Pooling

  1. 1 Eight groups, one population
  2. 3 When borrowing goes wrong
  3. 4 When the spread estimates to zero
  4. 5 The fewest groups that can borrow
  5. 6 Where the borrowing goes
  6. +1 more
6 essays · hierarchical
The normal density at sigma = 1.00. The bands hold 68.27%, 95.45%, 99.73% of the mass. Those figures are integrals of the curve drawn, not the memorised 68-95-99.7.

Bands

  1. 1 The shape, and where its mass is
  2. 2 Two standard deviations of what
  3. 3 A tenth as wide, and both of them right
  4. 4 Ninety-three observations, and nothing assumed
  5. 5 All of the next ten
5 essays · normal
What a positive test means, sensitivity 90%, specificity 95%. At a prevalence of one in a thousand, 98 of every hundred positives are false. At one in 10, 33 are. The test has not changed.

Baserate

  1. 1 What a positive test is worth
  2. 2 The base rate was always Bayes
  3. 3 The second test that is not a second opinion
  4. 4 The test is a point somebody chose
  5. 5 The prevalence the test has to estimate
5 essays · paradox
Least squares estimates persistence low, by an amount with a formula. 3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.

Bias

  1. 1 Correcting the persistence
  2. 2 The repair that moves the wrong number
  3. 3 Correcting the forecast instead
  4. 4 The correction that leaves the region
  5. 5 What the interval is short by
5 essays · evaluation
Power at an effect of 0.5 standard deviations. The curve is the non-central t on 2n − 2 degrees of freedom with δ = d√(n/2); the dots are 4,000 experiments run at each size. Reaching 80% power needs 64 per arm.

Power

  1. 2 What a p-value does not say
  2. 3 How many subjects
  3. 4 The spread a pilot supplies
  4. 5 The chance a trial succeeds
  5. 6 An outcome cut in two
5 essays · testing
Twenty samples of 40, every one of them genuinely normal. Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

Qq

  1. 2 What normal actually looks like
  2. 3 Twenty residual plots
  3. 4 The band the eye was standing in for
  4. 5 The plot is about the wrong quantity
  5. 6 Residuals are not the errors
5 essays · normal
Two measurements of the same thing, correlated 0.60. Pick the worst 15% on the first measurement and their average rises by 0.78 on the second. Pick the best and theirs falls by 0.48. No treatment was given to anybody.

Rtm

  1. 2 Regression to the mean
  2. 3 Two analyses of one baseline
  3. 4 The measurement that got them enrolled
  4. 5 A lead that a heavy tail keeps
  5. 6 The slope of a density nobody can see
5 essays · paradox
Twenty samples of 40, every one of them genuinely normal. Each panel is a quantile-quantile plot of 40 draws from a normal distribution. The worst point in the worst panel sits 0.87 standard deviations off the line. Anything a reader would reject here would be a false alarm.

Seeds

  1. 1 The seed is part of the figure
  2. 2 A coverage table with its own error
  3. 3 The same draws for both methods
  4. 4 A simulation that stops when it looks settled
  5. 5 The draws aimed at the tail
5 essays · method

All essays