Depth

Series

A field says what an essay is about. A series follows one idea essay by essay — from the question that introduces it to the one that assumes all the others.
Where a D-optimal design puts its runs. The D-optimal measure over 121 candidate settings on a square region. It keeps 9 of them and discards the rest, and the 9 it keeps are the settings a catalogue would have offered without any of this arithmetic. What the search adds is the weights: 0.1458, 0.0962, 0.0802, which nine equal runs cannot express.

Criterion

  1. 1 A design is a number
  2. 2 Four letters and two camps
  3. 3 The design that has to be integers
  4. 4 The family behind the letters
  5. 5 The two terms anybody wanted
  6. +28 more
33 essays · optimality
Where the bootstrap works and where it does not. Uniform data on [0, 1]. For the mean the percentile bootstrap covers 93.5%. For the maximum it covers 0.0%, because a resample can never contain a value larger than the largest one observed, so the interval cannot reach above it.

Bootstrap

  1. 3 Where the bootstrap lies
  2. 4 A distribution drawn from the null
  3. 5 A null with a model in it
  4. 6 Residuals that keep their own variance
  5. 7 Two defects and one resampling
  6. +19 more
24 essays · intervals
Every way of splitting 16 units into two halves. All 12,870 assignments, enumerated. The spread of the standardised imbalance is exactly 2/√n = 0.500, whatever the covariate's own distribution, and 33.3% of assignments differ by more than 0.5 standard deviations. Randomisation does not deliver balance; it delivers a known distribution of imbalance.

Randomisation

  1. 1 Randomisation is not balance
  2. 2 Randomising towards the winner
  3. 3 What the balanced trial is worth
  4. 4 The analysis and the shape
  5. 5 Where the guarantee is exactly zero
  6. +19 more
24 essays · design
What each criterion selects, at 50 observations. 700 series from an AR(2) with coefficients 0.6 and -0.3, every order from 0 to 8 fitted to the same 42 responses so the log-likelihoods are comparable. AIC finds the true order 55.1% of the time and lands above it 25.7%; BIC finds it 54.1% and lands above it 4.0%. The closed form for one extra lag is P(χ²₁ > 2) = 15.73% for AIC, which does not depend on n at all, and P(χ²₁ > ln n) = 4.79% for BIC at this size, which falls to zero. Under the true order is the other failure and it is BIC's: 41.9% against 19.1%.

Order-selection

  1. 1 Choosing the order
  2. 2 The interval after the choice
  3. 3 A criterion is a prediction of the hold-out
  4. 4 A penalty is a trace
  5. 5 The window that has to be chosen, and the term that was dropped
  6. +13 more
18 essays · forecast
Twenty series with a lag-one correlation of 0.8. Every series has a true mean of zero and 60 observations. The marks on the right are the twenty sample means. The variance of that mean is 8.3 times what 60 independent observations would give, so the series is worth about 7 of them.

Dependence

  1. 1 The observations that repeat each other
  2. 2 The check before the standard error
  3. 3 One number for a table of candidates
  4. 3 The model that corrects its error
  5. 4 The cost of differencing a pair
  6. +9 more
14 essays · timeseries
Testing at 0.05 every time the data is looked at. The null is true in every one of these trials and the test is correct every time it is run. Looking once rejects 4.9% of the time, as it should; looking ten times rejects 19.2% of the time. Nothing changed except permission to look.

Stopping

  1. 1 When the looking happens
  2. 2 Spending the error rate
  3. 3 Choosing n after looking
  4. 4 Dropping the losers
  5. 5 Stopping when it is precise enough
  6. +9 more
14 essays · sequential
What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.

Assignment

  1. 1 Balancing what is known in advance
  2. 2 The rule that can be guessed
  3. 3 The analysis has to know the rule
  4. 4 The reference the covariates supply
  5. 5 Three arms and three scores
  6. +7 more
12 essays · covadapt
The familywise error rate with no correction, α = 0.05. Two routes: the curve is 1 − (1 − α)^m and the points are counted over 6,000 simulated families of true nulls. With twenty tests the chance of at least one false positive is 64.1%.

Multiplicity

  1. 1 What the correction corrects
  2. 2 Two different promises
  3. 3 The price of control
  4. 4 One control, many arms
  5. 5 Eight forecasters and one benchmark
  6. +7 more
12 essays · multiplicity
The allocations the rule could have made, from these exact patients. One 200-patient trial allocated by response-adaptive randomisation, re-randomised 999 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the same rule over the same patients in the same order, so what is drawn is the set of experiments that could have happened rather than a sampling distribution. The observed |z| is 1.417, 258 of the 999 re-randomisations reach it, and the p-value is (1 + 258)/(1 + 999) = 0.2590. The curve is the half-normal the ordinary analysis reads the same statistic against; its 5% point is 1.96 and this distribution's is 2.101.

Reference

  1. 1 The experiments that could have happened
  2. 2 The test that needs the rule
  3. 3 The plus one and the round number
  4. 4 The corner the test is calibrated at
  5. 5 How long a block a multiplier shares
  6. +6 more
11 essays · exact
The profile a break point is chosen from. One sample of 120 rows under a break in the persistence, fitted as two first-order regimes at every admissible break point. The maximum is at row 78, where the true break is at 60. The shaded band is every break point within two log-likelihood units of the best one — 8 of the 73 positions searched, which is 11% of the range. The horizontal line is the one-regime fit the search is compared against; the whole profile is above it, at every position, which is the point: a maximum over 73 candidates is above the null by construction and not by evidence.

Break point

  1. 1 A break that was looked for
  2. 2 What a search costs in parameters
  3. 3 Choosing whether to break
  4. 4 Two searches, one sample
  5. 4 A second break on a flat profile
  6. +5 more
10 essays · charged
Every equation's adjustment speed, and the one number they make together. Each series gets its own equation, each is regressed on the same lagged disequilibrium, and what comes back is the whole vector α. Averaged over 400 systems at n = 300: α₁ = -0.154 against -0.15 generated, α₂ = 0.104 against 0.1 generated. The gap closes at the combination of them rather than at any one entry — 25% of any disagreement per step, a half-life of 2.41 steps, where the single equation that fits only the first series reports 4.27.

Adjustment

  1. 1 Which series does the moving
  2. 12 Two failures that cancel
  3. 13 A symmetry that was not enough
  4. 14 A copula that halves a marginal
  5. 15 The zero that survives both
  6. +4 more
9 essays · systems
Every split of 100 units, σ = 1 against 3. Each point is one integer split, with its variance computed exactly rather than simulated. The minimum is at 25:75, which is the ratio of the spreads 25:75, and equal allocation costs 25% more variance — the same as throwing away 20 of the 100 units. The shaded band is every split within 5% of the best, and it runs from 17% to 35%: sharp to state, flat to sit on.

Allocation

  1. 1 Not half and half
  2. 2 The cost of a unit
  3. 3 Allocating on a guess
  4. 4 Balancing towards unequal targets
  5. 5 When the constraints run out
  6. +4 more
9 essays · allocation
The same 40 units, arranged two ways. Both designs estimate the same effect of 0.5 and both are unbiased — 0.488 and 0.497. The blocked design's estimate has standard deviation 0.318 against 0.692, a variance ratio of 0.21 where the model predicts 0.20.

Blocking

  1. 1 The variance removed before the data
  2. 2 Balancing more than one number
  3. 3 Balanced on the wrong function
  4. 4 A threshold in the tail
  5. 5 Which shapes are worth protecting
  6. +4 more
9 essays · design
One forecast, and the band the arithmetic puts round it. An AR(1) with φ = 0.75, 60 observations, fitted by least squares and forecast 14 steps ahead. The point forecast decays towards the fitted mean at φ̂^h; the band is ±1.96 standard errors from σ̂²Σψ̂², which grows with the horizon and stops at the unconditional spread 1.72. The dashed pair is the same band computed at the true parameters, which nobody has. The marks past zero are what actually arrived: 12 of 14 inside the band this once, which is one draw and settles nothing.

Forecast

  1. 1 What the model says next
  2. 2 The interval that forgets it estimated
  3. 3 Which forecast is better
  4. 4 When one model contains the other
  5. 5 What the other forecast adds
  6. +4 more
9 essays · forecast
Where to look depends on the answer. The information a single run at time t carries about the rate of an exponential decay, (∂η/∂θ)² = t²·exp(−2θt), at three values of θ. Each curve has one maximum and it is at t = 1/θ exactly — marked, and found by a search over 8,001 settings that was never told the formula. Nothing in a linear model behaves this way: there the information matrix is X′X and the parameters are not in it, so a design can be chosen once and used whatever the answer turns out to be. Here the design is optimal at a guess, and the three curves are three different experiments for one model.

Local design

  1. 1 The design that needs the answer
  2. 2 The design that hedges
  3. 3 The design for the worst case
  4. 4 Where the minimum is attained
  5. 5 The design that stops guessing
  6. +3 more
8 essays · criteria
The cheap repair needs a number nobody has. The obvious alternative to re-randomising is to simulate the design under its null once and use the critical value that comes out — which is what the arm-dropping design does, where the critical value has to be solved for and is 2.313. It does not transfer here. The rule chases outcomes, so how imbalanced the allocation gets depends on how often anything succeeds, and the critical value moves from 1.668 at a success rate of 0.05 to 2.718 at 0.8. Calibrated at 0.3 and used at 0.8 the test's real size is 12.4%; used at 0.05 it is 0.12%. The randomisation test needs none of this, because it conditions on the outcomes that happened rather than on a rate they were supposed to come from.

Nuisance

  1. 1 What the exactness buys
  2. 2 What the blindfold costs
  3. 3 What a schedule actually buys
  4. 4 Which weights are the inverse variances
  5. 5 Weights that need only a ratio
  6. +3 more
8 essays · exact
Expected width against coverage, n = 30, p = 0.15. The Wald interval is the shortest and covers 94.2%. Clopper–Pearson covers 98.3% and is 13% wider. Shortness is not a virtue on its own — an interval of zero width is the shortest of all.

Width

  1. 3 The shortest interval is the one that misses
  2. 4 Two degrees of freedom, one total
  3. 5 A width promised for a difference
  4. 6 Blinded, and still exact
  5. 7 The bias that lands in the slope
  6. +3 more
8 essays · intervals
What each forecaster says, and what is true. The true probability of the event given a forecaster's signal, and what three forecasters report. The truth is Φ(-0.5 + 1.2w) and the honest forecaster reports it, so its curve and the truth are the same line. The loud forecaster pushes every probability towards the ends and the hedged one pulls every probability towards the middle; both have their mean report held at the base rate of 0.3744, so each crosses the truth exactly once and neither can be caught by checking its average. Their reliability terms are 0.008500 and 0.013025 against the honest forecaster's zero, and all three have the same area under the ROC curve, 0.868312.

Calibration

  1. 1 An identity in three terms
  2. 2 A curve that is a binning
  3. 3 Calibrated and useless
  4. 4 A score that rewards lying
  5. 5 The miscalibration a perfect forecaster shows
  6. +2 more
7 essays · calibrate
The same study read three ways. At time 2 the truth is 0.497. Kaplan–Meier gives 0.532; dropping the censored subjects gives 0.180; treating the censoring time as the event time gives 0.392. Both naive readings understate survival, because the subjects they mishandle are the ones doing well.

Censoring

  1. 1 The data that stops early
  2. 2 The curve that survives censoring
  3. 3 The interval at the end of the curve
  4. 4 A dropout the data cannot see
  5. 5 One minus Kaplan–Meier is not a risk
  6. +2 more
7 essays · survival
The same covariate, three ways round. Three worlds over a treatment, an outcome and a covariate, joined by the same three edges at the same three strengths — 0.90, 0.50 and 0.70 — differing only in which way the two edges touching the covariate point. In the first the covariate causes both and adjusting for it recovers the effect of 0.50 exactly. In the second the treatment causes the covariate, the effect is 1.13, and adjusting returns 0.50 — the direct edge alone, with the part that travels through the covariate deleted. In the third the treatment and the outcome both cause the covariate, the effect is 0.50, and adjusting returns -0.087. The regression that produces those three numbers is one formula, and nothing in the data says which panel it is being run in.

Conditioning

  1. 1 One arithmetic, three decisions
  2. 2 The two worlds that look the same
  3. 3 Adjusting for everything
  4. 4 A collider before the treatment
  5. 5 The sample is a condition
  6. +2 more
7 essays · collider

All essays