The collection

Every essay — page 9

Essays 193 to 216 of 436, in the same order.

Balancing on what was recorded first

The adaptive field's rules read outcomes and every one of them broke something. These read baseline covariates, and each holds one imbalance flat while the others grow behind it exactly as a coin's do. What follows is the opposite of the outcome-adaptive story at every step: the unadjusted analysis is too cautious rather than too eager, and the exact test that cost nineteen points of power there costs almost nothing here.

What each rule leaves behind, at 120 patients. Four allocation rules over the same cohorts and the same seeds, each scored on three imbalances: the number of patients in each arm, the worst of the nine factor levels, and the worst of the 24 cells of the cross-classification. No rule holds all three. Permuted blocks hold the totals exactly and leave the margins near a coin's. Blocks inside every cell hold the cells and let the totals drift, because 24 part-filled blocks do not have to end level. Minimisation holds the margins and the totals and is at 83% of a coin's cell imbalance. Each of the three columns is somebody's definition of a balanced trial.

Balancing what is known in advance

Four allocation rules, three definitions of balance, and no rule that holds more than one of them. Minimisation keeps the worst factor margin near three patients whether the trial has forty or six hundred and forty — and lets the imbalance in the cross-classified cells climb to 86% of a coin's, because the cells are not what it is watching.

9 figures · Assignment, part 1
What a rule gives away by being predictable. Minimisation run at every probability from a coin to fully deterministic, over 500 cohorts of 120 at each. The upper line is the share of assignments an investigator who knows the rule and the enrolled patients can name in advance: 49.7% at p = 0.5, which is a coin and cannot be beaten, and 87.6% at p = 1 — short of everything only where the two arms tie and the rule falls back on a coin. The lower line is what that is worth: an investigator who enrols a patient 0.5 of a standard deviation better than average whenever they predict their favoured arm produces a treatment effect of 0.75 where the truth is zero. Nothing about the randomisation was broken; the bias entered through who was enrolled, which is the one thing an allocation rule cannot control. The dashed line is the closed form 2δ(2g − 1).

The rule that can be guessed

A balancing rule improves as it becomes more deterministic, and a deterministic rule can be worked out in advance from information the person enrolling the patient already has. At full determinism 87.6% of assignments are guessable, and an investigator who acts on the guess produces a treatment effect of three quarters of a standard deviation where the truth is zero.

9 figures · Assignment, part 2
Four analyses of the same trials, with no treatment effect at all. 700 trials of 120 patients allocated by minimisation at p = 0.8, with the prognostic factors carrying a real effect on the outcome and no treatment effect — every rejection below is a false one. Two statistics, the plain difference and the same after adjusting for the balanced factors, each read against two reference distributions: a t table, and the set of allocations the rule could have produced from these covariates. The unadjusted comparison rejects 0.6% where it claims 5% — conservative, which is a loss of power rather than an error, and nothing on the output says so. Adjusting puts it back at 5.4%. Both re-randomised versions are at their nominal level by construction, whatever statistic goes into them.

The analysis has to know the rule

A trial balanced by minimisation and analysed by comparing the two arms' means rejects a true null 0.6% of the time where it claims 5%, and at full determinism 0.0%. That is not an error anybody complains about — it is a test that has stopped working, paid for by a balance the analysis then refused to use.

9 figures · Assignment, part 3
The allocations this trial could have made, and the ones it could not. One 120-patient trial allocated by minimisation at p = 1, re-randomised 399 times. No outcome is redrawn anywhere in this figure: each re-randomisation runs the rule again over the same patients in the same order with the same recorded factors, so what is drawn is the set of experiments that could have happened. The bars are that set; the outline is what shuffling the labels gives, which is the reference distribution of a coin and is what every off-the-shelf permutation routine assumes. The coin's is wider — its 5% point is 1.95 against the rule's 1.09 — because a coin's allocations are less balanced and a less balanced allocation gives a larger statistic. Reading this trial against it makes the test conservative rather than anti-conservative, which is the opposite error from the outcome-adaptive case and for the same structural reason.

The reference the covariates supply

Hold the outcomes fixed, re-run the rule that assigned them, count. The same construction cost nineteen points of power in the adaptive field, because its rule chased outcomes and its critical value depended on a rate nobody has. Here the rule reads only what was recorded before anything happened, and the same unadjusted statistic goes from 20.3% power to 55.0% by being read against the right distribution.

9 figures · Assignment, part 4

Comparing two forecasters

The forecast field ranks forecasts by mean squared error and stops. Whether one forecaster is really better is a test, its terms are dependent because neighbouring forecasts overlap, and where one model contains the other it fails completely — declaring the smaller model significantly better while the null it claims to test is true, and more certainly the more data it is given.

One comparison, and the two error bars it can be given. 60 rolling origins, a window of 60 observations, forecasts 4 steps ahead, at the persistence φ = 0.8256 where the two benchmarks have exactly equal population mean squared error. Each mark is one origin's difference in squared error; the horizontal line is their mean, 0.6522. The two vertical bars at the right are ±1.96 standard errors round that mean computed two ways — 0.5337 treating the differences as independent, 0.6880 allowing for the overlap between neighbouring forecasts. The null is true here by construction, so an interval that excludes zero is a mistake, and the narrow one does it far more often than the wide one.

Which forecast is better

Two forecasters, one series, and a difference in mean squared error. Whether that difference is real is a hypothesis test, its terms are not independent, and the standard error it needs is not the one a t-test computes.

9 figures · Forecast, part 3
A test between nested models, under a null that is true. 1000 comparisons: an AR(1) truth, forecast by a fitted AR(1) and by a fitted AR4 whose extra coefficients are zero. In population the two forecasts are identical, so every rejection is false. The larger model's mean squared error is 1.1663 against 1.0583 — worse, by exactly the noise in estimating coefficients that are not there — and the ordinary test therefore declares the smaller model significantly better 67.2% of the time. Read one-sided in the direction anybody asks about, it finds the larger model better 0.0% of the time. Adding the squared difference between the two forecasts back into the loss differential puts the level at 4.9%.

When one model contains the other

The comparison a forecaster most often wants is between a model and the same model with one more term. That is exactly the comparison the standard test cannot make — and it fails by declaring the smaller model significantly better, more confidently the more data it is given.

9 figures · Forecast, part 4
Least squares estimates persistence low, by an amount with a formula. 3000 series of 50 observations at each persistence. The lower curve is the counted bias of the least-squares estimate of φ, and the open marks on it are −(1 + 3φ)/n, computed rather than fitted. The upper curve is the bias left after adding that quantity back, evaluated at the estimate rather than at the truth nobody has: -0.0020 at φ = 0.3, -0.0023 at φ = 0.5, -0.0039 at φ = 0.7, -0.0059 at φ = 0.8, -0.0108 at φ = 0.9, -0.0165 at φ = 0.95. The formula is a leading-order expression and it understates the bias where the persistence is nearest one — -0.0882 counted against -0.0770 predicted at φ = 0.95, which is the corner of the parameter space every one of these approximations is worst in.

Correcting the persistence

Least squares estimates how much a series remembers of itself as smaller than it is, at every value it can take, by an amount with a closed form. Subtracting that amount back is one line of arithmetic, and what the line costs is variance.

8 figures · Bias, part 1
The correction does not arrive at the truth, it passes it. The average decay factor a forecast applies to the last observation, at φ = 0.85 and 50 observations, 3000 series per horizon. The middle curve is φʰ, what the model actually does. Below it is the uncorrected forecast, which uses φ̂ʰ and reverts too fast — 24.8% short at h = 4, 30.0% short at h = 6, 32.7% short at h = 8. Above it is the forecast built on the corrected estimate, which overshoots, and the reason is arithmetic rather than a bad correction: raising an unbiased estimate to a power does not give an unbiased estimate of the power, and the higher the power the more the spread of φ̂ is converted into overshoot.

The repair that moves the wrong number

Correcting the bias in a persistence parameter is one line of arithmetic that works. Feeding the corrected estimate into a forecast repairs the number everybody looks at, makes the forecast worse by squared error at moderate persistence, and improves the interval for a reason that has nothing to do with bias.

8 figures · Bias, part 2
The average decay factor each route produces, φ = 0.85, 6 steps ahead. The truth is φ^6 = 0.3771. no correction averages 0.2616 with a spread of 0.1646 and a squared forecast error of 3.2516; the formula, on the persistence averages 0.4213 with a spread of 0.2528 and a squared forecast error of 3.4827; the bootstrap, on the persistence averages 0.4355 with a spread of 0.2655 and a squared forecast error of 3.5120; the bootstrap, on the decay factor averages 0.3375 with a spread of 0.2278 and a squared forecast error of 3.4132. 800 series, 100 bootstrap refits each.

Correcting the forecast instead

The complaint against the usual repair is that a correction aimed at the persistence lands on the wrong quantity. Aiming it at the decay factor the forecast actually uses fixes exactly that — the error stops compounding with the horizon, 69.7% becomes 9.5% at twelve steps — and the forecast still gets worse.

7 figures · Bias, part 3
Five treatments of an estimate above one, φ = 0.95, n = 25. The correction exceeds one on 31.1% of series at this setting. left where it lands: squared forecast error 12.828, average decay factor 0.7974 against a true 0.7351; capped at 0.995: squared forecast error 5.680, average decay factor 0.5950 against a true 0.7351; capped at 1 − 1/n: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction scaled to fit: squared forecast error 5.535, average decay factor 0.5256 against a true 0.7351; correction refused where it leaves: squared forecast error 5.868, average decay factor 0.4423 against a true 0.7351.

The correction that leaves the region

The bias correction adds (1 + 3φ̂)/n whatever φ̂ is, so it pushes the estimate above one whenever φ̂ exceeds (n − 1)/(n + 3) — on 31.1% of series at φ = 0.95 and twenty-five observations. Five obvious things to do about it differ by a factor of 2.3 in squared forecast error, and none of them is documented as a choice.

6 figures · Bias, part 4
What the forecast interval is short by, φ = 0.85, 6 steps ahead. The plug-in interval covers 88.42% against a claimed 95%. Correcting the variance recovers 0.56 points, propagating the persistence's own standard error recovers 0.40, correcting the persistence recovers 2.66, and all three together recover 4.20 — leaving 2.38 points unaccounted for.

What the interval is short by

The forecast interval covers 88.42% where it claims 95%. Correcting the persistence recovers 2.66 points, correcting the innovation variance 0.56, propagating the persistence's own standard error 0.40 — and all three together recover 4.20 of the 6.58, leaving a residual none of the standard repairs reaches.

6 figures · Bias, part 5

A forecast that is a probability

The field before this one ranks point forecasts by squared error. A probability forecast can be held to something stronger and stranger: say 30% often enough and about three in ten of those days should happen, which is a claim anybody can check by counting. Calibration alone passes a forecaster that issues the base rate every time, the diagram that draws it charges a blameless forecaster in proportion to how finely it was binned, and an honest record of a hundred forecasts shows more apparent miscalibration than the amount routinely read as evidence of a problem. A score that is not proper pays a forecaster to answer only 0 or 1, and what that liar gives up in ranking and resolution is computed rather than described.

What each forecaster says, and what is true. The true probability of the event given a forecaster's signal, and what three forecasters report. The truth is Φ(-0.5 + 1.2w) and the honest forecaster reports it, so its curve and the truth are the same line. The loud forecaster pushes every probability towards the ends and the hedged one pulls every probability towards the middle; both have their mean report held at the base rate of 0.3744, so each crosses the truth exactly once and neither can be caught by checking its average. Their reliability terms are 0.008500 and 0.013025 against the honest forecaster's zero, and all three have the same area under the ROC curve, 0.868312.

An identity in three terms

Reliability minus resolution plus uncertainty is quoted as a rewriting of a probability score. It is an identity to 2.6·10⁻¹⁵ on the one grouping where reliability is the whole score and resolution exactly cancels uncertainty, and it is out by 0.004125 on the coarsest grouping anybody would actually draw.

6 figures · Calibration, part 1
One forecaster, one sample, and the bins it was read in. A reliability diagram of 500 forecasts in 10 equal-count bins, drawn over the curve the same forecaster has in population. The forecaster is honest, so its population curve is the diagonal exactly and every departure the points show is sampling. The sample's reliability term reads 0.002262 and its expected calibration error 0.0358, against a true reliability of 0.000000. Both are properties of this binning as much as of this forecaster: at 5 bins and at 50 the same honest forecaster reports 0.001429 and 0.014100.

A curve that is a binning

A forecaster with no miscalibration in it at all reads 0.001429 at five bins and 0.014100 at fifty, on the same five hundred forecasts. The closed form is K/n times the forecaster's own irreducible score, and subtracting it returns zero.

7 figures · Calibration, part 2
Six forecasters, all calibrated, not equally useful. The resolution of six forecasters that are all perfectly calibrated, each reporting the true probability of the event given a signal that carries more or less of the latent state. The largest reliability anywhere in the family is 2.0e-33, so a calibration check passes every one of them. They are not equally good: resolution runs from exactly 0.00 for the forecaster that issues the base rate every time to 0.092758 for the one that sees everything, and their Brier scores run from 0.234237 — which is the world's own uncertainty, and the score of a table of base rates — to 0.141479. Calibration is a necessary condition that a constant forecast satisfies exactly.

Calibrated and useless

Six forecasters that are calibrated to 2·10⁻³³ run from resolution exactly 0 to 0.092758, and three forecasters with reliabilities from 0 to 0.013025 have areas under the ROC curve identical to every bit a double carries. Each measure is exactly blind to what the other one sees.

6 figures · Calibration, part 3
Where each rule says to put the number. The expected score of reporting each probability on the axis when the event's true probability is 0.25, for four scoring rules. The Brier, logarithmic and spherical scores each bottom out at 0.25 — found by search rather than assumed, to 8 decimal places — which is what makes them proper: a forecaster with a genuine belief cannot improve its expected score by reporting anything else. The absolute-error score is a straight line in the reported value, p + r(1 − 2p), so it has no interior minimum at all; its optimum is 0.0, a distance of 0.250 from the truth, and taking it saves 0.125.

A score that rewards lying

An absolute-error score pays a forecaster exactly ⅛ of a point to replace a true quarter with a zero, and over two hundred records a liar beats a truthful forecaster on 200 of 200. A skill score against the forecaster's own average buys 0.012633 of reported skill for 0.002035 of real score.

6 figures · Calibration, part 4
How long an honest forecaster looks broken for. The calibration error shown by a forecaster with no miscalibration in it at all, at five record lengths, drawn against one over the square root of the length so that the closed form is a straight line through the origin. It is 0.1252 at fifty forecasts and 0.0090 at ten thousand, against a closed form of √(2K/πn) times the mean root bin variance which gives 0.1257 and 0.0089. The threshold drawn across it is 0.02, a figure routinely read as evidence that something is wrong; the mean falls under it at 1976 forecasts and the 95th percentile at about 4111. Below that, an honest forecaster and a miscalibrated one are being told apart by a statistic that is mostly the sample size.

The miscalibration a perfect forecaster shows

A forecaster whose true reliability is exactly zero shows a calibration error of 0.1252 on fifty forecasts and 0.0090 on ten thousand. Every one of 1,200 blameless hundred-forecast records exceeds the 0.02 routinely read as evidence of a problem, and the mean does not fall under it until 1,976 forecasts.

7 figures · Calibration, part 5
One curve, and two forecasters that are each a single point. The ROC curve of an honest forecaster whose signal has correlation 1 with the latent state — every threshold on its probability, from the bivariate normal — and two forecasters that only ever say 0 or 1. The one an absolute-error score pays for says 1 wherever the honest probability exceeds a half: it has a true-positive rate of 0.6742 and a false-positive rate of 0.1375, and its "curve" is the two straight segments through that point, with area (TPR + TNR)/2 = 0.7684. The honest curve's area is 0.8683. Thresholding at the base rate of 0.3744 instead of at a half gives the largest area any two-valued forecaster can have here, 0.7818, and it is still 0.0866 short.

The liar with two answers

The forecaster an absolute-error score pays for says only 0 or 1, and on the ROC square it is a single point: its area is (TPR + TNR)/2 = 0.7684, against the honest forecaster's 0.8683, and it falls below the honest one on 200 of 200 counted records. No relabelling of its two answers returns what it threw away — the best recovers a Brier score short of the honest one by exactly the 0.022154 of resolution lost — and below a signal correlation of 0.7332 the same score prefers saying no every time to an honest forecast.

6 figures · Calibration, part 6
The honest curve, and the same forecaster in three coarse vocabularies. The ROC curve of the honest probability at full signal, area 0.8683, beside the same forecaster rounded to the nearest whole number, area 0.7684 — the two-valued liar — to the nearest half, area 0.8205, and to the nearest tenth, area 0.8650. A vocabulary of v values is v points on the square joined by straight segments, and tied reports count half.

A forecaster that rounds

An honest probability issued in tenths loses 0.0033 of ROC area and 0.000708 of resolution — the variance its bands average away, and 89.5% of the 0.000792 it adds to the Brier score. Two hundred records of two thousand forecasts show that loss on 189; it takes about 3,300 forecasts to put it two standard errors from zero. And 3.207 in every thousand forecasts in tenths are a 0% on an event that happened, which a logarithmic score charges without limit.

6 figures · Calibration, part 7

A design that assumes less

A design for a non-linear model is optimal only at a guess about the answer. Two ways out: protect the worst parameter value in a range rather than the average, which is an optimum that sits on a tie rather than a slope; or stop guessing, run part of the experiment, and design the rest at the estimate — where the interesting question turns out to be what that does to the interval afterwards.

More arms than two

Minimisation balances a trial by making the arms' counts even inside every factor level. With two arms there is one way to measure how uneven two counts are. With three there are several, they are all called minimisation, and they send different patients to different arms — while the ratio a trial was designed to deliver quietly disappears unless the score was told about it.

FieldsThreadsSeriesConceptsFigure librarySearch