Concept

Profile likelihood — where it appears

The likelihood as a function of one parameter, with every other parameter re-estimated at each of its values rather than held fixed. It is how a discrete parameter such as the position of a break is estimated, since nothing can be differentiated with respect to it.

Named by 14 essays across 9 fields — each of them below, with the objects they name alongside it.

The profile a break point is chosen from. One sample of 120 rows under a break in the persistence, fitted as two first-order regimes at every admissible break point. The maximum is at row 78, where the true break is at 60. The shaded band is every break point within two log-likelihood units of the best one — 8 of the 73 positions searched, which is 11% of the range. The horizontal line is the one-regime fit the search is compared against; the whole profile is above it, at every position, which is the point: a maximum over 73 candidates is above the null by construction and not by evidence.

A break that was looked for

A two-regime whitening finds its change point by maximising a profile, and then reads a criterion that counts parameters. Under no break there is no parameter to count, because every position describes the same model.

charged · Break point
Where the general fit becomes the parametric one. The band family's objective at the autoregression's own geometric sequence, cut off at each width, on one sample of 60 rows. The horizontal line is the profile likelihood the parametric fit maximises, written independently through a different whitening. At the full width the two are the same number to 3e-14, which is what says the general construction contains the parametric one rather than resembling it. Below 10 lags there is no line at all: the geometric sequence cut off short is not a covariance matrix, so the objective has nothing to evaluate. Between the two the truncation is briefly above the parametric likelihood — a wrong covariance can fit one sample better than the right one, which is the whole reason a width has to be charged for rather than chosen.

A family before a fit

A regression's coefficients and one correlation can be maximised together. Replace the correlation with an estimated covariance and there is nothing left for "jointly" to mean — until a set of covariances is named, and the set turns out not to contain the truth.

family · Dependence
What each rung is made of. Each pair of searches, over 300 draws, split into the two effects its excess is the difference of. The overlap is what the second search loses by having the first already run at its own answer; the interaction is what the joint search finds by moving the first off it. They subtract to the excess exactly, on every draw, because the pinned supremum cancels. Two disjoint dictionaries of independent columns read an excess of 0.000011 and are made of 0.000514 and 0.000503. A break paired with a dictionary of step columns has an interaction of exactly 0 and is all overlap. And a break paired with an independent column has an overlap of -0.004395 against an interaction of 0.002364, which is what puts its excess below zero.

Two effects in one number

How much two searches over one sample share is measured as the net of two things — ground both of them find, and configurations only the joint search reaches. One extra supremum per draw separates them exactly.

separate · Break point
Most of the rise is the optimiser's, and under one law it is not. The rise in log-likelihood from the tapered plug-in to the maximum over the same eight-lag band, beside what the same optimiser produces on a sample generated from the plug-in's own covariance — where the family is correctly specified by construction and there is nothing to find. Under AR(1) at 0.8 the raw rise is 5.72 and the manufactured baseline is 4.79, leaving 0.93 at 1.8 standard errors; under long memory the excess is 0.14, at 0.2. Under the moving average it is 11.87 at 19.4 standard errors, on every draw. The taper is a shrinkage, and it costs nothing where the sequence decays smoothly and a great deal where it stops dead.

The plug-in and the maximum

A tapered covariance estimate sits five and a half log-likelihood units below the maximum of the likelihood it is substituted into. Four fifths of that is what the optimiser would have found if nothing were missing.

family · Dependence
Generality in the wrong direction buys nothing. Regret on a sample whose persistence changes from 0.95 to 0.65 at row 60, over 200 draws. The three stationary rules — told one number, told a window, told an order — are within 0.4 standard errors of each other, and all three stop in the same place: they are general in the lag direction, and the departure is in the other one. Letting the model change once, at a point estimated from the same residuals, is worth 0.05021 more at 4.5 paired standard errors — about as much again as the whole of the first repair. Being told where the break is adds 0.01926, and being told the entire covariance adds 0.02465.

Where the generality runs out

A covariance that changes half way through a sample is not one a window can estimate. One number, a window and an order are worth the same as each other on it — and letting the model change once, at a point nobody can locate, is worth as much again as all three.

general · Dependence
What a search costs is not a property of that search. The likelihood ratio a searched break in the regression reports, two ways, on every law. On its own — the whole rule being a split of the sample, at no whitening — it averages 34.70 under AR(1) at 0.8, against the 11.07 a chi-square on the five coefficients a split adds would use as a threshold. Inside a rule that also chooses a window from a list of eight, the same search adds only 15.40 — less than half. Most of what a break search finds under correlated errors is the correlation, and a whitening chosen from the same sample has taken it already. A charge measured for one search, carried into a rule that makes two, is not conservative in some harmless direction: it is measuring a different quantity.

A charge that depends on the rule

The break search's charge is 34.7 on its own and 15.4 once a window has been chosen from the same sample. Most of what a break search finds under correlated errors is the correlation, and a whitening has taken it already.

twice · Break point
The split depends on the order. How much two searches share, measured both ways round, over 300 draws. Pinning the first search at its own answer and searching the second gives one overlap; pinning the second and searching the first gives another. A break paired with an independent column reads -0.004395 one way and 0.002163 the other, at 8.70 paired standard errors and on opposite sides of zero. The excess the two components subtract to is the same in both orders by construction, so what changes is only how it is attributed. There is no order-free way to say which of two searches found ground both can reach, and the two orders bracket it.

A split that depends on the order

Run the second search first and pin that instead, and the same draw gives a different overlap and a different interaction — with the same difference. And one pair has no second order at all.

separate · Break point
One likelihood, three answers. The concentrated Gaussian log-likelihood of one sample of 120 rows under AR(1) at 0.8, as a function of the correlation the errors are whitened at. Three rules put three different numbers on this curve. The two-step rule reads the least-squares residuals and lands at 0.7616, giving up 0.304 of log-likelihood. Iterating moves it to 0.8080 and gives up 0.002. The maximum is at 0.8044. The curve is not flat between them: what a fixed point of the residual update finds is a solution of a different equation, and the difference is the Jacobian term ½log(1 − ρ²), which grows as the correlation does.

Iterating is not maximising

Re-reading a correlation from the generalised residuals and refitting converges in seven steps. What it converges to solves the first-order condition of a sum of squares, and the likelihood has one term more than that.

together · Dependence
A wider band is always a better fit. The likelihood maximised over the band, at five widths, averaged over 30 samples. A band at L lags is a band at L + 1 with the last entry held at zero, so the families are nested and the maximised likelihood cannot fall — it does not, on any draw. What it does is rise at 0.984 of log-likelihood a lag. A parameter that is doing nothing buys half a unit in expectation and Akaike's criterion charges one, so this is a criterion very nearly indifferent between every width on offer. The dashed line is what a charge of one unit a lag would exactly cancel. Nothing in the fit chooses a width, and what does choose one is a charge somebody has to pick.

Nothing in the fit picks the width

A wider band is always a better fit, and it is better by about one unit of log-likelihood a lag — which is the order of what a criterion charges for a parameter. Three defensible rules choose widths a factor of three apart.

family · Order-selection
What a second break adds. Over 200 draws, the likelihood ratio a search over one break point reports, and how much more a search over an ordered pair adds on top of it. Under AR(1) at 0.8, which has no break at all, the first search manufactures 5.697 and the second adds 4.278. Under a law with exactly one break — where a second one is as absent as the first was in the row above — the first search reports 9.442 and the second still adds 5.800. Searching for something that is not there costs the same whether or not something else was there to find.

A second break on a flat profile

Searching a hundred and twenty rows for one change point manufactures five units of likelihood. Searching for a second manufactures four more, on a series that has at most one — and on a profile whose whole range is under seven.

charged · Break point
Four intervals at one stopping time. 3,500 experiments under the sequential rule with a first stage of 5 and a required half-width of 0.4, which is a demand that knowing σ would meet with 24.0 observations and which the rule meets with 20.8. The rule's own interval covers 90.3%. Replacing the fixed width by a t interval on the same data gives 92.0%. Keeping the rule's own random sample size and drawing a fresh sample of that size gives 89.8% at the fixed width and 95.5% for a t interval — so the sample size being random costs nothing, and the sample size being chosen by the data the interval is built from costs the rest. The spread estimated at the stopping moment is 17.1% below the truth, which is the same fact one level down.

The interval after a stop it chose

A rule that stops when the estimated precision is good enough stops on the samples whose estimate was small. Its interval covers 90% and claims 95%, and a fresh sample of the same random size covers 95.4%.

guarantee · Stopping
Fitting them together is worth something under one law. Four fits of the same regression under four dependences: least squares, the two-step plug-in every whitened rule in this collection runs, the coefficients and the band maximised together, and a whitening at the law's own covariance that nobody has. Under the moving average — the one law the band family contains — the joint fit beats the two-step by 0.0077 at 3.3 paired standard errors. Under the autoregression, long memory and the break it is a tie: 0.4, 1.0, 0.3 standard errors. That is the same ordering the likelihood gap gave, arrived at through the coefficients rather than through the objective.

What fitting them together buys

Maximising over the coefficients and the covariance together beats the two-step under one of four dependences and ties under the other three. It is the one the band family contains, and the likelihood said so before any coefficient was compared.

family · Dependence
The interval every package reports first does not cover. Counted coverage of two 95% intervals for the 100-block return level of a normal parent, against the length of the record they were fitted from, over 300 records at each length. The level they are about is known in closed form, so this is coverage of a number rather than agreement between two estimates. The delta-method interval covers 80.3% at 25 blocks and reaches only 89.0% at 200; the profile-likelihood interval sits between 94.0% and 94.7% throughout. The gap is not a small-sample effect that lengthening the record removes — it narrows by 8.7 points for an eightfold longer record.

Two intervals for one return level

Two 95% intervals read off the same fits of the same records, against a level known in closed form. The symmetric one covers 80.3% at twenty-five blocks and reaches only 89.0% at two hundred — and 99.24% of its misses are the interval sitting entirely below the truth, which is not the endpoint anybody expects to fail.

extreme · Extremes
A peak where the recorded data have none. The profile log-likelihood of a selection model in cy, the coefficient that lets the chance of being recorded depend on the outcome itself, for one study of 800 rows whose missingness is at random, with residuals normal; every other parameter is maximised at each fixed value. The model assumes the outcome is normal given the covariates. The curve peaks at cy = 0.35, where the fitted slope is 0.839, and the values of cy within the 95% cut run from −0.13 to 0.75; the likelihood-ratio statistic against cy = 0 is 1.47. The study was drawn with cy = 0.00. With the outcome's law left free, every value of cy fits the recorded rows equally well and this curve would be flat: its curvature is the normal assumption.

The assumption that identifies the mechanism

A selection model estimates how strongly an outcome decides whether it is recorded — the quantity two identical datasets showed no statistic can see — and it does so by assuming the outcome is normal. Where that holds and the outcome does decide, it repairs a slope complete cases put at 0.4318 to 0.5795. Where the missingness is at random and the residual is merely skewed, it reports selection that is not there, moves the slope from 0.5971 to 1.0319, and rejects missingness at random in 72.5% of studies.

missing · Missingness

Named alongside it

The objects these essays reach for when they reach for this one.

WhiteningModel selectionStructural breakClosed formDependenceMonte CarloNuisance parameterSelection effectCovariance matrixIdentificationLikelihood ratioSpecification search

All concepts