Regression, and what the summary hides

The other half a trimmed fit rejected

The exact trimmed fit visits every half of the data, so it can report more than its winner: the best half whose line slopes the other way. Whenever the winner is reversed by a far cluster, that other half is the clean line. Printed beside the winner when its residual sum is within 1.70 times — a rule that raises five per cent of clean datasets — it flags 65% of the reversed fits at seven far rows of twenty, 55% at eight and 30% at nine, and no rule built on fit can do better at nine, where the far cluster has become the tighter half by every measure the data hold.

Worth reading first: A robust loss and a far x.

The start an efficient robust line inherits located the MM-estimator’s failure precisely. On twenty rows with a tight cluster of bad rows far out at x = 9, the exact least trimmed squares fit chooses the cluster’s half of the data on 54 datasets of two hundred at seven far rows, 111 at eight and 159 at nine; the MM-estimator started from it repairs none of those choices and commits to every one. Its breakdown point of one half is real, and it belongs to a start that, for a tight far cluster, gives way well before half the rows are bad.

The response that essay proposed is not a better estimator but a check on the start. The trimmed fit is computed by visiting every subset of eleven rows, so the enumeration that finds the winner has also looked at every alternative. If the best alternative tells a different story at nearly the same price, the data hold two competing lines, and the honest report is both of them.

The alternative worth reporting

The runner-up subset is not the alternative worth reporting. It differs from the winner by one row and draws nearly the same line, so printing it says only that the winner is not unique to the last row. The alternative that matters is the best subset whose line has the opposite sign of slope — the best case the data can make for the other story. Finding it costs nothing extra: the same walk over all 167,960 subsets of eleven rows keeps two running bests, one for each sign, and the exact pruning that makes the walk affordable still applies, since adding a row to a least-squares fit can never lower its residual sum.

Two halves of one dataset with 8 far rows of twenty: the trimmed fit's winner and the best half with the other sign. Twenty rows, 8 of them near x = 9 on a line with a reversed slope. The exact trimmed fit keeps the eleven rows whose own line has the smallest residual sum: 5 far rows and 6 clean ones, slope -0.629, residual sum 0.587. The best eleven rows whose line slopes the other way are all clean, slope 0.398, residual sum 0.657 — 1.12 times the winner's.
Fig. 1 One dataset with eight far rows. The trimmed fit’s winner threads five far rows and six clean ones, with slope −0.63 and residual sum 0.587; the best eleven rows whose line slopes upwards are all clean, with slope 0.40 and residual sum 0.657. The second is the truth, and costs 1.12 times as much.

The figure is one dataset from the two hundred at eight far rows, and it shows the thing in its plainest form. The winning half is a line through five of the far rows and six of the clean ones near x = 1, with slope −0.629 and residual sum 0.587. The best half with an upward slope is eleven clean rows, with slope 0.398 — the true slope is 0.4 — and residual sum 0.657, which is 1.12 times the winner’s. The trimmed objective prefers the wrong line by twelve per cent. Nothing in the winner says that a line nearly as good, through entirely different rows, slopes the other way.

Across all the datasets on which the winner is reversed, the opposite-sign half is the clean line in the way this one is. At seven, eight and nine far rows its median slope is 0.381, 0.393 and 0.396, and it contains no far row at all; the reversed winners it competes with hold a median of six, seven and seven far rows among their eleven. Whenever the trimmed fit picks the cluster, the clean line is sitting in its enumeration as the best case for the other sign.

How close the other half comes

The quantity a report would print is the ratio of the other half’s residual sum to the winner’s. A ratio near one says the two lines fit their halves equally well; a large ratio says the other story is expensive.

How close the best half with the other sign of slope comes to the trimmed fit's winner, by the count of far rows. Median ratio of the best opposite-sign half's residual sum to the winning half's, over two hundred datasets at each count. Where the winner held: 16.32 at 1, 10.07 at 2, 8.16 at 3, 5.20 at 4, 3.56 at 5, 2.74 at 6, 1.96 at 7, 1.51 at 8, 1.45 at 9. Where it was reversed, at the counts with at least four such datasets: 1.17 at 4, 1.42 at 5, 1.60 at 6, 1.35 at 7, 1.63 at 8, 2.15 at 9. The flag at 1.70 is set so that five per cent of clean datasets are flagged.
Fig. 2 The median ratio of the best opposite-sign half’s residual sum to the winner’s, against the count of far rows, split by whether the winner held or was reversed. The dashed line is the flag’s threshold, set so that five per cent of clean datasets cross it.

Where the winner held, the other half is the far cluster’s line, and it is expensive at first and cheap later. With one far row its median ratio is 16.32: one row cannot make a convincing line. With five it is 3.56, with seven 1.96, and by eight 1.51 — the cluster is close to matching the clean line’s fit, and the winner held by a margin the next draw could reverse.

Where the winner was reversed, the other half is the clean line, and its ratio sits between about 1.2 and 1.6 for most counts — 1.35 at seven far rows and 1.63 at eight — and then rises to 2.15 at nine. That last number is the important one. At nine far rows the reversed winner does not merely edge out the clean line; it beats it by more than a factor of two.

The two curves cross between seven and eight far rows, and the crossing is the useful reading of the figure. To the left of it, a reversal is a near thing — the clean line is the runner-up by a margin small enough to print — while a winner that held has beaten the cluster comfortably. To the right, the roles swap: the held winners are the near things and the reversals are decisive. The ratio is therefore most informative exactly where the trimmed fit is most at risk of breaking for the first time, and least informative once it has broken.

A flag, calibrated on clean data

A report could print both lines whenever the ratio is below some threshold. The threshold has to be set where it costs something known, and the natural cost is how often it fires on data with no bad rows at all. On the two hundred clean datasets, 114 have no subset of eleven rows with a downward slope — on a grid from 0.2 to 3.4 with a slope of 0.4, no eleven rows can be chosen to point the other way — so their ratio is infinite and they are never flagged. Among the rest, a threshold of 1.70 raises five per cent of all clean datasets, which includes all three on which noise alone reversed the winner.

What a report that prints both halves when they are close would flag, by the count of far rows. Of two hundred datasets at each count, the datasets on which the exact trimmed fit — and so the MM-estimator carried on from it — is reversed, and how many of them the flag at a ratio of 1.70 catches; and the datasets on which it held, and how many of those the flag also raises. 3 far rows: 1 of 2 reversed flagged, 11 of 198 held flagged; 4 far rows: 4 of 4 reversed flagged, 11 of 196 held flagged; 5 far rows: 6 of 8 reversed flagged, 26 of 192 held flagged; 6 far rows: 10 of 18 reversed flagged, 47 of 182 held flagged; 7 far rows: 35 of 54 reversed flagged, 62 of 146 held flagged; 8 far rows: 61 of 111 reversed flagged, 54 of 89 held flagged; 9 far rows: 48 of 159 reversed flagged, 32 of 41 held flagged.
Fig. 3 At each count of far rows, the datasets on which the trimmed fit — and so the MM-estimator carried on from it — is reversed, with the part the flag catches, and the datasets on which it held, with the part the flag also raises. The flag catches most reversals at seven far rows and a minority at nine.

At that threshold the flag catches 35 of the 54 reversed fits at seven far rows, 61 of the 111 at eight and 48 of the 159 at nine. Those are the datasets on which the efficient step went on to entrench the wrong line, so a report that printed both halves would have shown the clean line beside the entrenched one on about two of every three such datasets at seven far rows, a little over half at eight and under a third at nine.

It also fires on fits that held: 62 of 146 at seven far rows, 54 of 89 at eight and 32 of 41 at nine. On those the printed alternative is the cluster’s line, and a reader seeing two lines would have to decide between them without being told which is right. That is the correct outcome rather than a false alarm. Once a cluster of seven rows of twenty has made its own line fit nearly as well as the clean one, the data genuinely hold two lines, and a report that printed one would be making a choice the data do not support.

What the flag trades

The threshold of 1.70 was chosen by fixing the clean-data rate. Moving it trades that rate against the share of reversals caught.

What the flag on the two halves trades: clean datasets raised against reversed fits caught. Each curve moves the threshold on the ratio from 1 to 4. At the threshold that raises 5% of clean datasets, 1.70, it catches 64.8% of the reversed fits at seven far rows, 55.0% at eight and 30.2% at nine. To catch 80% at nine it has to raise about 18% of clean datasets.
Fig. 4 The share of reversed fits flagged against the share of clean datasets flagged, as the threshold moves from 1 to 4, at seven, eight and nine far rows. The curve for nine stays well below the others: catching most of its reversals means raising a large share of clean datasets.

At five per cent of clean datasets the flag catches 64.8% of reversals at seven far rows, 55.0% at eight and 30.2% at nine. Catching 80% of the reversals at nine far rows needs a threshold that raises about 18% of clean datasets — a flag that fires on nearly one clean study in five to find the cluster that has won. The seven-row curve rises steeply and reaches every reversal by a fifth of clean datasets; the nine-row curve rises slowly throughout, because the reversals there are mostly decisive.

Why no rule built on fit can tell at nine

The flag fails at nine far rows for a reason that applies to every diagnostic that measures how well a half fits. At nine of twenty, the reversed winner is typically seven of the far rows and four clean rows near the bottom of the range: a tight group scattered by 0.35 around a single point, and a few rows the line passes close to at the other end. The clean half is eleven rows scattered by 0.35 around a line over a range of 3.2 in x. The cluster’s half is the tighter half, and its residual sum is smaller by a factor of two in the median.

Any criterion that ranks halves by how well they fit — a residual sum, a scale, a likelihood, a count of rows within some band — ranks the cluster first. The other tempting measure, how many effective points each half’s slope rests on, ranks it first too, because a half built from two tight groups at opposite ends of x is the best-balanced design for a slope there is, which is the reason two points can hide each other from every diagnostic that looks for agreement. The information that separates the two halves is not in the rows’ fit. It is in knowledge outside the data: that the study’s settings were meant to lie between 0.2 and 3.4, that nine rows near x = 9 are an instrument fault rather than a regime, that a reversed slope contradicts the mechanism. A single summary cannot say which half of the data is right, and at nine far rows neither can a pair of them.

What a report can do at nine is print both lines anyway, and the cheapest version of that is to print the other half’s line and ratio unconditionally. At nine far rows the median ratio is 2.15 where the winner is reversed, which reads as “the alternative costs twice as much” — a statement that is true and that, on its own, would lead a reader to the wrong line.

A larger majority breaks sooner

One more lever exists inside the trimmed fit itself: how many rows it keeps. Eleven of twenty is the largest breakdown point a line can have; keeping more rows asks the clean data to be a larger majority, and it is tempting to think that a fit keeping thirteen or fifteen rows could no longer be captured by a cluster of nine, since the cluster cannot supply that many rows on its own.

Asking the trimmed fit for a larger majority, against the count of far rows. Datasets of two hundred on which the exact trimmed fit is reversed, keeping 11, 13 or 15 of twenty rows. 11 rows: 3, 1, 2, 2, 4, 8, 18, 54, 111, 159. 13 rows: 1, 0, 1, 0, 1, 7, 13, 56, 200, 200. 15 rows: 1, 0, 0, 0, 0, 4, 200, 200, 200, 200 — at 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 far rows.
Fig. 5 Datasets of two hundred on which the exact trimmed fit is reversed, keeping eleven, thirteen or fifteen of twenty rows, against the count of far rows. Keeping more rows protects nothing at seven far rows and breaks completely at eight or at six.

It is the reverse. Keeping thirteen rows, the fit is reversed on 56 datasets at seven far rows, against 54 at eleven, and on all 200 at eight. Keeping fifteen, it is reversed on all 200 from six far rows. A fit that must keep fifteen rows has to take some bad rows whenever six or more are bad, and once it must take them the cluster’s tight line is cheaper than a clean line dragged by the bad rows it is forced to keep. A larger coverage is a lower breakdown point, exactly as the arithmetic says, and in this design it buys nothing before it breaks — only fewer noise reversals with no bad rows at all, 1 instead of 3.

Carrying both halves through the efficient step

The flag hands the choice to a reader. A procedure can also make it, and the obvious way is to stop judging the two halves on eleven rows. Run the MM step from each — from the winner and from the best half with the other sign — and keep whichever converged line has the smaller M-scale over all twenty rows. That is the criterion an S-estimator minimises, the start the earlier essay noted most MM software actually uses, applied here to choose between two candidates the exact enumeration supplied.

Carrying both halves through the efficient step and keeping the one with the smaller scale, against the count of far rows. Datasets of two hundred on which the reported slope is reversed. The MM-estimator from the trimmed fit's winner: 0, 0, 0, 0, 3, 8, 18, 54, 111, 159. The rule that also runs it from the best opposite-sign half and keeps whichever fit has the smaller M-scale over all twenty rows: 0, 0, 0, 0, 0, 5, 5, 22, 93, 176 — at 0, 1, 2, 3, 4, 5, 6, 7, 8, 9 far rows. It rescues 3 at 4, 3 at 5, 13 at 6, 34 at 7, 28 at 8, 3 at 9 and spoils 2 at 7, 10 at 8, 20 at 9.
Fig. 6 Datasets of two hundred with the reported slope reversed, against the count of far rows: the MM-estimator from the trimmed fit’s winner, and the rule that also runs it from the best opposite-sign half and keeps the fit with the smaller scale over all twenty rows. The rule rescues most reversals at six and seven far rows and adds reversals at nine.

It helps exactly where the flag already worked, and it hurts where the flag was silent. With the winner alone, the efficient step is reversed on 18 datasets at six far rows, 54 at seven, 111 at eight and 159 at nine. Keeping the smaller scale, it is reversed on 5, 22, 93 and 176. At six far rows the rule rescues 13 reversals and spoils none; at seven it rescues 34 and spoils 2; at eight it rescues 28 and spoils 10. At nine it rescues 3 and spoils 20, and ends with more reversals than it started with. On clean data and up to four far rows it changes nothing.

The reason is the one that defeated the flag. A scale over twenty rows asks how tightly the line fits the rows it fits, after discounting the rest, and at seven far rows the clean line fits thirteen rows well while the cluster’s line fits seven far rows and a handful of clean ones: the clean line wins the larger count. At nine, the cluster is nine rows of tight fit against eleven clean rows of looser fit, and a criterion that rewards tightness picks the cluster, now even on some datasets where the trimmed fit had held. A different objective moves the point where the data stop being able to choose, from a little past six far rows to a little past seven. It does not remove that point, because no objective computed from the rows’ fit can.

That is also why the choice of start matters less than the robust loss’s own behaviour once the cluster is large. Starting from both halves is cheap — the enumeration has already found them — and on these designs it is the best use of the second half short of printing it. But it is a second objective chosen after seeing that the first can fail, and a reader should be told it was used, just as a search’s seed is part of the answer when the search is how the answer was found.

What a trimmed fit’s report should add

The best half with the other sign, and its ratio. It costs nothing beyond the enumeration that already ran, and on these datasets it is the clean line whenever the winner is reversed. A ratio under about 1.7 means two lines fit eleven rows nearly equally well, and on clean data of this design that happens one time in twenty.

The rows each half uses. The winner in the first figure uses five far rows and six clean ones; the alternative uses eleven clean ones. A reader who can see which rows support which line can bring the knowledge the fit does not have — and the line that one point drew is the reminder that this is always where the decision was made.

That a decisive winner is not a vindication. The flag’s silence at nine far rows is the cluster winning, not the clean line. A large ratio says the alternative fits badly; it does not say the winner is right.

Measured on what

The design of the two essays before it: twenty rows, eighteen to twenty clean ones on an even grid of x from 0.2 to 3.4 with slope 0.4 and noise 0.35, and up to nine bad rows near (9, −4.4) — two hundred datasets at each count, on the same seeds, so the reversal counts here are the counts there. Both halves are exact over every subset at every coverage. The flag is calibrated on this design’s clean datasets, and its threshold belongs to this design; a different grid, noise or count of rows would need its own calibration, and the pattern — most reversals caught when the cluster is just winning, few when it wins decisively — is what is expected to carry over. The vertical contamination of the earlier essays, rows eight below the line at an ordinary x, is not measured here, because the trimmed fit is almost never reversed by it.

Still open: a half the data were never meant to contain

The analysis above ends where the data’s information ends, and the next step has to bring in information of another kind. The far cluster is far in x, and x is usually a setting the study chose. A trimmed fit could be told the range the design intended — or, without being told, could penalise a half whose slope depends on rows outside the range most rows occupy, which is a leverage statement rather than a fit statement.

Whether such a penalty can be calibrated so that it costs little on clean data, and whether it separates the two halves at nine far rows where fit cannot, is a measurement these enumerations can make with one extra term in the objective. It is also where the subject borders the cost of a robust standard error: a penalty on leverage changes which estimate is reported, and a reader would need to know the penalty was there to read the estimate at all.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Breakdown pointExact enumerationFalse positive rateLeast squaresLeast trimmed squaresLeverageLocal minimumM estimatorMaskingModel diagnosticsOutliersRobust regression