QUICK REVIEW
[Paper Review] Discussion of "Breakdown and groups" by P. L. Davies and U. Gather
Marc G. Genton, André Lucas|arXiv (Cornell University)|Aug 25, 2005
History and advancements in chemistry2 references3 citations
TL;DR
This paper proposes a new definition of breakdown point for estimators in time series and dependent data settings, focusing on information loss rather than parameter space limits. By measuring the Lebesgue measure of intersection between badness sets under contamination and clean data, it formalizes when an estimator ceases to convey useful information, yielding asymptotic breakdown points of 0% for OLS, 22.1% for LMS, and 50% for DR estimators in AR(1) models.
ABSTRACT
Discussion of ``Breakdown and groups'' by P. L. Davies and U. Gather [math.ST/0508497]
Motivation & Objective
- To address limitations in traditional breakdown point definitions that rely heavily on group structure and equivariance, especially in dependent data settings.
- To formalize the intuitive notion that an estimator breaks down when it no longer conveys useful information about the uncontaminated data process.
- To develop a robust, information-theoretic breakdown criterion that avoids the 'small print' issues highlighted by Davies and Gather.
- To distinguish between robust estimators in time series, where classical breakdown points may be misleading, by quantifying information loss under contamination.
- To provide a framework applicable to asymptotic and continuous sample spaces, particularly for stationary time series models like AR(1).
Proposed method
- Introduces a new breakdown definition based on the measure of intersection between the badness set under contamination and the clean data set.
- Defines the badness set R∗(Zζk, Y′) as the range of estimator values over all possible uncontaminated samples Y′ ∈ Y.
- Uses Lebesgue measure µ to quantify the size of the intersection R∗(Zζk, Y′) ∩ R∗(0, Y′), with breakdown occurring when this measure is zero.
- Applies the definition to AR(1) models under additive i.i.d. outlier processes with contamination probability p and outlier magnitude ζ.
- Computes asymptotic breakdown points by analyzing the limiting behavior of the badness set as ζ → ∞.
- Employs explicit expressions for OLS, LMS, and deepest regression (DR) estimators to derive the badness sets and their limiting measures.
Experimental results
Research questions
- RQ1How can the breakdown point of an estimator be meaningfully defined in time series and dependent data settings where traditional definitions fail?
- RQ2To what extent does an estimator lose its ability to convey information about the uncontaminated process when contaminated by outliers?
- RQ3Why do classical breakdown point definitions based on group structure and equivariance fail to capture information loss in dependent data?
- RQ4What are the asymptotic breakdown points of OLS, LMS, and DR estimators in an AR(1) model under additive outlier contamination?
- RQ5Can a measure-theoretic approach based on badness set intersection provide a more robust and interpretable breakdown criterion than parameter-space-based definitions?
Key findings
- The OLS estimator for the AR(1) parameter θ has an asymptotic breakdown point of 0% under additive outlier contamination, as its badness set collapses to {0} when outlier size ζ → ∞.
- The LMS estimator has an asymptotic breakdown point of 22.1% in the AR(1) model, derived from the limiting behavior of its badness set under contamination.
- The deepest regression (DR) estimator has an asymptotic breakdown point of 50% in the AR(1) model, as its badness set collapses to {0} only when contamination probability p reaches 50%.
- The proposed breakdown definition successfully distinguishes between robust estimators in time series, revealing that LMS and DR have different breakdown thresholds despite both being 50% in simple linear regression.
- The definition is less dependent on group structure and equivariance than prior approaches, resolving concerns raised by Davies and Gather about 'void' definitions in non-equivariant settings.
- The method remains effective in asymptotic settings with continuous sample spaces, though it faces limitations when applied to finite samples or discrete parameter spaces.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.