[Paper Review] Covariance Estimation: Optimal Dimension-free Guarantees for Adversarial Corruption and Heavy Tails
This paper presents a robust covariance estimator that achieves optimal, dimension-free convergence rates in the operator norm under adversarial corruption and heavy-tailed distributions with only four moments. It provides sub-Gaussian tail bounds under weak norm equivalence assumptions, matching Gaussian performance while being resilient to outliers and scaling with effective rank rather than dimension.
We provide an estimator of the covariance matrix that achieves the optimal rate of convergence (up to constant factors) in the operator norm under two standard notions of data contamination: We allow the adversary to corrupt an $η$-fraction of the sample arbitrarily, while the distribution of the remaining data points only satisfies that the $L_{p}$-marginal moment with some $p \ge 4$ is equivalent to the corresponding $L_2$-marginal moment. Despite requiring the existence of only a few moments, our estimator achieves the same tail estimates as if the underlying distribution were Gaussian. As a part of our analysis, we prove a dimension-free Bai-Yin type theorem in the regime $p > 4$.
Motivation & Objective
- To develop a covariance estimator robust to adversarial corruption where up to an η-fraction of samples are arbitrarily corrupted.
- To achieve optimal convergence rates in the operator norm under only four moments of the distribution, avoiding logarithmic factors.
- To ensure dimension-free guarantees by scaling with the effective rank r(Σ) instead of the ambient dimension d.
- To provide high-probability bounds matching those of the Gaussian case under weak norm equivalence (Lp–L2) assumptions.
- To unify robustness to outliers and heavy-tailed behavior with optimal statistical rates without requiring log factors or strong distributional assumptions.
Proposed method
- The estimator is constructed via a median-of-means approach combined with robust mean estimation techniques tailored for sub-exponential and heavy-tailed distributions.
- It leverages a non-asymptotic, dimension-free Bai–Yin type theorem for p > 4, establishing operator norm convergence without logarithmic factors.
- The method uses a quantile-based trimming mechanism to handle outliers, ensuring stability under η-corruption.
- It applies a reduction from covariance estimation to mean estimation of squared inner products, using robust median-of-means estimation in the sub-exponential regime.
- The analysis relies on a norm equivalence assumption: (E|⟨X,v⟩|^q)^{1/q} ≤ κ(q)(E|⟨X,v⟩|^2)^{1/2} for 2 ≤ q ≤ p and p ≥ 4.
- The estimator is tuned via confidence level δ and corruption level η, achieving high-probability bounds that match Gaussian performance.
Experimental results
Research questions
- RQ1Can a covariance estimator achieve optimal operator norm convergence rates under adversarial corruption with only four moments?
- RQ2Can such an estimator provide dimension-free guarantees scaling with effective rank r(Σ) rather than dimension d?
- RQ3Can the estimator achieve sub-Gaussian tail bounds under weak Lp–L2 norm equivalence without assuming Gaussian or log-concave distributions?
- RQ4Is the η log(1/η) dependence on corruption level optimal in the sub-Gaussian regime?
- RQ5Can the estimator be constructed without logarithmic factors in the convergence rate, matching classical asymptotic results?
Key findings
- The proposed estimator achieves the optimal rate of convergence √(r(Σ)/N) in the operator norm, matching the classical Bai–Yin result, under only p ≥ 4 moments.
- The convergence rate is dimension-free and depends only on the effective rank r(Σ), not the ambient dimension d.
- The estimator provides sub-Gaussian tail bounds—matching the performance of the sample covariance under Gaussian assumptions—under weak Lp–L2 norm equivalence.
- The method achieves optimal dependence on the corruption level η, with a tight η log(1/η) term in the sub-Gaussian case, confirming the sharpness of the bound.
- The analysis establishes a non-asymptotic, dimension-free Bai–Yin type theorem for p > 4, extending classical results to heavy-tailed and corrupted settings.
- The estimator is robust to adversarial corruption: even with ηN corrupted points, the error remains bounded with high probability, provided N ≥ c(p)(r(Σ) + log(1/δ)).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.