[论文解读] Structured Recovery with Heavy-tailed Measurements: A Thresholding Procedure and Optimal Rates
本文提出了一种正则化阈值最小二乘估计器,用于从仅具有有限矩的重尾测量中恢复结构化信号。在最小矩假设下——线性链接函数下为(20+ε)阶矩,一般链接函数下为(4+ε)阶矩——该方法在高概率下实现了最优的样本复杂度和误差率,即使在传统次高斯假设不成立的情况下依然有效。
This paper introduces a general regularized thresholded least-square procedure estimating a structured signal $θ_*\in\mathbb{R}^d$ from the following observations: $y_i = f(\langle\mathbf{x}_i, θ_* angle, ξ_i),~i\in\{1,2,\cdots,N\}$, with i.i.d. heavy-tailed measurements $\{(\mathbf{x}_i,y_i)\}_{i=1}^N$. A general framework analyzing the thresholding procedure is proposed, which boils down to computing three critical radiuses of the bounding balls of the estimator. Then, we demonstrate these critical radiuses can be tightly bounded in the following two scenarios: (1) The link function $f(\cdot)$ is linear, i.e. $y = \langle\mathbf{x},θ_* angle + ξ$, with $θ_*$ being a sparse vector and $\{\mathbf{x}_i\}_{i=1}^N$ being general heavy-tailed random measurements with bounded $(20+ε)$-moments. (2) The function $f(\cdot)$ is arbitrary unknown (possibly discontinuous) and $\{\mathbf{x}_i\}_{i=1}^N$ are heavy-tailed elliptical random vectors with bounded $(4+ε)$-moments. In both scenarios, we show under these rather minimal bounded moment assumptions, such a procedure and corresponding analysis lead to optimal sample and error bounds with high probability in terms of the structural properties of $θ_*$.
研究动机与目标
- 弥合理论模型中假设测量为次高斯分布与现实世界中具有重尾噪声和异常值的数据之间的差距。
- 在仅存在有限矩(如4+ε或20+ε阶矩)的条件下,为结构化信号开发一种鲁棒估计框架。
- 在弱矩条件下,建立稀疏和一般结构化恢复的最优样本复杂度与误差界。
- 提出一种通用的阈值化程序,实现最优速率,而无需假设设计向量为各向同性次高斯分布。
提出的方法
- 提出一种正则化阈值最小二乘估计器,用于从重尾观测中恢复结构化信号。
- 引入一个通用框架,通过三个关键半径来界定参数空间中估计器的支撑范围。
- 利用矩假设来界定这些关键半径:线性链接函数下使用(20+ε)阶矩,任意未知链接函数下使用(4+ε)阶矩。
- 利用高斯宽度和经验过程技术,在弱矩条件下控制估计误差。
- 应用阈值化机制以抑制测量过程中重尾异常值的影响。
- 在弱矩假设下,利用集中不等式推导估计误差和样本复杂度的高概率界。
实验结果
研究问题
- RQ1当测量为重尾分布且仅具有有限矩时,能否实现结构化信号的最优恢复速率?
- RQ2在最小矩假设下,基于阈值化的正则化最小二乘程序是否能实现最优样本复杂度和误差界?
- RQ3能否在重尾设计下绕过或弱化限制性等距性质(RIP),转而采用基于矩的界?
- RQ4在不假设设计向量为各向同性次高斯分布的条件下,能否在稀疏恢复中实现最优速率?
- RQ5在一般(可能不连续)链接函数和重尾椭球设计下,误差和样本复杂度如何变化?
主要发现
- 对于线性链接函数和稀疏信号,在设计向量满足(20+ε)阶矩假设下,该方法实现了最优误差率 $\sqrt{\frac{s\log d}{N}}$ 的量级。
- 在任意未知链接函数的情况下,该方法仅在设计向量满足(4+ε)阶矩假设下即能达到最优误差界。
- 阈值化过程确保了即使在噪声或设计具有重尾特性时,估计器的误差仍以高概率有界。
- 通过基于矩的集中不等式,紧密界定了估计器的关键半径,从而在无需次高斯假设的前提下实现了最优速率。
- 样本复杂度与稀疏恢复的信息论下界 $N \gtrsim s\log d$ 一致,即使在弱矩条件下也成立。
- 分析表明,仅需满足RIP的下界即可实现恢复,且该条件可在远弱于次高斯性的矩假设下得到保证。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。