Skip to main content
QUICK REVIEW

[论文解读] Penalized Maximum Tangent Likelihood Estimation and Robust Variable Selection

Yichen Qin, Shaobo Li|arXiv (Cornell University)|Aug 17, 2017
Statistical Methods and Inference参考文献 28被引用 4
一句话总结

该论文提出了一种惩罚最大切线似然估计(penalized MTE),一种鲁棒的高维回归方法,通过调优参数 t 将最小 Kullback-Leibler 散度与 ℓ₂ 距离估计相结合。该方法实现了 √(ln d / n) 的最优收敛速度,表现出 oracle 性质,并在数据污染和高维设置下保持强性能。

ABSTRACT

We introduce a new class of mean regression estimators -- penalized maximum tangent likelihood estimation -- for high-dimensional regression estimation and variable selection. We first explain the motivations for the key ingredient, maximum tangent likelihood estimation (MTE), and establish its asymptotic properties. We further propose a penalized MTE for variable selection and show that it is $\sqrt{n}$-consistent, enjoys the oracle property. The proposed class of estimators consists penalized $\ell_2$ distance, penalized exponential squared loss, penalized least trimmed square and penalized least square as special cases and can be regarded as a mixture of minimum Kullback-Leibler distance estimation and minimum $\ell_2$ distance estimation. Furthermore, we consider the proposed class of estimators under the high-dimensional setting when the number of variables $d$ can grow exponentially with the sample size $n$, and show that the entire class of estimators (including the aforementioned special cases) can achieve the optimal rate of convergence in the order of $\sqrt{\ln(d)/n}$. Finally, simulation studies and real data analysis demonstrate the advantages of the penalized MTE.

研究动机与目标

  • 开发一种鲁棒的均值回归估计器,使其在模型污染和高维设置下仍保持高效性。
  • 解决传统惩罚似然方法(如 Lasso)在高维数据中对异常值敏感的问题。
  • 建立所提估计器的理论性质,包括一致性与 oracle 性质。
  • 在 d 随 n 指数增长的超高维设置下,证明该方法的最优性。
  • 通过切线似然函数,为 MLE、MTE 及其他估计器提供统一框架,使其成为特例。

提出的方法

  • 提出使用对数似然函数的截断泰勒展开定义的极大切线似然估计(MTE),即 lnₜ(u),作为标准对数似然的鲁棒替代。
  • MTE 估计器等价于求解加权似然方程,其中密度值较低的观测获得较小权重,从而增强鲁棒性。
  • 通过在 MTE 目标函数中添加非凸惩罚(如 SCAD 或自适应 Lasso),提出惩罚 MTE,实现变量选择。
  • 惩罚 MTE 同时最小化 Kullback-Leibler 散度与 ℓ₂ 距离,平衡鲁棒性与效率。
  • 该方法被证明等价于最小化 Kullback-Leibler 与 ℓ₂ 距离的混合形式,权重由调优参数 t 决定。
  • 理论分析表明,在固定维渐近下,惩罚 MTE 具备 √n 一致性与 oracle 性质;在超高维设置下,达到最优收敛速度 √(ln d / n)。

实验结果

研究问题

  • RQ1能否构建一种鲁棒回归估计器,在模型污染下仍保持高效性,并在高维设置中实现最优收敛速度?
  • RQ2惩罚 MTE 估计器在预测变量数量发散的高维线性回归中是否具备 oracle 性质?
  • RQ3与现有鲁棒估计器(如 LAD、Huber、Lasso)相比,所提出的 MTE 方法在变量选择准确性和预测性能方面表现如何?
  • RQ4当预测变量数量 d 随样本量 n 指数增长时,惩罚 MTE 的理论收敛速度是什么?
  • RQ5MTE 框架能否统一 MLE、最小距离估计与截尾似然估计器等现有估计器,形成单一鲁棒估计框架?

主要发现

  • 在固定维渐近下,惩罚 MTE 估计器实现 √n 一致性与 oracle 性质,确保正确的变量选择与渐近正态性。
  • 在 d 随 n 指数增长的超高维设置下,惩罚 MTE 达到最优收敛速度 √(ln d / n)。
  • 模拟研究显示,惩罚 MTE 在变量选择准确性和预测误差方面优于 Lasso、LAD-Lasso 和 Huber-Lasso,尤其在数据污染条件下表现更优。
  • 在波士顿房价数据集中,惩罚 MTE 选择了稳定且模型大小变异小的变量集合,预测性能优于其他方法。
  • 在 eQTL 数据集中,惩罚 MTE 实现了最低的均方预测误差(0.565)与最小的模型大小标准差(1.210),在 100 次划分中表现一致。
  • 该方法在 eQTL 数据中选择了四个探针集,其中三个也被 LASSO 选中,表明在低污染条件下与现有方法具有高度一致性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。